Methods and compositions for characterizing proteins
Fusion proteins with proximity labeling enzymes efficiently map transcriptional components in cancer cells, addressing the challenge of identifying short-lived interactions to understand cancer mechanisms and identify therapeutic targets.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- CAMBRIDGE GENETIX LTD
- Filing Date
- 2025-11-14
- Publication Date
- 2026-05-21
Smart Images

Figure IB2025061673_21052026_PF_FP_ABST
Abstract
Description
Attorney Docket No. 2013763-0005METHODSAND COMPOSITIONS FOR CHARACTERIZING PROTEINSCROSS REFERENCE TO RELATED APPLICATIONS
[0001] This application claims the benefit of U.S. Provisional Application No.63 / 721,176 filed November 15, 2024, the contents of which is hereby incorporated by reference in its entirety.BACKGROUND
[0002] Proteins are complex molecules that perform their functions solely or by interacting with other companion proteins and / or DNA or RNA complexes. Certain protein interactions implicate diseases, such as cancer. These interactions in living cells may be short lived (10‘9seconds), and thus difficult to identify and map. There exists a need to identify these protein interactions involved in processes including transcription, translation, replication, and protein degradation pathways, to provide better insight into potential drug targets in, for example, cancer.SUMMARY
[0003] Among other things, in some embodiments, the present disclosure provides fusion proteins comprising a protein of interest or a binding-agent that targets a protein of interest, that is able to, via proximity labeling (PL), identify interacting components involved in one or more cellular processes including transcription, translation, replication, and protein degradation pathways in various cell types. Interactions include protein-DNA, proteinprotein and protein-RNA interactions.
[0004] Fusion proteins and systems, as described herein, provide a new and efficient mechanism to understand and map transcriptional patterns and core transcriptional components. Such methods may be applied to identify core transcriptional components and understand transcriptional mechanisms implicated in cancer. Methods of mapping core transcriptional components in certain cancer cells may be used to understand metastasis and acquisition of drug resistance in certain cancers. Accordingly, fusion proteins and methods of using the same, as described herein, provide a way to rapidly accelerate our understanding of transcriptional mechanisms specific to cancer and increase discovery of therapeutic targets.Page 1 of 16513092399vlAttorney Docket No. 2013763-0005
[0005] Fusion proteins described herein include a protein of interest or a binding agent that binds to a protein of interest, a linker, and a proximity labeling enzyme. Methods in the present disclosure use proximity labeling to, among other things, map protein interactions with a protein of interest. In some embodiments, a protein of interest is a core transcriptional component (e.g., TATA-binding protein) and the fusion protein can identify and map actively transcribing genes, in particular cancer cells. In some embodiments, such information provides for the ability to track changes in a cancer cell’s transcriptional patterns, for example associated with metastasis and / or acquisition of drug resistance and / or responsiveness to therapy. Thus, in some embodiments, provided technologies permit assessment of how particular cancer cells metastasize, how they acquire drug resistance, and / or how they respond to one or more particular interventions (e.g., to identify, monitor, or otherwise characterize useful or effective interventions.
[0006] Fusion proteins, as described herein, can include a binding agent such as a nanobody, specific for a protein of interest, to provide for selective labeling of transcriptional components. For example, a nanobody utilized in a fusion protein as described herein may be designed so that the fusion protein, by way of the nanobody, selectively binds transcription components that are active.
[0007] Fusion proteins, as described herein, can be used to identify and map transcriptional components in various cell types, e.g., cancer cells or rapidly dividing cells. In some embodiments, fusion proteins, as described herein, can be used to identify and map transcriptional components in cell types that are not rapidly dividing or are terminally differentiated. In some embodiments, methods of identifying transcriptional components in a specific cell type (e.g., a cancer cell) include comparing transcriptional components identified from that specific cell to a different cell type (e.g., a cell isolated from a healthy tissue or a terminally differentiated cell) to identify transcriptional components specific to the cell population. Such methods can distinguish core transcriptional machinery involved in, e.g., cancer.
[0008] The present disclosure provides, among other things, a system comprising: a first fusion protein comprising first protein of interest (POI) and a first protein fragment, wherein the first protein fragment comprises a fragment of a functional indicator protein; and Page 2 of 16513092399vlAttorney Docket No. 2013763-0005a second fusion protein comprising a proximity labeling enzyme and a second protein fragment, wherein the second protein fragment comprises a fragment of the functional indicator protein.
[0009] In some embodiments, a system described herein further comprises a third fusion protein comprising: a second POI and a third protein fragment, wherein the third protein fragment comprises a fragment of the functional indicator protein. In some embodiments, a first fusion protein, second fusion protein, and third fusion protein, when in proximity to one another, associate to form a complex via the association of the first protein fragment, second protein fragment, and third protein fragment, thereby forming the functional indicator protein.
[0010] In some embodiments, (i) the first POI and the first protein fragment are connected via a linker; (ii) the proximity labelling enzyme and the second protein fragment are connected via linker; and / or (iii) the second POI and the third protein fragment are connected via a linker. In some embodiments, the first POI and the second POI are the same. In some embodiments, the first POI and the second POI are the different. In some embodiments, the proximity labeling enzyme is characterized by its ability to label proteins that interact with the first POI and / or second POI when they are in proximity therewith.
[0011] In some embodiments, the first POI and / or the second POI is a transcriptional machinery component. In some embodiments, the transcriptional machinery component is a transcriptional activator or a transcriptional repressor. In some embodiments, the transcriptional machinery component comprises TATA-binding protein (TBP). In some embodiments, when the system is delivered to a population of cells, the system is capable of labeling proteins that interact with the protein of interest.
[0012] In another aspect, the present disclosure provides a system comprising: (i) a first fusion protein comprising a POI and a first protein fragment, wherein the first protein fragment comprises a fragment of a functional indicator protein; (ii) a second fusion protein comprising a DNA-targeting moiety that binds to a target genomic locus and a second protein fragment, wherein the second protein fragment comprises another fragment of a functional indicator protein ; and (iii) a third fusion protein comprising a proximity labeling enzyme andPage 3 of 16513092399vlAttorney Docket No. 2013763-0005a third protein fragment, wherein the third protein fragment comprises another fragment of a functional indicator protein.
[0013] In some embodiments, the POI and the first protein fragment are connected via a linker; (ii) the DNA-targeting moiety and the second protein fragment are connected via linker; and / or (iii) the proximity labeling enzyme and the third protein fragment are connected via a linker.
[0014] In some embodiments, the proximity labeling enzyme is characterized by its ability to label proteins that interact with the POI and / or target genomic locus when they are in proximity therewith. In some embodiments, the first POI is a transcriptional machinery component. In some embodiments, the transcriptional machinery component is a transcriptional activator or a transcriptional repressor. In some embodiments, the transcriptional machinery component is TATA-binding protein (TBP).
[0015] In some embodiments, the DNA-targeting moiety comprises a dCas9 protein or a locked nucleic acid (LNA). In some embodiments, the target genomic locus comprises a TATA box sequence.
[0016] In some embodiments, the proximity labelling enzyme comprises a biotin ligase or a peroxidase. In some embodiments, the proximity labelling enzyme comprises an enzyme selected from a group consisting of: APEX, APEX2, Bio-ID, Turbo-ID, and mini Turbo-ID. In some embodiments, the proximity labelling enzyme comprises APEX2. In some embodiments, when the system is delivered to a population of cells, the system is capable of labeling proteins that interact with the protein of interest and / or the target genomic locus.
[0017] In some embodiments, the first fusion protein and the second fusion protein dimerize via the first protein fragment and the second protein fragment. In some embodiments, the first fusion protein, the second fusion protein, and the third fusion protein, when in proximity to one another (e.g., at the target genomic locus or in proximity to the POI), associate to form a complex via the association of the first protein fragment, second protein fragment, and third protein fragment, thereby forming the functional indicator protein.Page 4 of 16513092399vlAttorney Docket No. 2013763-0005
[0018] In some embodiments, when the first, second, and third protein fragments are not a functional indicator protein unless they are associated in a complex. In some embodiments, the indicator protein comprises a fluorescent protein. In some embodiments, the fluorescent protein comprises a green fluorescent protein (GFP). In some embodiments, the first protein fragment comprises a GFP 10, a GFP 11, or a GFPl-9r protein fragment. In some embodiments, the second protein fragment comprises a GFP 10, a GFP11, or a GFPl-9r protein fragment. In some embodiments, the third protein fragment comprises a GFP 10, a GFP11, or a GFPl-9r protein fragment. In some embodiments, the first protein fragment comprises GFP11. In some embodiments, the second protein fragment comprises GFP10. In some embodiments, the third protein fragment comprises GFPl-9r.
[0019] In another aspect, the present disclosure provides, a fusion protein comprising: a proximity labeling enzyme APEX2 and a GFP ( I -9r) protein, wherein the proximity labeling enzyme is connected to the GFP ( I -9r) protein via a linker.
[0020] In another aspect, the present disclosure provides a fusion protein comprising: a DNA-targeting moiety fused to a protein fragment GFPl-9r. In some embodiments, the DNA-targeting moiety comprises a dCas9 protein or a locked nucleic acid (LNA). In some embodiments, the DNA-targeting moiety and the first protein fragment are connected via a linker. In some embodiments, the DNA-targeting moiety targets a target genomic loci.
[0021] In another aspect, the present disclosure provides, a method of isolating a repressive or transcriptional activation complex associated with a target genomic locus, the method comprising: delivering the system described herein to a cell, maintaining the system under conditions and for a time sufficient such that one or more of the fusion proteins interacts with the target genomic locus; exposing the system to conditions under which the proximity labeling enzyme labels components in proximity with the target genomic locus; isolating the complex comprising the labeled component(s) comprising the repressive or transcriptional activation complex interacting with the target genomic locus from the system.
[0022] In another aspect, the present disclosure provides, a method of characterizing a protein that interacts with a target protein of interest or a target genomic locus of interest, the method comprising steps of: delivering the system described herein to a cell; maintaining the system under conditions and for a time sufficient such that one or more of the fusion proteins Page 5 of 16513092399vlAttorney Docket No. 2013763-0005interacts with the target protein of interest or target genomic locus of interest; exposing the system to conditions under which the proximity labeling enzyme labels components in proximity with the target protein of interest or target genomic locus; separating labeled component(s) from the system; and characterizing the labeled component(s).
[0023] In another aspect, the present disclosure provides, a method comprising a step of: characterizing a labeled interacting component that was generated by: delivering the system described herein to a cell; maintaining the system under conditions and for a time sufficient such that one or more of the fusion proteins interacts with the target genomic locus; exposing the system to conditions under which the proximity labeling enzyme labels components in proximity with the target protein of interest; separating labeled component(s) from the system; and characterizing the labeled component(s).
[0024] In some embodiments, the first fusion protein, the second fusion protein, and the third fusion protein co-localize at the target genomic locus. In some embodiments, the labelled cell can be visualized by fluorescence microscopy.
[0025] In some embodiments, the cell comprises a cancer cell. In some embodiments, the cell is a HEK293 cell. In some embodiments, the cell comprises a terminally differentiated cell. In some embodiments, the cell comprises a non-dividing cell. In some embodiments, the non-dividing cells is an oocyte.
[0026] In some embodiments, exposing the system to conditions under which the proximity labeling enzyme labels components in proximity with the target protein of interest or target genomic locus comprises adding biotin to the first and second systems. In some embodiments, the labeled component(s) comprise biotinylated proteins. In some embodiments, separating the labeled components comprises extracting protein from the system and contacting the extracted protein with an anti -biotin antibody or streptavidin beads. In some embodiments, characterizing the labeled component(s) comprises analyzing the component(s) using LC MS / MS analysis.
[0027] The present disclosure provides, in one aspect, a fusion protein comprising: (i) a protein of interest, which protein of interest is a transcriptional machinery component; (ii) a linker; and (iii) a proximity labeling enzyme; wherein the linker connects the protein of Page 6 of 16513092399vlAttorney Docket No. 2013763-0005interest to the proximity labeling enzyme; and wherein the proximity labeling enzyme is characterized by its ability to label proteins that interact with the protein of interest when they are in proximity therewith.
[0028] In some embodiments, the transcriptional machinery component is TATA-binding protein 2 (TBP2). In some embodiments, the protein of interest comprises an amino acid sequence that has at least 90% identity to SEQ ID NO: 29, or to a functional portion thereof, and / or that shares a characteristic sequence element therewith. In some embodiments, the protein of interest comprises an amino acid sequence according to SEQ ID NO: 29, or a functional portion thereof, and / or that shares a characteristic sequence element therewith. In some embodiments, the protein of interest is encoded by a nucleic acid sequence that has at least 90% identity to SEQ ID NO: 5. In some embodiments, the protein of interest is encoded by a nucleic acid sequence according to SEQ ID NO: 5.
[0029] In some embodiments, the linker comprises a length that is within a range of about lOnm-lOOnm. In some embodiments, the linker comprises an amino acid sequence that has at least 90% identity to SEQ ID NO: 24 [GGNNGGNNGGNNGG]. In some embodiments, the linker comprises an amino acid sequence according to SEQ ID NO: 24. In some embodiments, the linker is encoded by a nucleic acid sequence having at least 90% identity to SEQ ID NO: 4 [GGAGGCAATAACGGCGGAAACAATGGAGGCAACAATGGAGGC], In some embodiments, the linker is encoded by a nucleic acid sequence according to SEQ ID NO: 4 [GGAGGCAATAACGGCGGAAACAATGGAGGCAACAATGGAGGC], In some embodiments, the linker is connected to the N-terminus of the protein of interest. In some embodiments, the linker is connected to the C-terminus of the protein of interest.
[0030] In some embodiments, the proximity labelling enzyme comprises a biotin ligase or a peroxidase. In some embodiments, the proximity labelling enzyme comprises an enzyme selected from a group consisting of: APEX, APEX2, Bio-ID, Turbo-ID, and mini Turbo-ID. In some embodiments, the proximity labelling enzyme comprises APEX2. In some embodiments, the proximity labelling enzyme comprises Turbo-ID.
[0031] In some embodiments, the proximity labelling enzyme comprises an amino acid sequence that has at least 90% identity to SEQ ID NO: 31. In some embodiments, the Page 7 of 16513092399vlAttorney Docket No. 2013763-0005proximity labelling enzyme comprises an amino acid sequence according to SEQ ID NO: 31. In some embodiments, the proximity labelling enzyme is encoded by a nucleic acid sequence that has at least 90% identity to SEQ ID NO: 3. In some embodiments, the proximity labelling enzyme is encoded by a nucleic acid sequence according to SEQ ID NO: 3.
[0032] In another aspect, the present disclosure provides a vector comprising a nucleic acid sequence encoding a fusion protein described herein. In some embodiments, the vector is a lentiviral vector.
[0033] In some embodiments, a fusion protein described herein comprises an amino acid sequence having at least 85% identity to SEQ ID NO: 36. In some embodiments, the fusion protein comprises an amino acid sequence according to SEQ ID NO: 36. In some embodiments, the fusion protein is delivered to a population of cells, the fusion protein is capable of labeling proteins that interact with the protein of interest. In some embodiments, labeling proteins that interact with the protein of interest comprises labeling the proteins with biotin.
[0034] In some embodiments, the population of cells comprises cells isolated from healthy tissue. In some embodiments, the population of cells comprises cells isolated from tumor tissue or cells comprising a tumor or cancer cell line.
[0035] In one aspect, the present disclosure provides a fusion protein comprising: (i) a nanobody that specifically binds to a protein of interest, which protein of interest is a transcriptional machinery component; (ii) a linker; and (iii) a proximity labeling enzyme; wherein the linker connects the nanobody to the proximity labeling enzyme; and wherein the proximity labeling enzyme is characterized by its ability to label proteins that interact with the protein of interest when they are in proximity therewith.
[0036] In some embodiments, the nanobody comprises a modification-specific intracellular antibody (mintbody). In some embodiments, the protein of interest is RNA polymerase II (Pol II) or TBP (e.g., TBP2 and mediator complex). In some embodiments, the nanobody specifically binds to a Pol-II Ser 2 phosphorylation. In some embodiments, the nanobody does not specifically bind to a Pol-II Ser 5 phosphorylation. In some embodiments, the nanobody comprises an amino acid sequence having at least 90% identity Page 8 of 16513092399vlAttorney Docket No. 2013763-0005to SEQ ID NO: 30. In some embodiments, in the nanobody comprises an amino acid sequence according to SEQ ID NO: 30.
[0037] In some embodiments, the nanobody is a bispecific nanobody. In some embodiments, the linker comprises a length that is within a range of about lOnm-lOOnm.
[0038] In some embodiments, the linker comprises an amino acid sequence that has at least 90% identity to SEQ ID NO: 25 [DPPVAT], In some embodiments, the linker comprises an amino acid sequence according to SEQ ID NO: 25. In some embodiments, the linker is encoded by a nucleic acid sequence having at least 90% identity to SEQ ID NO: 8 [GACCCACCGGTCGCCACC] . In some embodiments, the linker is encoded by a nucleic acid sequence according to SEQ ID NO: 8 [GACCCACCGGTCGCCACC]. In some embodiments, the linker is connected to the N-terminus of the nanobody. In some embodiments, the linker is connected to the C-terminus of the nanobody.
[0039] In some embodiments, the proximity labelling enzyme comprises a biotin ligase or a peroxidase. In some embodiments, the proximity labelling enzyme comprises an enzyme selected from a group consisting of: APEX, APEX2, Bio-ID, Turbo-ID, and mini Turbo-ID. In some embodiments, the proximity labelling enzyme comprises APEX2. In some embodiments, the proximity labelling enzyme comprises Turbo-ID. In some embodiments, the proximity labelling enzyme comprises an amino acid sequence that has at least 90% identity to SEQ ID NO: 31. In some embodiments, the proximity labelling enzyme comprises an amino acid sequence according to SEQ ID NO: 31. In some embodiments, the proximity labelling enzyme is encoded by a nucleic acid sequence that has at least 90% identity to SEQ ID NO: 9. In some embodiments, the proximity labelling enzyme is encoded by a nucleic acid sequence according to SEQ ID NO: 9.
[0040] In some embodiments, the present disclosure provides a vector comprising a nucleic acid sequence encoding a fusion protein described herein. In some embodiments, the vector is a lentiviral vector.
[0041] In some embodiments, the fusion protein comprises an amino acid sequence having at least 85% identity to SEQ ID NO: 35. In some embodiments, the fusion protein comprises an amino acid sequence according to SEQ ID NO: 35. In some embodiments, the Page 9 of 16513092399vlAttorney Docket No. 2013763-0005fusion protein is delivered to a population of cells, the fusion protein is capable of labeling proteins that interact with the protein of interest. In some embodiments, labeling proteins that interact with the protein of interest comprises labeling the proteins with biotin. In some embodiments, the population of cells comprises cells isolated from healthy tissue. In some embodiments, the population of cells comprises cells isolated from tumor tissue or cells comprising a tumor or cancer cell line.
[0042] In another aspect, the present disclosure provides a complex comprising: a fusion protein associated with at least one labeled interacting component; wherein the fusion protein comprises (i) a protein of interest or a nanobody that specifically binds to a protein of interest, which protein of interest is a transcriptional machinery component; (ii) a linker; and (iii) a proximity labeling enzyme; wherein the linker connects the protein of interest or nanobody to the proximity labeling enzyme; and wherein the proximity labeling enzyme is characterized by its ability to label components that interact with the protein of interest when they are in proximity therewith.
[0043] In some embodiments, the complex is inside of a cell. In some embodiments, the labeled interacting component comprises a protein that interacts with the protein of interest. In some embodiments, the labeled interacting component comprises a transcriptional machinery component. In some embodiments, the cell is isolated from healthy tissue. In some embodiments, the cell is isolated from a tumor tissue.
[0044] In another aspect, the present disclosure provides, a method of characterizing a protein that interacts with a target protein of interest, the method comprising steps of: providing a system comprising: the fusion protein described herein; and components that interact with the target protein of interest; maintaining the system under conditions and for a time sufficient such that one or more of the components interacts with the target protein of interest: exposing the system to conditions under which the proximity labeling enzyme labels components in proximity with the target protein of interest; separating labeled component(s) from the system; and characterizing the labeled component(s).
[0045] In some embodiments, the system is or comprises a cell in which the fusion protein is present. In some embodiments, the system comprises a first and second system, wherein the first system is or comprises a cancer cell and the second system is or comprises a Page 10 of 16513092399vlAttorney Docket No. 2013763-0005comparable non-cancer cell. In some embodiments, the method further comprising the step of: comparing the separated labeled components from the first and second system, such that cancer-associated interacting components are characterized.
[0046] In one aspect, the present disclosure provides, a method comprising a step of: characterizing a labeled interacting component that was generated by: providing a system comprising: the fusion protein described herein; and components that interact with the target protein of interest; maintaining the system under conditions and for a time sufficient such that one or more of the components interacts with the target protein of interest: exposing the system to conditions under which the proximity labeling enzyme labels components in proximity with the target protein of interest; and separating labeled component(s) from the system; and characterizing the labeled component(s).
[0047] In some embodiments, the system is or comprises a cell in which the fusion protein is present. In some embodiments, the system comprises a first and second system, wherein the first system is or comprises a cancer cell and the second system is or comprises a comparable non-cancer cell.
[0048] In some embodiments, characterizing the labeled component(s) comprises comparing the separated labeled components from the first and second systems, such that cancer-associated interacting component(s) are characterized. In some embodiments, exposing the system to conditions under which the proximity labeling enzyme labels components in proximity with the target protein of interest comprises adding biotin to the first and second systems. In some embodiments, exposing the system to conditions under which the proximity labeling enzyme labels components in proximity with the target protein of interest further comprises adding hydrogen peroxide to the first and second systems after adding biotin.
[0049] In some embodiments, the labeled component(s) comprise biotinylated proteins. In some embodiments, separating the labeled components comprises extracting protein from the first and second systems and contacting the extracted protein with an antibiotin antibody or streptavidin beads. In some embodiments, characterizing the labeled component(s) comprises analyzing the component(s) using LC MS / MS analysis. In somePage 11 of 16513092399vlAttorney Docket No. 2013763-0005embodiments, characterizing the labeled component(s) comprises characterizing the identified proteins based on cell atlas location.
[0050] In some embodiments, the comparable non-cancer cell is a HEK293 cell. In some embodiments, the comparable non-cancer cell comprises a terminally differentiated cell. In some embodiments, the comparable non-cancer cell comprises a non-dividing cell. In some embodiments, the non-dividing cells is an oocyte. In some embodiments, the cancer cell comprises a cell of a neuroblastoma. In some embodiments, the cancer cell comprises a cell of a SH5Y5 cell line. In some embodiments, the cancer cell comprises a cell isolated from a tumor tissue .
[0051] In some embodiments, the tumor is a neuroblastoma or a neuroendocrine tumor. In some embodiments, the neuroendocrine tumor is a tumor resulting from adrenal cancer, a carcinoid tumor, a merkel cell carcinoma, a pancreatic neuroendocrine tumor, a paraganglioma, or a pheochromocytoma. In some embodiments, the cancer cell is a SH5Y 5 cell and the comparable non-cancer cells is a HEK293 cell.
[0052] In some embodiments, the cancer-associated interacting component is a mutant of Stab-001, and wherein Stab-001 is separated from the second system and is not separated from the first system. In some embodiments, when the Stab-001 mutant is expressed in SH5Y5 cells, terminal differentiation of the cells is induced. In some embodiments, the cancer cell is a cell isolated from a tumor or a tumor cell line and the comparable non-cancer cell is an oocyte or a terminally differentiated cell.
[0053] In some embodiments, a cancer-associated interacting component is TBP-002, wherein TBP-002 is separated from the second system but not from the first system. In some embodiments, when TBP-002 expression is induced in myoblast cells [C2C12] cells, rapid differentiation of the myoblasts is induced. In some embodiments, the cancer-associated interacting component is a transcriptional machinery component.
[0054] In another aspect, the present disclosure provides a method of treating a subject suffering from a cancer or tumor, the method comprising: administering to the subject a therapy that targets the cancer-associated interacting component described herein.BRIEF DESCRIPTION OF THE DRAWINGPage 12 of 16513092399vlAttorney Docket No. 2013763-0005
[0055] Figure 1 shows a schematic of these exemplary methods for identification of global transcriptional interactome through core transcriptional machinery proximity labeling using the fusion protein described above. Green (normal) and Red (Tumor cells) are derived from patient specific samples and then infected with a Lentivirus encoding hTBP-APEX and hTBP-TurboID enzymes. Stable clones are then selected through FACS or Puromycin to further be used for the interactome analysis. The cells are generally grown in their growth media, however in some cases to discover the drug resistance mechanisms, the cells expressing the TBP-APEX in that case can be grown in the presence of SILAC amino acids, which can provide a high temporal resolution of transcriptional response to a drug stimulus. The resultant proteins are identified using the mass spectrometry (LC MS-MS).
[0056] Figure 2 shows results from a Western Blot using an anti-TBP antibody and confirms that the TBP-APEX2 construct encodes a fusion protein in a dose-dependent manner. Figure 2A is a schematic diagram of the experimental strategy for testing the expression of TBP-APEX2 fusion in HEK293T cells. Figure 2B shows the expression of TBP-APEX2 fusion in HEK293T cells in a dose dependent manner.
[0057] Figure 3 shows experimental strategy and results from a localization analysis using Chromatin Immunoprecipitation. Figure 3A shows schematic diagram showing the experimental setup for performing immunostaining of HEK293T cells. Figure 3B shows the immunostaining DAPI (blue )m Anti -V5 (Green) and Anti-TBP (Red) after Dox induction. Anti-TBP antibody and Anti -V5 antibody shows the localization of TBP-APEX2 fusion proteins in the nucleus.
[0058] Figure 4 shows experimental strategy and results from ChlP-qPCR assays to map changes in expression of genes associated with TPB binding in cells contacted with the TPB-APEX2 fusion protein. Figure 4A shows a schematic diagram of the experimental strategy which involves contacting cells with the TBP-APEX2 complex and induction of cells with 2ug / ml of Dox for 48 hours, after which cells were subjected to lysis and ChlP-qPCR.Figures 4B and 4C shows by Anti-TBP Chromatin immunoprecipitation the relative percentage of DNA as compared to input material. Immunoprecipitated DNA was subjected to qPCRto specifically amplify TATA boxes in RPS9 and RAB5B in both induced (FigurePage 13 of 16513092399vlAttorney Docket No. 2013763-00054B) and uninduced (Figure 4C) conditions. Each experiment contained 3 replicates with 2xl06cells. Error bars shows SEM of n=3. *SEM p<0.05, **SEM p< 0.05 and *** p<0.02.
[0059] Figure 5 shows a schematic diagram showing the experimental strategy for transcriptional profding of the SH-5YSY cell line and normal human neuroblast cells.SILAC based proximity labelling was performed in the cells and subjected to Mass spectrometry.
[0060] Figure 6 shows characterization of proteins identified from proximity labeling of SH-5YSY cell line and normal human neuroblast cells with TPB-APEX2 fusion proteins.Figure 6A shows a silver stain gel (10% Tris glycine gel) of the biotinylated protein pull down from proximity labeling with TBP-APEX2 fusion protein in neuroblastoma cells. Figure 6B shows a scattered plot showing the differentially expressed genes among Neuroblastoma SH-5YSY cells (Red) and normal human neuroblast cells (Green). Data was normalized and fdtered with respect to the APEX only and -Dox controls. Genes in black are expressed in both normal and diseased cells, while genes in Red and Green are differentially expressed or silenced in diseased and normal cells respectively. Plot shows an average value plotted for each gene from n=9.
[0061] Figure 7 shows mapping of proteins interacting with TBP-APEX2 fusion protein in non-dividing cell of Xenopus oocytes. Figure 7A shows characterization maps that include the classification of the identified proteins / transcription. The grey plot (top) shown in Figure 7A shows proteins in non-transcription factor families, i.e., that are not directly involved in transcription. The orange plot (bottom) of Figure 7A shows the transcription factors. Figure 7B characterizes the transcription factors and non-transcription factors by their cell atlas location and their abundance in the GV of the oocytes is listed. Figure 7C shows the second phase of characterization, which identifies transcription factors according to their cell atlas location.
[0062] Figure 8 shows an experiment where mouse C2C12 cells were engineered to express TBP-002. Results showed that when C2C12 cells were induced to express TBP-002 (using doxycycline induction), C2C12 cells underwent differentiation in about 48 hours without any serum starvation or without any addition of signaling molecules like insulin compared to 7-10 days of C2C12 cells not engineered to express TBP-002.Page 14 of 16513092399vlAttorney Docket No. 2013763-0005
[0063] Figure 9 shows expression of differentiation markers in C2C12 cells that have been engineered to express TBP-002 using qPCR. The results show the early upregulation of the genes involved in cell cycle exit and for the stable expression of myogenic factors compared with cells that have not been engineered to express TBP-002.
[0064] Figure 10 shows the binding mechanism of an RNA-Pol-II-Ser2 mintbody that selectively labels the RNA-Pol- II elongation complex in normal and tumor cells for the identification of chromatin binding partner proteins.
[0065] Figure 11 shows the differential expression pattern of the proteins from neuroblastoma and neuroblast cells identified via proximity labeling with an RNA-Pol-II-Ser2 mintbody. LC MS / MS data was analyzed initially through max quant with the peptide count threshold of 25, and false discovery rate to 1%. Differentially expressed protein interactors were plotted on the fold change detection of X axis and -logP value on the Y axis. The data shown here represent 3 independent experiments (n=3).
[0066] FIG. 12 shows a schematic of an exemplary tripartite split GFP system (Figure adapted from S Castillo et al., 2023). In this exemplary system, Protein A and Protein B (exemplary interacting proteins) are fused with GFP Beta 10 and GFP Beta 11, respectively. GFP Beta 11 and GFP Beta 12 are 20 amino acid tags. Using this system, when the two proteins (A and B) are in close proximity or are interacting with one another, the GFP Beta 10 and GFP Beta 11 proteins dimerize, and when the system is supplemented with GFP l-9r, the GFP fragments complex to make a functional reconstituted GFP that can be visualized.
[0067] FIG. 13 shows a schematic diagram of interaction-induced proximity labelling coupled with fluorescence visualization. FIG. 13A shows Proteins A and B fused with GFP 10 and GFP 11 fragments, respectively, and a GFPl-9r (GFP detection fragment) fused with APEX2. Upon interaction of all three fragments, GFP is reconstituted and can be visualized (FIG. 13B). Proximity labelling of interacting components may be achieved by addition biotin phenol and hydrogen peroxide into the cultured cells or tissues.
[0068] FIG. 14 shows a schematic diagram showing the split GFP system with dCas9 as described above (FIG. 14A), and an exemplary method of site specific pull down and Page 15 of 16513092399vlAttorney Docket No. 2013763-0005labelling of DNA elements and their interacting components (e.g., transcriptional complexes) (FIG. 14B). Binding of Protein A to the target locus (i.e., a target DNA sequence) and tethered by dCas9 induces dimerization, which is then detected by GFP-APEX2 fusion protein. The GFP signal can be visualized by a fluorescent microscopy and a proximity labelling reaction can be induced by addition of hydrogen peroxide and biotin phenol (FIG.14B).
[0069] FIG. 15 shows a schematic of the components and functionality of exemplary TetOn and TetOff systems. FIG. 15A shows a construct which contains the Tet-element at the 5’ end followed by a mCherry reporter. FIG. 15B shows functionality of the described TetOn and TetOff systems as inducible gene expression systems that allow for precise temporal and spatial control of gene expression.
[0070] FIG. 16 shows a fragment map of the protein sequence of GFP 10 fusion with APEX2.
[0071] FIG. 17 shows an exemplary vector map of the fusion protein including DNMT3 A and a GFP 11 fragment.
[0072] FIG. 18A shows a schematic diagram of Xenopus oocytes used for the injection of dCas9 mRNA followed by the GV injection of FF-L DNA. Oocytes were then incubated further overnight before processing for ChlP. FIG. 18B shows qPCR mediated amplification of DNA pulled down through dCas9 ChlP. Amplification of Ebox and nonEbox control region was performed with sequence specific primers. Error Bars show the SEM (n=5) each experiment contained triplicates with each sample having 15 oocytes pooled together. ** (p<0.007) ***(P<0.001) Student t Test was used forthe statistical analysis.DEFINITIONS
[0073] In this application, unless otherwise clear from context, (i) the term “a” may be understood to mean “at least one”; (ii) the term “or” may be understood to mean “and / or”; (iii) the terms “comprising” and “including” may be understood to encompass itemized components or steps whether presented by themselves or together with one or more additional components or steps; and (iv) the terms “about” and “approximately” may be Page 16 of 16513092399vlAttorney Docket No. 2013763-0005understood to permit standard variation as would be understood by those of ordinary skill in the art; and (v) where ranges are provided, endpoints are included.
[0074] About: The term “about”, when used herein in reference to a value, refers to a value that is similar, in context to the referenced value. In general, those skilled in the art, familiar with the context, will appreciate the relevant degree of variance encompassed by “about” in that context. For example, in some embodiments, the term “about” may encompass a range ofvalues that within 25%, 20%, 19%, 18%, 17%, 16%, 15%, 14%, 13%, 12%, 11%, 10%, 9%, 8%, 7%, 6%, 5%, 4%, 3%, 2%, 1%, or less of the referred value.
[0075] Affinity. As is known in the art, “affinity” is a measure of the tightness with which two or more binding partners associate with one another. Those skilled in the art are aware of a variety of assays that can be used to assess affinity, and will furthermore be aware of appropriate controls for such assays. In some embodiments, affinity is assessed in a quantitative assay. In some embodiments, affinity is assessed over a plurality of concentrations (e.g., of one binding partner at a time). In some embodiments, affinity is assessed in the presence of one or more potential competitor entities (e.g., that might be present in a relevant - e.g., physiological - setting). In some embodiments, affinity is assessed relative to a reference (e.g., that has a known affinity above a particular threshold (a “positive control” reference) or that has a known affinity below a particular threshold (a “negative control” reference”). In some embodiments, affinity may be assessed relative to a contemporaneous reference; in some embodiments, affinity may be assessed relative to a historical reference. Typically, when affinity is assessed relative to a reference, it is assessed under comparable conditions.
[0076] Affinity matured" (or "affinity matured antibody”)'. as used herein, refers to an antibody (e.g., nanobody) with one or more alterations in one or more CDRs thereof which result an improvement in the affinity of the antibody for antigen, compared to a parent antibody which does not possess those alteration(s). In some embodiments, affinity matured antibodies will have nanomolar or even picomolar affinities for a target antigen. Affinity matured antibodies may be produced by any of a variety of procedures known in the art. Marks et al., BioTechnology 10:779-783 (1992) describes affinity maturation by VH and VL domain shuffling. Random mutagenesis of CDR and / or framework residues is described by: Barbas et al. Proc. Nat. Acad. Sci. U.S.A 91:3809-3813 (1994); Schier et al., Gene 169: 147-Page 17 of 16513092399vlAttorney Docket No. 2013763-0005155 (1995); Yelton et al., J. Immunol. 155: 1994-2004 (1995); Jackson et al., J. Immunol. 154(7): 3310-9 (1995); and Hawkins et al., J. Mol. Biol. 226:889-896 (1992).
[0077] Agent : In general, the term “agent”, as used herein, is used to refer to an entity (e.g., for example, a lipid, metal, nucleic acid, polypeptide, polysaccharide, small molecule, etc., or complex, combination, mixture or system [e.g., cell, tissue, organism] thereof), or phenomenon (e.g., heat, electric current or field, magnetic force or field, etc.). In appropriate circumstances, as will be clear from context to those skilled in the art, the term may be utilized to refer to an entity that is or comprises a cell or organism, or a fraction, extract, or component thereof. Alternatively or additionally, as context will make clear, the term may be used to refer to a natural product in that it is found in and / or is obtained from nature. In some instances, again as will be clear from context, the term may be used to refer to one or more entities that is man-made in that it is designed, engineered, and / or produced through action of the hand of man and / or is not found in nature. In some embodiments, an agent may be utilized in isolated or pure form; in some embodiments, an agent may be utilized in crude form. In some embodiments, potential agents may be provided as collections or libraries, for example, that may be screened to identify or characterize active agents within them.
[0078] Antibody: As used herein, the term “antibody” refers to a polypeptide that includes canonical immunoglobulin sequence elements sufficient to confer specific binding to a particular target antigen. As is known in the art, intact antibodies as produced in nature are approximately 150 kD tetrameric agents comprised of two identical heavy chain polypeptides (about 50 kD each) and two identical light chain polypeptides (about 25 kD each) that associate with each other into what is commonly referred to as a ‘Y-shaped” structure. Each heavy chain is comprised of at least four domains (each about 110 amino acids long)- an amino-terminal variable (VH) domain (located at the tips of the Y structure), followed by three constant domains: CHI, CH2, and the carboxy -terminal CH3 (located at the base of the Y’s stem). A short region, known as the “switch”, connects the heavy chain variable and constant regions. The “hinge” connects CH2 and CH3 domains to the rest of the antibody. Two disulfide bonds in this hinge region connect the two heavy chain polypeptides to one another in an intact antibody. Each light chain is comprised of two domains - an aminoterminal variable (VL) domain, followed by a carboxy-terminal constant (CL) domain, Page 18 of 16513092399vlAttorney Docket No. 2013763-0005separated from one another by another “switch”. Intact antibody tetramers are comprised of two heavy chain-light chain dimers in which the heavy and light chains are linked to one another by a single disulfide bond; two other disulfide bonds connect the heavy chain hinge regions to one another, so that the dimers are connected to one another and the tetramer is formed. Naturally-produced antibodies are also glycosylated, typically on the CH2 domain. Each domain in a natural antibody has a structure characterized by an “immunoglobulin fold” formed from two beta sheets (e.g., 3-, 4-, or 5-stranded sheets) packed against each other in a compressed antiparallel beta barrel. Each variable domain contains three hypervariable loops known as “complementarity determining regions” (CDR1, CDR2, and CDR3) and four somewhat invariant “framework” regions (FR1, FR2, FR3, and FR4). Those skilled in the art are familiar with technologies for identifying CDR and / or FR region sequences (e.g., using Kabat, Chothia, and / or IMGT methodologies), and furthermore appreciate that precise boundaries of CDR and / or FR elements defined using these different technologies may vary somewhat even when applied to the same sequence; regardless, those skilled in the art are able to recognize, when comparing two or more antibody sequences, whether the same or different CDRs are present. When natural antibodies fold, the FR regions form the beta sheets that provide the structural framework for the domains, and the CDR loop regions from both the heavy and light chains are brought together in three-dimensional space so that they create a single hypervariable antigen binding site located at the tip of the Y structure. The Fc region of naturally-occurring antibodies binds to elements of the complement system, and also to receptors on effector cells, including for example effector cells that mediate cytotoxicity. As is known in the art, affinity and / or other binding attributes of Fc regions for Fc receptors can be modulated through glycosylation or other modification. In some embodiments, antibodies produced and / or utilized in accordance with the present disclosure include glycosylated Fc domains, including Fc domains with modified or engineered such glycosylation. In some embodiments, antibodies produced and / or utilized in accordance with the present disclosure include one or more modifications on an Fc domain, e.g., an effector null mutation, e.g., a LALA, LAGA, FEGG, AAGG, or AAGA mutation. For purposes of the present disclosure, in certain embodiments, any polypeptide or complex of polypeptides that includes sufficient immunoglobulin domain sequences as found in natural antibodies can be referred to and / or used as an “antibody”, whether such polypeptide is naturally produced (e.g., generated by an organism reacting to an antigen), or produced by recombinant Page 19 of 16513092399vlAttorney Docket No. 2013763-0005engineering, chemical synthesis, or other artificial system or methodology. In some embodiments, an antibody is polyclonal; in some embodiments, an antibody is monoclonal. In some embodiments, an antibody has constant region sequences that are characteristic of dog, cat, mouse, rabbit, primate, or human antibodies. In some embodiments, antibody sequence elements are human, humanized, primatized, chimeric, etc., as is known in the art. Moreover, the term “antibody” as used herein, can refer in appropriate embodiments (unless otherwise stated or clear from context) to any of the art-known or developed constructs or formats for utilizing antibody structural and functional features in alternative presentation. For example, in some embodiments, an antibody utilized in accordance with the present invention is in a format selected from, but not limited to, intact IgA, IgG, IgE or IgM antibodies; bi- or multi- specific antibodies (e.g., Zybodies®, etc.); antibody fragments such as Fab fragments, Fab’ fragments, F(ab’)2 fragments, Fd’ fragments, Fd fragments, and isolated CDRs or sets thereof; single chain Fvs; polypeptide-Fc fusions; single domain antibodies, alternative scaffolds or antibody mimetics (e.g., anticalins, FN3 monobodies, DARPins, Affibodies, Affilins, Affimers, Affitins, Alphabodies, Avimers, Fynomers, Im7, VLR, VNAR, Trimab, CrossMab, Trident); nanobodies, mintbodies binanobodies, F(ab’)2, Fab’, di-sdFv, single domain antibodies, trifunctional antibodies, diabodies, and minibodies. In some embodiments, relevant formats may be or include: Adnectins®; Affibodies®;Affilins®; Anticalins®; Avimers®; BiTE®s; cameloid antibodies; Centyrins®; ankyrin repeat proteins or DARPINs®; dual-affinity re-targeting (DART) agents; Fynomers®; shark single domain antibodies such as IgNAR; immune mobilixing monoclonal T cell receptors against cancer (ImmTACs); KALBITOR®s; MicroProteins; Nanobodies® minibodies; masked antibodies (e.g., Probodies®); Small Modular ImmunoPharmaceuticals (“SMIPs™ ); single chain or Tandem diabodies (TandAb®); TCR-like antibodies;, Trans-bodies®;TrimerX®; VHHs. In some embodiments, an antibody may lack a covalent modification (e.g., attachment of a glycan) that it would have if produced naturally. In some embodiments, an antibody may contain a covalent modification (e.g., attachment of a glycan, a payload [e.g., a detectable moiety, a therapeutic moiety, a catalytic moiety, etc.], or other pendant group [e.g., poly-ethylene glycol, etc.]).
[0079] Antibody agent: As used herein, the term “antibody agent” refers to an agent that specifically binds to a particular antigen. In some embodiments, the term encompasses a Page 20 of 16513092399vlAttorney Docket No. 2013763-0005polypeptide or polypeptide complex that includes immunoglobulin structural elements sufficient to confer specific binding. For example, in some embodiments, an antibody agent is or comprises a polypeptide whose amino acid sequence includes one or more structural elements recognized by those skilled in the art as a complementarity determining region (CDR); in some embodiments an antibody agent is or comprises a polypeptide whose amino acid sequence includes at least one CDR (e.g., at least one heavy chain CDR and / or at least one light chain CDR) that is substantially identical to one found in a reference antibody. In some embodiments an included CDR is substantially identical to a reference CDR in that it is either identical in sequence or contains between 1-5 amino acid substitutions as compared with the reference CDR. In some embodiments an included CDR is substantially identical to a reference CDR in that it shows at least 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity with the reference CDR. In some embodiments an included CDR is substantially identical to a reference CDR in that it shows at least 96%, 96%, 97%, 98%, 99%, or 100% sequence identity with the reference CDR. In some embodiments an included CDR is substantially identical to a reference CDR in that at least one amino acid within the included CDR is deleted, added, or substituted as compared with the reference CDR but the included CDR has an amino acid sequence that is otherwise identical with that of the reference CDR. In some embodiments an included CDR is substantially identical to a reference CDR in that 1-5 amino acids within the included CDR are deleted, added, or substituted as compared with the reference CDR but the included CDR has an amino acid sequence that is otherwise identical to the reference CDR. In some embodiments an included CDR is substantially identical to a reference CDR in that at least one amino acid within the included CDR is substituted as compared with the reference CDR but the included CDR has an amino acid sequence that is otherwise identical with that of the reference CDR. In some embodiments an included CDR is substantially identical to a reference CDR in that 1-5 amino acids within the included CDR are deleted, added, or substituted as compared with the reference CDR but the included CDR has an amino acid sequence that is otherwise identical to the reference CDR. In some embodiments, an antibody agent is or comprises a polypeptide whose amino acid sequence includes structural elements recognized by those skilled in the art as an immunoglobulin variable domain. In some embodiments, an antibody agent in or comprises a polypeptide whose amino acid sequence includes structural elements recognized by those skilled in the art to correspond to CDRsl, 2,Page 21 of 16513092399vlAttorney Docket No. 2013763-0005and 3 of an antibody variable domain; in some such embodiments, an antibody agent in or comprises a polypeptide or set of polypeptides whose amino acid sequence(s) together include structural elements recognized by those skilled in the art to correspond to both heavy chain and light chain variable region CDRs, e.g., heavy chain CDRs 1, 2, and / or 3 and light chain CDRs 1, 2, and / or 3. In some embodiments, an antibody agent is a polypeptide protein having a binding domain which is homologous or largely homologous to an immunoglobulin-binding domain. In some embodiments, an antibody agent may be or comprise a polyclonal antibody preparation. In some embodiments, an antibody agent may be or comprise a monoclonal antibody preparation. In some embodiments, an antibody agent may include one or more constant region sequences that are characteristic of a particular organism, such as a camel, human, mouse, primate, rabbit, rat; in many embodiments, an antibody agent may include one or more constant region sequences that are characteristic of a human. In some embodiments, an antibody agent may include one or more sequence elements that would be recognized by one skilled in the art as a humanized sequence, a primatized sequence, a chimeric sequence, etc. In some embodiments, an antibody agent may be a canonical antibody (e.g., may comprise two heavy chains and two light chains). In some embodiments, an antibody agent may be in a format selected from, but not limited to, intact IgA, IgG, IgE or IgM antibodies; bi- or multi- specific antibodies (e.g., Zybodies®, etc.); antibody fragments such as Fab fragments, Fab’ fragments, F(ab’)2 fragments, Fd’ fragments, Fd fragments, and isolated CDRs or sets thereof; single chain Fvs; polypeptide-Fc fusions; single domain antibodies (e.g., shark single domain antibodies such as IgNAR or fragments thereof); cameloid antibodies; masked antibodies (e.g., Probodies®); Small Modular ImmunoPharmaceuticals (“SMIPs™ ); single chain or Tandem diabodies (TandAb®); VHHs; Anticalins®; Nanobodies® minibodies; mintbodies, BiTE®s; ankyrin repeat proteins or DARPINs®; Avimers®; DARTs; TCR-like antibodies; Adnectins®; Affilins®; Transbodies®; Affibodies®; TrimerX®; MicroProteins; Fynomers®, Centyrins®; and KALBITOR®s. In some embodiments, an antibody may lack a covalent modification (e.g., attachment of a glycan) that it would have if produced naturally. In some embodiments, an antibody may contain a covalent modification (e.g., attachment of a glycan, a payload [e.g., a detectable moiety, a therapeutic moiety, a catalytic moiety, etc.], or other pendant group [e.g., poly-ethylene glycol, etc.].Page 22 of 16513092399vlAttorney Docket No. 2013763-0005
[0080] Antibody fragment. As used herein, an “antibody fragment” refers to a portion of an antibody or antibody agent as described herein, and typically refers to a portion that includes an antigen-binding portion or variable region thereof. An antibody fragment may be produced by any means. For example, in some embodiments, an antibody fragment may be enzymatically or chemically produced by fragmentation of an intact antibody or antibody agent. Alternatively, in some embodiments, an antibody fragment may be recombinantly produced (i.e., by expression of an engineered nucleic acid sequence. In some embodiments, an antibody fragment may be wholly or partially synthetically produced. In some embodiments, an antibody fragment (particularly an antigen-binding antibody fragment) may have a length of at least about 50, 60, 70, 80, 90, 100, 110, 120, 130, 140, 150, 160, 170, 180, 190 amino acids or more, in some embodiments at least about 200 amino acids.
[0081] Antigen . The term “antigen”, as used herein, refers to an agent that elicits an immune response; and / or (ii) an agent that binds to a T cell receptor (e.g., when presented by an MHC molecule) or to an antibody. In some embodiments, an antigen elicits a humoral response (e.g., including production of antigen-specific antibodies); in some embodiments, an elicits a cellular response (e.g., involving T-cells whose receptors specifically interact with the antigen). In some embodiments, an antigen binds to an antibody and may or may not induce a particular physiological response in an organism. In general, an antigen may be or include any chemical entity such as, for example, a small molecule, a nucleic acid, a polypeptide, a carbohydrate, a lipid, a polymer (in some embodiments other than a biologic polymer [e.g., other than a nucleic acid or amino acid polymer) etc.. In some embodiments, an antigen is or comprises a polypeptide. In some embodiments, an antigen may be a protein of interest as described herein (e.g., a transcriptional machinery component or any other component that is to be proximity labeled). In some embodiments, an antigen is or comprises a glycan. Those of ordinary skill in the art will appreciate that, in general, an antigen may be provided in isolated or pure form, or alternatively may be provided in crude form (e.g., together with other materials, for example in an extract such as a cellular extract or other relatively crude preparation of an antigen-containing source). In some embodiments, antigens utilized in accordance with the present invention are provided in a crude form. In some embodiments, an antigen is a recombinant antigen.Page 23 of 16513092399vlAttorney Docket No. 2013763-0005
[0082] Approximately: As used herein, the term “approximately” or “about,” as applied to one or more values of interest, refers to a value that is similar to a stated reference value. In certain embodiments, the term “approximately” or “about” refers to a range of values that fall within 25%, 20%, 19%, 18%, 17%, 16%, 15%, 14%, 13%, 12%, 11%, 10%, 9%, 8%, 7%, 6%, 5%, 4%, 3%, 2%, 1%, or less in either direction (greater than or less than) of the stated reference value unless otherwise stated or otherwise evident from the context (except where such number would exceed 100% of a possible value).
[0083] Associated: Two events or entities are “associated” with one another, as that term is used herein, if the presence, level, degree, type and / or form of one is correlated with that of the other. For example, a particular entity (e.g., protein of interest, polypeptide, genetic signature, metabolite, microbe, etc.) is considered to be associated with a particular disease, disorder, or condition, if its presence, level and / or form correlates with incidence of, susceptibility to, severity of, stage of, etc. the disease, disorder, or condition (e.g., across a relevant population). In some embodiments, two or more entities are physically “associated” with one another if they interact, directly or indirectly, so that they are and / or remain in physical proximity with one another. In some embodiments, two or more entities that are physically associated with one another are covalently linked to one another; in some embodiments, two or more entities that are physically associated with one another are not covalently linked to one another but are non-covalently associated, for example by means of hydrogen bonds, van der Waals interaction, hydrophobic interactions, magnetism, and combinations thereof. In some embodiments, two or more entities are physically associated with each other such that they form a “complex”.
[0084] Binding. It will be understood that the term “binding”, as used herein, typically refers to a non-covalent association between or among two or more entities.“Direct” binding involves physical contact between entities or moieties; indirect binding involves physical interaction by way of physical contact with one or more intermediate entities. Binding between two or more entities can typically be assessed in any of a variety of contexts - including where interacting entities or moieties are studied in isolation or in the context of more complex systems (e.g., while covalently or otherwise associated with a carrier entity and / or in a biological system or cell). Binding between two entities may be considered “specific” if, under the conditions assessed, the relevant entities are more likely to associate with one another than with other available binding partners.Page 24 of 16513092399vlAttorney Docket No. 2013763-0005
[0085] Binding agent. In general, the term “binding agent” is used herein to refer to any entity that binds to a target of interest (e.g., a protein of interest) as described herein. In many embodiments, a binding agent of interest is one that binds specifically with its target in that it discriminates its target from other potential binding partners in a particular interaction context. In general, a binding agent may be or comprise an entity of any chemical class (e.g., polymer, non-polymer, small molecule, polypeptide, carbohydrate, lipid, nucleic acid, etc.). In some embodiments, a binding agent is a single chemical entity. In some embodiments, a binding agent is a complex of two or more discrete chemical entities associated with one another under relevant conditions by non-covalent interactions. For example, those skilled in the art will appreciate that in some embodiments, a binding agent may comprise a “generic” binding moiety (e.g., one of biotin / avidin / streptavidin and / or a class-specific antibody) and a “specific” binding moiety (e.g., an antibody or aptamers with a particular molecular target) that is linked to the partner of the generic biding moiety. In some embodiments, such an approach can permit modular assembly of multiple binding agents through linkage of different specific binding moieties with the same generic binding moiety partner. In some embodiments, binding agents are or comprise polypeptides (including, e.g., antibodies, antibody agents, or antibody fragments). In some embodiments, binding agents are or comprise small molecules. In some embodiments, binding agents are or comprise nucleic acids. In some embodiments, binding agents are aptamers. In some embodiments, binding agents are polymers; in some embodiments, binding agents are not polymers. In some embodiments, binding agents are non-polymeric in that they lack polymeric moieties. In some embodiments, binding agents are or comprise carbohydrates. In some embodiments, binding agents are or comprise lectins. In some embodiments, binding agents are or comprise peptidomimetics. In some embodiments, binding agents are or comprise scaffold proteins. In some embodiments, binding agents are or comprise mimeotopes. In some embodiments, binding agents are or comprise stapled peptides. In certain embodiments, binding agents are or comprise nucleic acids, such as DNA or RNA.
[0086] Cancer. The terms “cancer”, “malignancy”, “neoplasm”, “tumor”, and “carcinoma”, are used herein to refer to cells that exhibit relatively abnormal, uncontrolled, and / or autonomous growth, so that they exhibit an aberrant growth phenotype characterized by a significant loss of control of cell proliferation. In some embodiments, a tumor may be or comprise cells that are precancerous (e.g., benign), malignant, pre-metastatic, metastatic, Page 25 of 16513092399vlAttorney Docket No. 2013763-0005and / or non-metastatic. The present disclosure specifically identifies certain cancers to which its teachings may be particularly relevant. In some embodiments, a relevant cancer may be characterized by a solid tumor. In some embodiments, a relevant cancer may be characterized by a hematologic tumor. In general, examples of different types of cancers known in the art include, for example, hematopoietic cancers including leukemias, lymphomas (Hodgkin’s and non-Hodgkin’s), myelomas and myeloproliferative disorders; sarcomas, melanomas, adenomas, carcinomas of solid tissue, squamous cell carcinomas of the mouth, throat, larynx, and lung, liver cancer, genitourinary cancers such as prostate, cervical, bladder, uterine, and endometrial cancer and renal cell carcinomas, bone cancer, pancreatic cancer, skin cancer, cutaneous or intraocular melanoma, cancer of the endocrine system, cancer of the thyroid gland, cancer of the parathyroid gland, head and neck cancers, breast cancer, gastro-intestinal cancers and nervous system cancers, benign lesions such as papillomas, and the like.
[0087] CDR as used herein, refers to a complementarity determining region within an antibody variable region. There are three CDRs in each of the variable regions of the heavy chain and the light chain, which are designated CDR1, CDR2 and CDR3, for each of the variable regions. A “set of CDRs" or “CDR set” refers to a group of three or six CDRs that occur in either a single variable region capable of binding the antigen or the CDRs of cognate heavy and light chain variable regions capable of binding the antigen. Certain systems have been established in the art for defining CDR boundaries (e.g., Kabat, Chothia, etc.); those skilled in the art appreciate the differences between and among these systems and are capable of understanding CDR boundaries to the extent required to understand and to practice the claimed invention.
[0088] Cellular lysate: As used herein, the term “cellular lysate” or “cell lysate” refers to a fluid containing contents of one or more disrupted cells (i.e., cells whose membrane has been disrupted). In some embodiments, a cellular lysate includes both hydrophilic and hydrophobic cellular components. In some embodiments, a cellular lysate includes predominantly hydrophilic components; in some embodiments, a cellular lysate includes predominantly hydrophobic components. In some embodiments, a cellular lysate is a lysate of one or more cells selected from the group consisting of plant cells, microbial (e.g., bacterial or fungal) cells, animal cells (e.g., mammalian cells), human cells, and combinations thereof. In some embodiments, a cellular lysate is a lysate of one or more abnormal cells, Page 26 of 16513092399vlAttorney Docket No. 2013763-0005such as cancer cells. In some embodiments, a cellular lysate is a crude lysate in that little or no purification is performed after disruption of the cells; in some embodiments, such a lysate is referred to as a “primary” lysate. In some embodiments, one or more isolation or purification steps is performed on a primary lysate; however, the term “lysate” refers to a preparation that includes multiple cellular components and not to pure preparations of any individual component.
[0089] Comparable: As used herein, the term “comparable” refers to two or more agents, entities, situations, sets of conditions, etc., that may not be identical to one another but that are sufficiently similar to permit comparison therebetween so that one skilled in the art will appreciate that conclusions may reasonably be drawn based on differences or similarities observed. In some embodiments, comparable sets of conditions, circumstances, individuals, or populations are characterized by a plurality of substantially identical features and one or a small number of varied features. Those of ordinary skill in the art will understand, in context, what degree of identity is required in any given circumstance for two or more such agents, entities, situations, sets of conditions, etc. to be considered comparable. For example, those of ordinary skill in the art will appreciate that sets of circumstances, individuals, or populations are comparable to one another when characterized by a sufficient number and type of substantially identical features to warrant a reasonable conclusion that differences in results obtained or phenomena observed under or with different sets of circumstances, individuals, or populations are caused by or indicative of the variation in those features that are varied.
[0090] Composition: Those skilled in the art will appreciate that the term “composition” may be used to refer to a discrete physical entity that comprises one or more specified components. In general, unless otherwise specified, a composition may be of any form - e.g., gas, gel, liquid, solid, etc.
[0091] Determine: Many methodologies described herein include a step of “determining”. Those of ordinary skill in the art, reading the present specification, will appreciate that such “determining” can utilize or be accomplished through use of any of a variety of techniques available to those skilled in the art, including for example specific techniques explicitly referred to herein. In some embodiments, determining involves manipulation of a physical sample. In some embodiments, determining involves consideration and / or manipulation of data or information, for example utilizing a computer or Page 27 of 16513092399vlAttorney Docket No. 2013763-0005other processing unit adapted to perform a relevant analysis. In some embodiments, determining involves receiving relevant information and / or materials from a source. In some embodiments, determining involves comparing one or more features of a sample or entity to a comparable reference.
[0092] Domain: The term “domain” as used herein refers to a section or portion of an entity. In some embodiments, a “domain” is associated with a particular structural and / or functional feature of the entity so that, when the domain is physically separated from the rest of its parent entity, it substantially or entirely retains the particular structural and / or functional feature. Alternatively or additionally, a domain may be or include a portion of an entity that, when separated from that (parent) entity and linked with a different (recipient) entity, substantially retains and / or imparts on the recipient entity one or more structural and / or functional features that characterized it in the parent entity. In some embodiments, a domain is a section or portion of a molecule (e.g., a small molecule, carbohydrate, lipid, nucleic acid, or polypeptide). In some embodiments, a domain is a section of a polypeptide; in some such embodiments, a domain is characterized by a particular structural element (e.g., a particular amino acid sequence or sequence motif, a-helix character, [3-sheet character, coiled-coil character, random coil character, etc.), and / or by a particular functional feature (e.g., binding activity, enzymatic activity, folding activity, signaling activity, etc.).
[0093] Engineered: In general, the term “engineered” refers to the aspect of having been manipulated by the hand of man. For example, a polynucleotide is considered to be “engineered” when two or more sequences that are not linked together in that order in nature are manipulated by the hand of man to be directly linked to one another in the engineered polynucleotide and / or when a particular residue in a polynucleotide is non-naturally occurring and / or is caused through action of the hand of man to be linked with an entity or moiety with which it is not linked in nature. For example, in some embodiments described and / or utilized herein, an engineered polynucleotide comprises a regulatory sequence that is found in nature in operative association with a first coding sequence but not in operative association with a second coding sequence, is linked by the hand of man so that it is operatively associated with the second coding sequence. Comparably, a cell or organism is considered to be “engineered” if it has been subjected to a manipulation, so that it’s genetic, epigenetic, and / or phenotypic identity is altered relative to an appropriate reference cell such as otherwise identical cell that has not been so manipulated. In some embodiments, the Page 28 of 16513092399vlAttorney Docket No. 2013763-0005manipulation is or comprises a genetic manipulation, so that its genetic information is altered (e.g., new genetic material not previously present has been introduced, for example by transformation, mating, somatic hybridization, transfection, transduction, or other mechanism, or previously present genetic material is altered or removed, for example by substitution or deletion mutation, or by mating protocols). In some embodiments, an engineered cell is one that has been manipulated so that it contains and / or expresses a particular agent of interest (e.g., a protein, a nucleic acid, and / or a particular form thereof) in an altered amount and / or according to altered timing relative to such an appropriate reference cell. As is common practice and is understood by those in the art, progeny of an engineered polynucleotide or cell are typically still referred to as “engineered” even though the actual manipulation was performed on a prior entity.
[0094] Epitope: As used herein, the term “epitope” refers to a moiety that is specifically recognized by an immunoglobulin (e.g., antibody or receptor) binding component. In some embodiments, an epitope is comprised of a plurality of chemical atoms or groups on an antigen. In some embodiments, such chemical atoms or groups are surface-exposed when the antigen adopts a relevant three-dimensional conformation. In some embodiments, such chemical atoms or groups are physically near to each other in space when the antigen adopts such a conformation. In some embodiments, at least some such chemical atoms are groups are physically separated from one another when the antigen adopts an alternative conformation (e.g., is linearized).
[0095] Expression : As used herein, the term “expression” of a nucleic acid sequence refers to the generation of any gene product from the nucleic acid sequence. In some embodiments, a gene product can be a transcript. In some embodiments, a gene product can be a polypeptide. In some embodiments, expression of a nucleic acid sequence involves one or more of the following: (1) production of an RNA template from a DNA sequence (e.g. , by transcription); (2) processing of an RNA transcript (e.g., by splicing, editing, etc.); (3) translation of an RNA into a polypeptide or protein; and / or (4) post-translational modification of a polypeptide or protein.
[0096] Framework" or "framework region, as used herein, refers to the sequences of a variable region minus the CDRs. Because a CDR sequence can be determined by different systems, likewise a framework sequence is subject to correspondingly different interpretations. The six CDRs divide the framework regions on the heavy and light chains Page 29 of 16513092399vlAttorney Docket No. 2013763-0005into four sub-regions (FR1, FR2, FR3 and FR4) on each chain, in which CDR1 is positioned between FR1 and FR2, CDR2 between FR2 and FR3, and CDR3 between FR3 and FR4. Without specifying the particular sub-regions as FR1, FR2, FR3 or FR4, a framework region, as referred by others, represents the combined FRs within the variable region of a single, naturally occurring immunoglobulin chain. As used herein, a FR represents one of the four sub-regions, FR1, for example, represents the first framework region closest to the amino terminal end of the variable region and 5' with respect to CDR1, and FRs represents two or more of the sub-regions constituting a framework region.
[0097] Functional: As used herein, a “functional” biological molecule is a biological molecule in a form in which it exhibits a property and / or activity by which it is characterized.
[0098] Fragment: A “fragment” of a material or entity as described herein has a structure that includes a discrete portion of the whole, but lacks one or more moieties found in the whole. In some embodiments, a fragment consists of such a discrete portion. In some embodiments, a fragment consists of or comprises a characteristic structural element or moiety found in the whole. In some embodiments, a polymer fragment comprises or consists of at least 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 110, 120, 130, 140, 150, 160, 170, 180, 190, 200, 210, 220, 230, 240, 250, 275, 300, 325, 350, 375, 400, 425, 450, 475, 500 or more monomeric units (e.g., residues) as found in the whole polymer. In some embodiments, a polymer fragment comprises or consists of at least about 5%, 10%, 15%, 20%, 25%, 30%, 25%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% or more of the monomeric units (e.g., residues) found in the whole polymer. The whole material or entity may in some embodiments be referred to as the “parent” of the fragment.
[0099] Fusion Protein : As used herein, the term “fusion protein” refers to a fusion including at least two segments. Typically, a protein containing at least two such segments is considered to be a fusion protein if the two segments are moieties that (1) are not included in nature in the same peptide, and / or (2) have not previously been linked to one another in a single polypeptide, and / or (3) have been linked to one another through action of the hand of man. In some embodiments, a fusion protein described herein comprises a protein of interest or binding agent targeting a protein of interest, a linker, and a proximity labeling enzyme. A fusion protein as described herein may also contain other sequences such as targeting or localization sequences and / or tag sequences.Page 30 of 16513092399vlAttorney Docket No. 2013763-0005
[0100] Gene. As used herein, the term “gene” refers to a DNA sequence in a chromosome that codes for a product (e.g., an RNA product and / or a polypeptide product). In some embodiments, a gene includes coding sequence (i.e., sequence that encodes a particular product); in some embodiments, a gene includes non-coding sequence. In some particular embodiments, a gene may include both coding (e.g., exonic) and non-coding (e.g., intronic) sequences. In some embodiments, a gene may include one or more regulatory elements that, for example, may control or impact one or more aspects of gene expression (e.g., cell-type-specific expression, inducible expression, etc.).
[0101] Genome. As used herein, the term “genome” refers to the total genetic information carried by an individual organism or cell, represented by the complete DNA sequences of its chromosomes.
[0102] High affinity binding. The term “high affinity binding”, as used herein refers to a high degree of tightness with which a particular ligand binds to its partner. Affinities can be measured by any available method, including those known in the art. In some embodiments, binding is considered to be high affinity if the Kd is about 500 pM or less (e.g., below about 400 pM, about 300 pM, about 200 pM, about 100 pM, about 90 pM, about 80 pM, about 70 pM, about 60 pM, about 50 pM, about 40 pM, about 30 pM, about 20 pM, about 10 pM, about 5 pM, about 4 pM, about 3 pM, about 2 pM, etc.) in binding assays. In some embodiments, binding is considered to be high affinity if the affinity is stronger (e.g., the Ka is lower) for a polypeptide of interest than for a selected reference polypeptide. In some embodiments, binding is considered to be high affinity if the ratio of the Kd for a polypeptide of interest to the Kd for a selected reference polypeptide is 1 : 1 or less (e.g., 0.9:1, 0.8:1, 0.7:1, 0.6:1, 0.5:1. 0.4:1, 0.3:1, 0.2:1, 0.1:1, 0.05:1, 0.01:1, or less). In some embodiments, binding is considered to be high affinity if the Kd for a polypeptide of interest is about 100% or less (e.g., about 99%, about 98%, about 97%, about 96%, about 95%, about 90%, about 85%, about 80%, about 75%, about 70%, about 65%, about 60%, about 55%, about 50%, about 45%, about 40%, about 35%, about 30%, about 25%, about 20%, about 15%, about 10%, about 5%, about 4%, about 3%, about 2%, about 1% or less) of the Kd for a selected reference polypeptide.
[0103] Homology . As used herein, the term “homology” refers to overall relatedness between polymeric molecules, e.g., between nucleic acid molecules (e.g., DNA molecules and / or RNA molecules) and / or between polypeptide molecules. In some embodiments, Page 31 of 16513092399vlAttorney Docket No. 2013763-0005polymeric molecules are considered to be “substantially homologous” to one another if their sequences are at least 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, or 99% homologous, meaning that identical or homologous residues are present in corresponding positions of both molecules. Calculation of percent homology of two nucleic acid or polypeptide sequences, for example, can be performed by aligning two sequences for optimal comparison purposes (e.g., gaps can be introduced in one or both of a first and a second sequences for optimal alignment and non-identical sequences can be disregarded for comparison purposes). In some embodiments, length of a sequence aligned for comparison purposes is at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, or substantially 100% of length of a reference sequence; residues at corresponding positions are then compared. When a position in the first sequence is occupied by the same residue (e.g., nucleotide or amino acid) as a corresponding position in the second sequence, then the two molecules (i.e., first and second) are identical at that position. When a position in the first sequence is occupied by the same residue or by a structurally and / or functionally related residue (as will be understood by those skilled in the art, in context), then the two molecules are considered “homologous” at that position.Percent homology between two sequences is a function of the number of homologous positions shared by the two sequences being compared, taking into account the number of gaps, and the length of each gap, which needs to be introduced for optimal alignment of the two sequences. Comparison of sequences and determination of percent homology between two sequences can be accomplished using a mathematical algorithm. For example, percent homology between two nucleotide sequences can be determined using the algorithm of Meyers and Miller (CABIOS, 1989, 4: 11-17, which is herein incorporated by reference in its entirety), which has been incorporated into the ALIGN program (version 2.0).
[0104] Host cell, as used herein, refers to a cell into which exogenous DNA (recombinant or otherwise) has been introduced. Persons of skill upon reading this disclosure will understand that such terms refer not only to the particular subject cell, but also to the progeny of such a cell. Because certain modifications may occur in succeeding generations due to either mutation or environmental influences, such progeny may not, in fact, be identical to the parent cell, but are still included within the scope of the term “host cell” as used herein. In some embodiments, host cells include prokaryotic and eukaryotic cells selected from any of the Kingdoms of life that are suitable for expressing an exogenous DNA Page 32 of 16513092399vlAttorney Docket No. 2013763-0005(e.g., a recombinant nucleic acid sequence). Exemplary cells include those of prokaryotes and eukaryotes (single-cell or multiple-cell), bacterial cells (e.g., strains of / '.', coli, Bacillus spp., Streptomyces spp., etc.), mycobacteria cells, fungal cells, yeast cells (e.g., .S', cerevisiae, S. pombe, P. pastoris, P. methanolica, etc.), plant cells, insect cells (e.g., SF-9, SF-21, baculovirus-infected insect cells, Trichoplusia ni, etc.), non-human animal cells, human cells, or cell fusions such as, for example, hybridomas or quadromas. In some embodiments, the cell is a human, monkey, ape, hamster, rat, or mouse cell. In some embodiments, the cell is eukaryotic and is selected from the following cells: CHO (e.g., CHO KI, DXB-1 1 CHO, Veggie-CHO), COS (e.g., COS-7), retinal cell, Vero, CV1, kidney (e.g., HEK293, 293 EBNA, MSR 293, MDCK, HaK, BHK), HeLa, HepG2, WI38, MRC 5, Colo205, HB 8065, HL-60, (e.g., BHK21), Jurkat, Daudi, A431 (epidermal), CV-1, U937, 3T3, L cell, C127 cell, SP2 / 0, NS-0, MMT 060562, Sertoli cell, BRL 3 A cell, HT1080 cell, myeloma cell, tumor cell, a cancer cell, a dendritic cell, and a cell line derived from an aforementioned cell. In some embodiments, the cell comprises one or more viral genes.
[0105] Human antibody: as used herein, is intended to include antibodies having variable and constant regions generated (or assembled) from human immunoglobulin sequences. In some embodiments, antibodies (or antibody components) may be considered to be “human” even though their amino acid sequences include residues or elements not encoded by human germline immunoglobulin sequences (e.g., include sequence variations, for example that may (originally) have been introduced by random or site -specific mutagenesis in vitro or by somatic mutation in vivo), for example in one or more CDRs and in particular CDR3.
[0106] Humanized: as is known in the art, the term "humanized" is commonly used to refer to antibodies (or antibody components) whose amino acid sequence includes VH and VL region sequences from a reference antibody raised in a non-human species (e.g., a mouse), but also includes modifications in those sequences relative to the reference antibody intended to render them more "human-like" , i.e., more similar to human germline variable sequences. In some embodiments, a "humanized" antibody (or antibody component) is one that immunospecifically binds to an antigen of interest and that has a framework (FR) region having substantially the amino acid sequence as that of a human antibody, and a complementary determining region (CDR) having substantially the amino acid sequence as that of a non-human antibody. A humanized antibody comprises substantially all of at least Page 33 of 16513092399vlAttorney Docket No. 2013763-0005one, and typically two, variable domains (Fab, Fab', F(ab')2, FabC, Fv) in which all or substantially all of the CDR regions correspond to those of a non-human immunoglobulin (i.e., donor immunoglobulin) and all or substantially all of the framework regions are those of a human immunoglobulin consensus sequence. In some embodiments, a humanized antibody also comprises at least a portion of an immunoglobulin constant region (Fc), typically that of a human immunoglobulin constant region. In some embodiments, a humanized antibody contains both the light chain as well as at least the variable domain of a heavy chain. The antibody also may include a CHI, hinge, CH2, CH3, and, optionally, a CH4 region of a heavy chain constant region. In some embodiments, a humanized antibody only contains a humanized VL region. In some embodiments, a humanized antibody only contains a humanized VH region. In some certain embodiments, a humanized antibody contains humanized VH and VL regions.
[0107] Identity . As used herein, the term “identity” refers to overall relatedness between polymeric molecules, e.g., between nucleic acid molecules (e.g., DNA molecules and / or RNA molecules) and / or between polypeptide molecules. In some embodiments, polymeric molecules are considered to be “substantially identical” to one another if their sequences are at least 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, or 99% identical. Calculation of percent identity of two nucleic acid or polypeptide sequences, for example, can be performed by aligning two sequences for optimal comparison purposes (e.g., gaps can be introduced in one or both of a first and a second sequences for optimal alignment and non-identical sequences can be disregarded for comparison purposes). In some embodiments, length of a sequence aligned for comparison purposes is at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, or substantially 100% of length of a reference sequence; residues at corresponding positions are then compared. When a position in the first sequence is occupied by the same residue (e.g., nucleotide or amino acid) as a corresponding position in the second sequence, then the two molecules (i.e., first and second) are identical at that position. Percent identity between two sequences is a function of the number of identical positions shared by the two sequences being compared, taking into account the number of gaps, and the length of each gap, which needs to be introduced for optimal alignment of the two sequences. Comparison of sequences and determination of percent identity between two sequences can be accomplished using a mathematical algorithm. For example, percent Page 34 of 16513092399vlAttorney Docket No. 2013763-0005identity between two nucleotide sequences can be determined using the algorithm of Meyers and Miller (CABIOS, 1989, 4: 11-17, which is herein incorporated by reference in its entirety), which has been incorporated into the ALIGN program (version 2.0). In some embodiments, nucleic acid sequence comparisons made with the ALIGN program use a PAM120 weight residue table, a gap length penalty of 12 and a gap penalty of 4.
[0108] “Improve, ” “increase” , “inhibit” or “reduce”: As used herein, the terms “improve”, “increase”, “inhibit’, “reduce”, or grammatical equivalents thereof, indicate values that are relative to a baseline or other reference measurement. In some embodiments, an appropriate reference measurement may be or comprise a measurement in a particular system (e.g., in a single individual) under otherwise comparable conditions absent presence of (e.g., prior to and / or after) a particular agent or treatment, or in presence of an appropriate comparable reference agent. In some embodiments, an appropriate reference measurement may be or comprise a measurement in comparable system known or expected to respond in a particular way, in presence of the relevant agent or treatment.
[0109] Interacting Component: The term “interacting component” as used herein refers to a component that interacts or is associated with a protein of interest. In some embodiments, a protein of interest is a protein associated with active transcriptional machinery that is specific to a certain cancer cell and interacting components are cancer-associated interacting components. In some embodiments, interacting components are identified by proximity labelling.
[0110] In vitro: The term “in vitro” as used herein refers to events that occur in an artificial environment, e.g., in a test tube or reaction vessel, in cell culture, etc., rather than within a multi-cellular organism.[oni] In vivo: as used herein refers to events that occur within a multi-cellular organism, such as a human and a non-human animal. In the context of cell-based systems, the term may be used to refer to events that occur within a living cell (as opposed to, for example, in vitro systems).
[0112] KD: as used herein, refers to the dissociation constant of a binding agent (e.g., an antibody or binding component thereof) from a complex with its partner (e.g., the epitope to which the antibody or binding component thereof binds).
[0113] Linker: as used herein, is used to refer to that portion of a multi -element agent that connects different elements to one another (e.g., in a fusion protein). For example, those Page 35 of 16513092399vlAttorney Docket No. 2013763-0005of ordinary skill in the art appreciate that a polypeptide whose structure includes two or more functional or organizational domains often includes a stretch of amino acids between such domains that links them to one another. In some embodiments, a polypeptide comprising a linker element has an overall structure of the general form S1-L-S2, wherein SI and S2 may be the same or different and represent two domains associated with one another by the linker. In some embodiments, a polypeptide linker is at least 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100 or more amino acids in length. In some embodiments, a linker is characterized in that it tends not to adopt a rigid three-dimensional structure, but rather provides flexibility to the polypeptide. A variety of different linker elements that can appropriately be used when engineering polypeptides (e.g., fusion proteins) known in the art (see e.g., Holliger, P., et al. (1993) Proc. Natl. Acad. Sci. USA 90:6444-6448; Poljak, R. J., et al. (1994) Structure 2: 1 121-1123).
[0114] Mutant: As used herein, the term “mutant” refers to an entity that shows significant structural identity with a reference entity but differs structurally from the reference entity in the presence or level of one or more chemical moieties as compared with the reference entity. In many embodiments, a mutant also differs functionally from its reference entity. In general, whether a particular entity is properly considered to be a “mutant” of a reference entity is based on its degree of structural identity with the reference entity. As will be appreciated by those skilled in the art, any biological or chemical reference entity has certain characteristic structural elements. A mutant, by definition, is a distinct chemical entity that shares one or more such characteristic structural elements. To give but a few examples, a small molecule may have a characteristic core structural element (e.g., a macrocycle core) and / or one or more characteristic pendent moieties so that a mutant of the small molecule is one that shares the core structural element and the characteristic pendent moieties but differs in other pendent moieties and / or in types of bonds present (single vs double, E vs Z, etc.) within the core, a polypeptide may have a characteristic sequence element comprised of a plurality of amino acids having designated positions relative to one another in linear or three-dimensional space and / or contributing to a particular biological function, a nucleic acid may have a characteristic sequence element comprised of a plurality of nucleotide residues having designated positions relative to another in linear or three-dimensional space. For example, a mutant polypeptide may differ from a reference polypeptide as a result of one or more Page 36 of 16513092399vlAttorney Docket No. 2013763-0005differences in amino acid sequence and / or one or more differences in chemical moieties (e.g., carbohydrates, lipids, etc.) covalently attached to the polypeptide backbone. In some embodiments, a mutant polypeptide shows an overall sequence identity with a reference polypeptide that is at least 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, or 99%. Alternatively or additionally, in some embodiments, a mutant polypeptide does not share at least one characteristic sequence element with a reference polypeptide. In some embodiments, the reference polypeptide has one or more biological activities. In some embodiments, a mutant polypeptide shares one or more of the biological activities of the reference polypeptide. In some embodiments, a mutant polypeptide lacks one or more of the biological activities of the reference polypeptide. In some embodiments, a mutant polypeptide shows a reduced level of one or more biological activities as compared with the reference polypeptide.
[0115] Mintbody: As used herein, a “mintbody” refers to a “modification-specific intracellular antibody” that functions as a genetically encoded probe. A mintbody may include an antibody fragment such as single-chain variable fragment scFv or a nanobody tagged with a fluorescent protein (see e.g., Sato et al., J. of Mol. Biol, 428(20): 2013). In some embodiments, a mintbody is capable of being expressed and functional in the cytoplasm of a host cell. In some embodiments, a mintbody comprises a human antibody fragment. In some embodiments, a mintbody comprises a heavy chain variable domain and / or a light chain variable domain of an antibody.
[0116] Nanobody: As used herein, a “nanobody” or “single -domain antibody” (“sdAb”) is an antibody fragment that includes a single antibody variable domain. In some embodiments, a variable domain is a heavy chain variable domain. In some embodiments, a variable domain is a light chain variable domain. In some embodiments, an antibody variable domain is engineered from a human antibody variable heavy chain domain (human “VH”). In some embodiments, an antibody variable domain is engineered from camelid variable heavy chain (“VHH”).
[0117] Nucleic acid. As used herein, in its broadest sense, refers to any compound and / or substance that is or can be incorporated into an oligonucleotide chain. In some embodiments, a nucleic acid is a compound and / or substance that is or can be incorporated into an oligonucleotide chain via a phosphodiester linkage. As will be clear from context, in some embodiments, "nucleic acid" refers to an individual nucleic acid residue (e.g., a Page 37 of 16513092399vlAttorney Docket No. 2013763-0005nucleotide and / or nucleoside); in some embodiments, "nucleic acid" refers to an oligonucleotide chain comprising individual nucleic acid residues. In some embodiments, a "nucleic acid" is or comprises RNA; in some embodiments, a "nucleic acid" is or comprises DNA. In some embodiments, a nucleic acid is, comprises, or consists of one or more natural nucleic acid residues. In some embodiments, a nucleic acid is, comprises, or consists of one or more nucleic acid analogs. In some embodiments, a nucleic acid analog differs from a nucleic acid in that it does not utilize a phosphodiester backbone. For example, in some embodiments, a nucleic acid is, comprises, or consists of one or more "peptide nucleic acids", which are known in the art and have peptide bonds instead of phosphodiester bonds in the backbone, are considered within the scope of the present invention. Alternatively or additionally, in some embodiments, a nucleic acid has one or more phosphorothioate and / or 5'-N-phosphoramidite linkages rather than phosphodiester bonds. In some embodiments, a nucleic acid is, comprises, or consists of one or more natural nucleosides (e.g., adenosine, thymidine, guanosine, cytidine, uridine, deoxyadenosine, deoxythymidine, deoxy guanosine, and deoxy cytidine). In some embodiments, a nucleic acid is, comprises, or consists of one or more nucleoside analogs (e.g., 2-aminoadenosine, 2-thiothymidine, inosine, pyrrolo-pyrimidine, 3 -methyl adenosine, 5-methylcytidine, C-5 propynyl-cytidine, C-5 propynyl-uridine, 2-aminoadenosine, C5 -bromouridine, C5-fluorouridine, C5 -iodouridine, C5-propynyl-uridine, C5 -propynyl-cytidine, C5 -methylcytidine, 2-aminoadenosine, 7-deazaadenosine, 7-deazaguanosine, 8-oxoadenosine, 8-oxoguanosine, 0(6)-methylguanine, 2-thiocytidine, methylated bases, intercalated bases, and combinations thereof). In some embodiments, a nucleic acid comprises one or more modified sugars (e.g., 2'-fluororibose, ribose, 2'-deoxyribose, arabinose, and hexose) as compared with those in natural nucleic acids. In some embodiments, a nucleic acid has a nucleotide sequence that encodes a functional gene product such as an RNA or protein. In some embodiments, a nucleic acid includes one or more introns. In some embodiments, nucleic acids are prepared by one or more of isolation from a natural source, enzymatic synthesis by polymerization based on a complementary template (in vivo or in vitro), reproduction in a recombinant cell or system, and chemical synthesis. In some embodiments, a nucleic acid is at least 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 1 10, 120, 130, 140, 150, 160, 170, 180, 190, 20, 225, 250, 275, 300, 325, 350, 375, 400, 425, 450, 475, 500, 600, 700, 800, 900, 1000, 1500, 2000, 2500, 3000, 3500, 4000, 4500, 5000 or more residues long. In Page 38 of 16513092399vlAttorney Docket No. 2013763-0005some embodiments, a nucleic acid is partly or wholly single stranded; in some embodiments, a nucleic acid is partly or wholly double stranded. In some embodiments a nucleic acid has a nucleotide sequence comprising at least one element that encodes, or is the complement of a sequence that encodes, a polypeptide. In some embodiments, a nucleic acid has enzymatic activity.
[0118] Operably linked, as used herein, refers to a juxtaposition wherein the components described are in a relationship permitting them to function in their intended manner. A control element "operably linked’ to a functional element is associated in such a way that expression and / or activity of the functional element is achieved under conditions compatible with the control element. In some embodiments, "operably linked” control elements are contiguous (e.g., covalently linked) with the coding elements of interest; in some embodiments, control elements act in trans to or otherwise at a from the functional element of interest.
[0119] Peptide: The term “peptide” as used herein refers to a polypeptide that is typically relatively short, for example having a length of less than about 100 amino acids, less than about 50 amino acids, less than about 40 amino acids less than about 30 amino acids, less than about 25 amino acids, less than about 20 amino acids, less than about 15 amino acids, or less than 10 amino acids.
[0120] Pharmaceutical composition: As used herein, the term “pharmaceutical composition” refers to an active agent, formulated together with one or more pharmaceutically acceptable carriers. In some embodiments, active agent is present in unit dose amount appropriate for administration in a therapeutic regimen that shows a statistically significant probability of achieving a predetermined therapeutic effect when administered to a relevant population (e.g., a cell population).
[0121] Pharmaceutically acceptable carrier: As used herein, the term “pharmaceutically acceptable carrier” means a pharmaceutically-acceptable material, composition or vehicle, such as a liquid or solid filler, diluent, excipient, or solvent encapsulating material, involved in carrying or transporting a subject compound. Each carrier must be “acceptable” in the sense of being compatible with the other ingredients of the formulation and not injurious to the patient. Some examples of materials which can serve as pharmaceutically-acceptable carriers include: sugars, such as lactose, glucose and sucrose; starches, such as com starch and potato starch; cellulose, and its derivatives, such as sodium Page 39 of 16513092399vlAttorney Docket No. 2013763-0005carboxymethyl cellulose, ethyl cellulose and cellulose acetate; powdered tragacanth; malt; gelatin; talc; excipients, such as cocoa butter and suppository waxes; oils, such as peanut oil, cottonseed oil, safflower oil, sesame oil, olive oil, com oil and soybean oil; glycols, such as propylene glycol; polyols, such as glycerin, sorbitol, mannitol and polyethylene glycol; esters, such as ethyl oleate and ethyl laurate; agar; buffering agents, such as magnesium hydroxide and aluminum hydroxide; alginic acid; pyrogen-free water; isotonic saline; Ringer’s solution; ethyl alcohol; pH buffered solutions; polyesters, polycarbonates and / or polyanhydrides; and other non-toxic compatible substances employed in pharmaceutical formulations.
[0122] Physiological conditions: “Physiological conditions” as used herein, has its art-understood meaning referencing conditions under which cells or organisms live and / or reproduce. In some embodiments, the term refers to conditions of the external or internal mileu that may occur in nature for an organism or cell system. In some embodiments, physiological conditions are those conditions present within the body of a human or nonhuman animal, especially those conditions present at and / or within a surgical site.Physiological conditions typically include, e.g., a temperature range of 20 - 40°C, atmospheric pressure of 1, pH of 6-8, glucose concentration of 1-20 mM, oxygen concentration at atmospheric levels, and gravity as it is encountered on earth. In some embodiments, conditions in a laboratory are manipulated and / or maintained at physiologic conditions. In some embodiments, physiological conditions are encountered in an organism.
[0123] Polypeptide: As used herein refers to a polymeric chain of amino acids. In some embodiments, a polypeptide has an amino acid sequence that occurs in nature. In some embodiments, a polypeptide has an amino acid sequence that does not occur in nature. In some embodiments, a polypeptide has an amino acid sequence that is engineered in that it is designed and / or produced through action of the hand of man. In some embodiments, a polypeptide may comprise or consist of natural amino acids, non-natural amino acids, or both. In some embodiments, a polypeptide may comprise or consist of only natural amino acids or only non-natural amino acids. In some embodiments, a polypeptide may comprise D-amino acids, L-amino acids, or both. In some embodiments, a polypeptide may comprise only D-amino acids. In some embodiments, a polypeptide may comprise only L-amino acids. In some embodiments, a polypeptide may include one or more pendant groups or other modifications, e.g., modifying or attached to one or more amino acid side chains, at the polypeptide’s N-terminus, at the polypeptide’s C-terminus, or any combination thereof. In Page 40 of 16513092399vlAttorney Docket No. 2013763-0005some embodiments, such pendant groups or modifications may be selected from the group consisting of acetylation, amidation, lipidation, methylation, pegylation, etc., including combinations thereof. In some embodiments, a polypeptide may be cyclic, and / or may comprise a cyclic portion. In some embodiments, a polypeptide is not cyclic and / or does not comprise any cyclic portion. In some embodiments, a polypeptide is linear. In some embodiments, a polypeptide may be or comprise a stapled polypeptide. In some embodiments, the term “polypeptide” may be appended to a name of a reference polypeptide, activity, or structure; in such instances it is used herein to refer to polypeptides that share the relevant activity or structure and thus can be considered to be members of the same class or family of polypeptides. For each such class, the present specification provides and / or those skilled in the art will be aware of exemplary polypeptides within the class whose amino acid sequences and / or functions are known; in some embodiments, such exemplary polypeptides are reference polypeptides for the polypeptide class or family. In some embodiments, a member of a polypeptide class or family shows significant sequence homology or identity with, shares a common sequence motif (e.g., a characteristic sequence element) with, and / or shares a common activity (in some embodiments at a comparable level or within a designated range) with a reference polypeptide of the class; in some embodiments with all polypeptides within the class). For example, in some embodiments, a member polypeptide shows an overall degree of sequence homology or identity with a reference polypeptide that is at least about 30-40%, and is often greaterthan about 50%, 60%, 70%, 80%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more and / or includes at least one region (e.g., a conserved region that may in some embodiments be or comprise a characteristic sequence element) that shows very high sequence identity, often greaterthan 90% or even 95%, 96%, 97%, 98%, or 99%. Such a conserved region usually encompasses at least 3-4 and often up to 20 or more amino acids; in some embodiments, a conserved region encompasses at least one stretch of at least 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15 or more contiguous amino acids. In some embodiments, a relevant polypeptide may comprise or consist of a fragment of a parent polypeptide. In some embodiments, a useful polypeptide as may comprise or consist of a plurality of fragments, each of which is found in the same parent polypeptide in a different spatial arrangement relative to one another than is found in the polypeptide of interest (e.g., fragments that are directly linked in the parent may be spatially separated in the polypeptide of interest or vice versa, and / or fragments may be present in a different order in Page 41 of 16513092399vlAttorney Docket No. 2013763-0005the polypeptide of interest than in the parent), so that the polypeptide of interest is a derivative of its parent polypeptide.
[0124] Protein : As used herein, the term “protein” refers to a polypeptide (z.e., a string of at least two amino acids linked to one another by peptide bonds). Proteins may include moieties other than amino acids (e.g, may be glycoproteins, proteoglycans, etc.) and / or may be otherwise processed or modified. Those of ordinary skill in the art will appreciate that a “protein” can be a complete polypeptide chain as produced by a cell (with or without a signal sequence), or can be a characteristic portion thereof. Those of ordinary skill will appreciate that a protein can sometimes include more than one polypeptide chain, for example linked by one or more disulfide bonds or associated by other means. Polypeptides may contain L-amino acids, D-amino acids, or both and may contain any of a variety of amino acid modifications or analogs known in the art. Useful modifications include, e.g., terminal acetylation, amidation, methylation, etc. In some embodiments, proteins may comprise natural amino acids, non-natural amino acids, synthetic amino acids, and combinations thereof. The term “peptide” is generally used to refer to a polypeptide having a length of less than about 100 amino acids, less than about 50 amino acids, less than 20 amino acids, or less than 10 amino acids. In some embodiments, proteins are antibodies, antibody fragments, biologically active portions thereof, and / or characteristic portions thereof.
[0125] Protein of Interest: As used herein, the term “protein of interest” refers to a target protein that interacts with or is associated with interacting components. A protein of interest may be included in a fusion protein as described herein or a fusion protein may include a binding agent that specifically binds the protein of interest. Interacting components of a protein of interest are identified and mapped using a fusion protein described herein. In some embodiments, a protein of interest is a protein involved in one or more of transcription, translation, replication (e.g., involved in DNA damage and / or repair), tumor metastasis and protein degradation.
[0126] Proximity, as used herein, “proximity” refers to a distance between a protein of interest and an interacting component (e.g, a transcriptional component) within a cell. In some embodiments, proximity is determined by the length of a linker that connects a proximity labelled enzyme and a protein of interest or a binding agent that targets a protein of interest. In some embodiments, such a distance may be within a range of about 0.5 nm-100Page 42 of 16513092399vlAttorney Docket No. 2013763-0005nm e.g., with a range of about lOnm-lOOnm, about 5-50 nm, between about 10-30 nm, or between about 15-20 nm).
[0127] Proximity Labeling. As used herein, “proximity labeling” refers to a method of using specific enzymes to label an interacting component (e.g., biomolecules, such as proteins or RNA) that are interacting with or are within a certain spatial proximity with a protein of interest in a cell. Proximity labeling can be performed using a fusion protein comprising a protein of interest or a binding agent that targets a protein of interest and a proximity labeling enzyme. A particular spatial proximity can be controlled by, among other things, a linker connecting the protein of interest or binding agent targeting the protein of interest and the proximity labeling enzyme. When applied to a cell or system, a proximity labeling enzyme is capable of labeling biomolecules spatially proximal to the protein of interest which can then be selectively marked with an agent such as biotin for pulldown and analysis.
[0128] Recombinant. As used herein, “recombinant” is intended to refer to polypeptides that are designed, engineered, prepared, expressed, created, manufactured, and / or or isolated by recombinant means, such as polypeptides expressed using a recombinant expression vector transfected into a host cell, polypeptides isolated from a recombinant, combinatorial human polypeptide library (e.g., Hoogenboom, TIB Tech 15:62, 1997; Azzazy Clin. Biochem. 35:425, 2002; Gavilondo BioTechniques 29:128, 2002;Hoogenboom Immunology Today 21:311, 2000), antibodies isolated from an animal (e.g., a mouse) that is transgenic for human immunoglobulin genes (see e.g., Taylor Nuc. Acids Res.20:6287, 1992; Little Immunology Today 12:364, 2000; Kellermann Curr. Opin. Biotechnol 13:593, 2002; Murphy Proc. NatlAcadSci USA 111:5153, 2104) or polypeptides prepared, expressed, created or isolated by any other means that involves splicing selected sequence elements to one another. In some embodiments, one or more of such selected sequence elements is found in nature. In some embodiments, one or more of such selected sequence elements is designed in silico. In some embodiments, one or more such selected sequence elements results from mutagenesis (e.g., in vivo or in vitro) of a known sequence element, e.g., from a natural or synthetic source. For example, in some embodiments, a recombinant antibody polypeptide is comprised of sequences found in the germline of a source organism of interest (e.g., human, mouse, etc.). In some embodiments, a recombinant antibody has an amino acid sequence that resulted from mutagenesis (e.g., in vitro or in vivo, for example in a Page 43 of 16513092399vlAttorney Docket No. 2013763-0005transgenic animal), so that the amino acid sequences of the VH and VL regions of the recombinant antibodies are sequences that, while originating from and related to germline VH and VL sequences, may not naturally exist within the germline antibody repertoire in vivo.
[0129] Recovering: as used herein, refers to the process of rendering an agent or entity substantially free of other previously-associated components, for example by isolation, e.g., using purification techniques known in the art. In some embodiments, an agent or entity is recovered from a natural source and / or a source comprising cells.
[0130] Reference: As used herein, “reference” describes a standard or control relative to which a comparison is performed. For example, in some embodiments, an agent, animal, individual, population, sample, cell, sequence or value of interest is compared with a reference or control agent, animal, individual, population, sample, cell, sequence or value. In some embodiments, a reference or control is tested and / or determined substantially simultaneously with the testing or determination of interest. In some embodiments, a reference or control is a historical reference or control, optionally embodied in a tangible medium. Typically, as would be understood by those skilled in the art, a reference or control is determined or characterized under comparable conditions or circumstances to those under assessment. Those skilled in the art will appreciate when sufficient similarities are present to justify reliance on and / or comparison to a particular possible reference or control.
[0131] Sample: As used herein, the term “sample” typically refers to an aliquot of material obtained or derived from a source of interest, as described herein. In some embodiments, a source of interest is a biological or environmental source. In some embodiments, a source of interest may be or comprise a cell or an organism, such as a microbe, a plant, or an animal (e.g., a human). In some embodiments, a source of interest is or comprises biological tissue or fluid. In some embodiments, a biological tissue or fluid may be or comprise amniotic fluid, aqueous humor, ascites, bile, bone marrow, blood, breast milk, cerebrospinal fluid, cerumen, chyle, chime, ejaculate, endolymph, exudate, feces, gastric acid, gastric juice, lymph, mucus, pericardial fluid, perilymph, peritoneal fluid, pleural fluid, pus, rheum, saliva, sebum, semen, serum, smegma, sputum, synovial fluid, sweat, tears, urine, vaginal secreations, vitreous humour, vomit, and / or combinations or component(s) thereof. In some embodiments, a biological fluid may be or comprise an intracellular fluid, an extracellular fluid, an intravascular fluid (blood plasma), an interstitial fluid, a lymphatic fluid, and / or a transcellular fluid. In some embodiments, a biological fluid may be or Page 44 of 16513092399vlAttorney Docket No. 2013763-0005comprise a plant exudate. In some embodiments, a biological tissue or sample may be obtained, for example, by aspirate, biopsy (e.g., fine needle or tissue biopsy), swab (e.g., oral, nasal, skin, or vaginal swab), scraping, surgery, washing or lavage (e.g., brocheoalvealar, ductal, nasal, ocular, oral, uterine, vaginal, or other washing or lavage). In some embodiments, a biological sample is or comprises cells obtained from an individual. In some embodiments, a sample is a “primary sample” obtained directly from a source of interest by any appropriate means. In some embodiments, as will be clear from context, the term “sample” refers to a preparation that is obtained by processing (e.g., by removing one or more components of and / or by adding one or more agents to) a primary sample. For example, filtering using a semi-permeable membrane. Such a “processed sample” may comprise, for example nucleic acids or proteins extracted from a sample or obtained by subjecting a primary sample to one or more techniques such as amplification or reverse transcription of nucleic acid, isolation and / or purification of certain components, etc. In some embodiments, a sample may be a “crude” sample in that it has been subjected to relatively little processing and / or is complex in that it includes components of relatively varied chemical classes.
[0132] Solid Tumor. As used herein, the term “solid tumor” refers to an abnormal mass of tissue that usually does not contain cysts or liquid areas. In some embodiments, a solid tumor may be benign; in some embodiments, a solid tumor may be malignant. Those skilled in the art will appreciate that different types of solid tumors are typically named for the type of cells that form them. Examples of solid tumors are carcinomas, lymphomas, and sarcomas. In some embodiments, solid tumors may be or comprise adrenal, bile duct, bladder, bone, brain, breast, cervix, colon, endometrium, esophagum, eye, gall bladder, gastrointestinal tract, kidney, larynx, liver, lung, nasal cavity, nasopharynx, oral cavity, ovary, penis, pituitary, prostate, retina, salivary gland, skin, small intestine, stomach, testis, thymus, thyroid, uterine, vaginal, and / or vulval tumors.
[0133] Specific binding: As used herein, the term “specific binding” refers to an ability to discriminate between possible binding partners in the environment in which binding is to occur. A binding agent that interacts with one particular target when other potential targets are present is said to "bind specifically” to the target with which it interacts. In some embodiments, specific binding is assessed by detecting or determining degree and / or rate of association between the binding agent and its partner (e.g., a protein of interest); in some embodiments, specific binding is assessed by detecting or determining degree and / or rate of dissociation of a Page 45 of 16513092399vlAttorney Docket No. 2013763-0005binding agent-partner complex; in some embodiments, specific binding is assessed by detecting or determining ability of the binding agent to compete an alternative interaction between its partner and another entity. In some embodiments, specific binding is assessed by performing such detections or determinations across a range of concentrations.
[0134] Specificity. As is known in the art, “specificity” is a measure of the ability of a particular ligand to distinguish its binding partner from other potential binding partners.
[0135] Subject: As used herein, the term “subject” or “test subject” refers to any organism from which cell populations are obtained and subjected to proximity labeling methods as described herein. Typical subjects include animals (e.g., mammals such as mice, rats, rabbits, non-human primates, and humans; etc.) and plants. In some embodiments, a subject may be suffering from, and / or susceptible to a disease, disorder, and / or condition (e.g., cancer).
[0136] Substantially: As used herein, the term “substantially” refers to the qualitative condition of exhibiting total or near-total extent or degree of a characteristic or property of interest. One of ordinary skill in the biological arts will understand that biological and chemical phenomena rarely, if ever, go to completion and / or proceed to completeness or achieve or avoid an absolute result. The term “substantially” is therefore used herein to capture the potential lack of completeness inherent in many biological and chemical phenomena.
[0137] Substantial identity: as used herein refers to a comparison between amino acid or nucleic acid sequences. As will be appreciated by those of ordinary skill in the art, two sequences are generally considered to be "substantially identical" if they contain identical residues in corresponding positions. As is well known in this art, amino acid or nucleic acid sequences may be compared using any of a variety of algorithms, including those available in commercial computer programs such as BLASTN for nucleotide sequences and BLASTP, gapped BLAST, and PSI-BLAST for amino acid sequences. Exemplary such programs are described in Altschul et al., Basic local alignment search tool, J. Mol. Biol., 215(3): 403-410, 1990; Altschul et al., Methods in Enzymology; Altschul et al., Nucleic Acids Res. 25:3389-3402, 1997; Baxevanis et al., Bioinformatics: A Practical Guide to the Analysis of Genes and Proteins, Wiley, 1998; and Misener, et al, (eds.), Bioinformatics Methods and Protocols (Methods in Molecular Biology, Vol. 132), Humana Press, 1999. In addition to identifying identical sequences, the programs mentioned above typically provide an indication of the Page 46 of 16513092399vlAttorney Docket No. 2013763-0005degree of identity. In some embodiments, two sequences are considered to be substantially identical if at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more of their corresponding residues are identical over a relevant stretch of residues. In some embodiments, the relevant stretch is a complete sequence. In some embodiments, the relevant stretch is at least 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 125, 150, 175, 200, 225, 250, 275, 300, 325, 350, 375, 400, 425, 450, 475, 500 or more residues.
[0138] In the context of a CDR, reference to "substantial identity" typically refers to a CDR having an amino acid sequence at least 80%, preferably at least 85%, at least 90%, at least 95%, at least 98% or at least 99% identical to that of a reference CDR.
[0139] Susceptible to: An individual who is “susceptible to” a disease, disorder, and / or condition is one who has a higher risk of developing the disease, disorder, and / or condition than does a member of the general public. In some embodiments, an individual who is susceptible to a disease, disorder and / or condition may not have been diagnosed with the disease, disorder, and / or condition. In some embodiments, an individual who is susceptible to a disease, disorder, and / or condition may exhibit symptoms of the disease, disorder, and / or condition. In some embodiments, an individual who is susceptible to a disease, disorder, and / or condition may not exhibit symptoms of the disease, disorder, and / or condition. In some embodiments, an individual who is susceptible to a disease, disorder, and / or condition will develop the disease, disorder, and / or condition. In some embodiments, an individual who is susceptible to a disease, disorder, and / or condition will not develop the disease, disorder, and / or condition.
[0140] Three dimensional representation . As used herein, the term “three dimensional representation” refers to converting the lists of structure coordinates into structural models or graphical representation in three-dimensional space. In some embodiments, the three dimensional structure may be displayed or used to performing computer modeling or fitting operations. In some embodiments, the structure coordinates themselves, without the displayed model, may be used to perform computer-based modeling and fitting operations. In some embodiments, a three dimensional representation may represent a cell (e.g., and location of transcriptional machinery components within a cell).Page 47 of 16513092399vlAttorney Docket No. 2013763-0005
[0141] Transcriptional Machinery Component: As used herein, “transcriptional machinery” or “transcriptional machinery component” refers to any component involved in the process of transcription. General examples of transcriptional machinery components include, e.g., promoters and general transcription factors. Specific examples include, e.g., RNA polymerase (e.g., RNApol II) which comprises three main components that function in the transcription: (1) The RNA-Pol II enzyme itself, (2) TBP (TATA-binding protein) / TAF (TBP-associated factor) complex SL1 (selectivity factor 1) / TIF-IB (transcription initiation factor-IB) and (3) transactivator protein UBF (upstream binding factor) (see Russell et al., 2006, Biochem Soc Symp. (73): 203-216).
[0142] Tumor. As used herein, the term “tumor” refers to an abnormal growth of cells or tissue. In some embodiments, a tumor may comprise cells that are precancerous (e.g., benign), malignant, pre-metastatic, metastatic, and / or non-metastatic. In some embodiments, a tumor is associated with, or is a manifestation of, a cancer. In some embodiments, a tumor may be a disperse tumor or a liquid tumor. In some embodiments, a tumor may be a solid tumor.
[0143] Variant: As used herein in the context of molecules, e.g., nucleic acids, proteins, or small molecules, the term “variant” refers to a molecule that shows significant structural identity with a reference molecule but differs structurally from the reference molecule, e.g., in the presence or absence or in the level of one or more chemical moieties as compared to the reference entity. In some embodiments, a variant also differs functionally from its reference molecule. In general, whether a particular molecule is properly considered to be a “variant” of a reference molecule is based on its degree of structural identity with the reference molecule. As will be appreciated by those skilled in the art, any biological or chemical reference molecule has certain characteristic structural elements. A variant, by definition, is a distinct molecule that shares one or more such characteristic structural elements but differs in at least one aspect from the reference molecule. To give but a few examples, a polypeptide may have a characteristic sequence element comprised of a plurality of amino acids having designated positions relative to one another in linear or three-dimensional space and / or contributing to a particular structural motif and / or biological function; a nucleic acid may have a characteristic sequence element comprised of a plurality of nucleotide residues having designated positions relative to another in linear or three-dimensional space. In some embodiments, a variant polypeptide or nucleic acid may differ Page 48 of 16513092399vlAttorney Docket No. 2013763-0005from a reference polypeptide or nucleic acid as a result of one or more differences in amino acid or nucleotide sequence and / or one or more differences in chemical moieties (e.g., carbohydrates, lipids, phosphate groups) that are covalently components of the polypeptide or nucleic acid (e.g., that are attached to the polypeptide or nucleic acid backbone). In some embodiments, a variant polypeptide or nucleic acid shows an overall sequence identity with a reference polypeptide or nucleic acid that is at least 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, or 99%. In some embodiments, a variant polypeptide or nucleic acid does not share at least one characteristic sequence element with a reference polypeptide or nucleic acid. In some embodiments, a reference polypeptide or nucleic acid has one or more biological activities. In some embodiments, a variant polypeptide or nucleic acid shares one or more of the biological activities of the reference polypeptide or nucleic acid. In some embodiments, a variant polypeptide or nucleic acid lacks one or more of the biological activities of the reference polypeptide or nucleic acid. In some embodiments, a variant polypeptide or nucleic acid shows a reduced level of one or more biological activities as compared to the reference polypeptide or nucleic acid. In some embodiments, a polypeptide or nucleic acid of interest is considered to be a “variant” of a reference polypeptide or nucleic acid if it has an amino acid or nucleotide sequence that is identical to that of the reference but for a small number of sequence alterations at particular positions. Typically, fewer than about 20%, about 15%, about 10%, about 9%, about 8%, about 7%, about 6%, about 5%, about 4%, about 3%, or about 2% of the residues in a variant are substituted, inserted, or deleted, as compared to the reference. In some embodiments, a variant polypeptide or nucleic acid comprises about 10, about 9, about 8, about 7, about 6, about 5, about 4, about 3, about 2, or about 1 substituted residues as compared to a reference. Often, a variant polypeptide or nucleic acid comprises a very small number (e.g., fewer than about 5, about 4, about 3, about 2, or about 1) number of substituted, inserted, or deleted, functional residues (i.e., residues that participate in a particular biological activity) relative to the reference. In some embodiments, a variant polypeptide or nucleic acid comprises not more than about 5, about 4, about 3, about 2, or about 1 addition or deletion, and, in some embodiments, comprises no additions or deletions, as compared to the reference. In some embodiments, a variant polypeptide or nucleic acid comprises fewer than about 25, about 20, about 19, about 18, about 17, about 16, about 15, about 14, about 13, about 10, about 9, about 8, about 7, about 6, and commonly fewer than about 5, about 4, about 3, or about 2 additions Page 49 of 16513092399vlAttorney Docket No. 2013763-0005or deletions as compared to the reference. In some embodiments, a reference polypeptide or nucleic acid is one found in nature. In some embodiments, a reference polypeptide or nucleic acid is a human polypeptide or nucleic acid.
[0144] Vector, as used herein, refers to a nucleic acid molecule capable of transporting another nucleic acid to which it has been linked. One type of vector is a “plasmid”, which refers to a circular double stranded DNA loop into which additional DNA segments may be ligated. Another type of vector is a viral vector, wherein additional DNA segments may be ligated into the viral genome. Certain vectors are capable of autonomous replication in a host cell into which they are introduced (e.g., bacterial vectors having a bacterial origin of replication and episomal mammalian vectors). Other vectors (e.g., non-episomal mammalian vectors) can be integrated into the genome of a host cell upon introduction into the host cell, and thereby are replicated along with the host genome.Moreover, certain vectors are capable of directing the expression of genes to which they are operatively linked. Such vectors are referred to herein as "expression vectors "
[0145] Standard techniques may be used for recombinant DNA, oligonucleotide synthesis, and tissue culture and transformation (e.g., electroporation, lipofection).Enzymatic reactions and purification techniques may be performed according to manufacturer's specifications or as commonly accomplished in the art or as described herein. The foregoing techniques and procedures may be generally performed according to conventional methods well known in the art and as described in various general and more specific references that are cited and discussed throughout the present specification. See e.g., Sambrook et al., Molecular Cloning: A Laboratory Manual (2d ed., Cold Spring Harbor Laboratory Press, Cold Spring Harbor, N.Y. (1989)), which is incorporated herein by reference for any purpose.DETAILED DESCRIPTION
[0146] In some embodiments, the present disclosure provides technologies that relate to fusion proteins that are capable of proximity labeling interacting components of a protein of interest (POI) - including, for example, providing such fusion proteins themselves, and / or providing methods and / or reagents for identifying, characterizing, manufacturing and / or using them and / or compositions that comprise and / or deliver them. The present disclosure also provides systems comprising one or more fusion proteins described herein capable of Page 50 of 16513092399vlAttorney Docket No. 2013763-0005labelling and / or isolating POI or gene of interest (GOI), or a target genomic locus. Systems described herein are useful in isolating (i.e., “pulling down”) a genomic locus of interest, including interacting proteins, e.g., transcriptional components. In some embodiments, a transcriptional machinery component comprises a repressor and / or activator of transcription.A. Fusion Proteins
[0147] Among other things, the present disclosure provides fusions proteins and compositions comprising the same to identify protein complexes interacting with a protein of interest (e.g., a transcriptional machinery component) in various cell types (e.g., cancer cells). A fusion protein, as described herein, may comprise (i) a protein of interest (e.g., a transcriptional machinery component), (ii) a linker, and (iii) a proximity labeling enzyme, where the linker connects the protein of interest to the proximity labeling enzyme. A fusion protein, as described herein, may use a proximity labeling enzyme to label proteins that interact with the protein of interest when they are in proximity therewith.
[0148] In some embodiments, a fusion protein, as described herein, may comprise (i) a binding agent (e.g., a nanobody) that specifically binds to a protein of interest, (ii) a linker, and (iii) a proximity labeling enzyme, where the linker connects the binding agent to the proximity labeling enzyme.
[0149] In some embodiments, a fusion protein may be part of a proximity labelling system (e.g., a split proximity labeling system), that contains one or more fusion proteins. In some embodiments, a fusion protein comprises a protein of interest (POI) and a protein fragment, wherein the protein fragment comprises a fragment of a functional indicator protein (e.g., a green fluorescent protein (GFP) fragment, e.g., GFP10 or GFP11). In some embodiments, a fusion protein comprises a proximity labelling enzyme and a protein fragment, wherein the protein fragment comprises another fragment of a functional indicator protein. In some embodiments, a system comprises two or more fusion proteins, where each fusion protein contains a fragment of an indicator protein, where the fragments associate when in proximity to one another and form a functional indicator protein (e.g., GFP). In some embodiments, a functional indicator protein is GFP that fluoresces when its fragments are associated and it is fully assembled. Such fusion proteins allow for conditionalPage 51 of 16513092399vlAttorney Docket No. 2013763-0005visualization of complexes that have been associated and where proximity labeling has occurred.
[0150] In some embodiments, a fusion protein may be part of a proximity labelling system (e.g., a split proximity labeling system), that contains one or more fusion proteins and targets a particular target genomic locus. In some embodiments, a fusion protein comprises a POI and a protein fragment, wherein the first protein fragment comprises a fragment of a functional indicator protein (e.g., a green fluorescent protein (GFP) fragment, e.g., GFP10 or GFP11). In some embodiments, a fusion protein comprises a DNA-targeting moiety (e.g., dCas9 or LNA) that binds to a target genomic locus and a protein fragment, wherein the protein fragment comprises fragment of a functional indicator protein (e.g., a green fluorescent protein (GFP) fragment, e.g., GFP 10 or GFP 11). In some embodiments, a fusion protein comprises a proximity labeling enzyme (e.g., APEX2) and a protein fragment, wherein the protein fragment comprises another fragment of a functional indicator protein (e.g., a GFP fragment, e.g., GFPl-9r).
[0151] Choosing a particular protein of interest that is part of transcriptional machinery to include in, or be targeted by a binding agent of, allows for identification, through proximity labelling, of proteins that interact with transcriptional machinery.Comparing proteins labelled from, for example, cancer vs. non-cancer cell populations provides insight into transcriptional patterns and composition of transcriptional apparatus specific to cancer cells.1. Proximity Labeling Enzymes
[0152] Proximity labeling (PL) methods, as described herein, are used to introduce a covalent biotin tag to proteins interacting with a protein of interest in a cell. In a typical proximity labelling reaction, proximity enzymes convert a supplemented substrate (e.g., biotin-phenol supplemented with hydrogen peroxide into a reactive biotinylated intermediate (e.g., biotin-phenol free radical) that then transfers biotin to amino acid side chains of interacting proteins in proximity to a protein of interest.
[0153] Certain PL enzymes can be classified into groups: peroxidases, biotin ligases, and other enzymes, which differ in substrates, kinetics, labeling radius, labeling conditions,Page 52 of 16513092399vlAttorney Docket No. 2013763-0005and ability to function in different systems and organisms (see Mair, Andrea, and Dominique C. Bergmann. Plant physiology 188.2 (2022): 756-768, which is herein incorporated by reference in its entirety).
[0154] Examples of biotin ligases include BioID, BioID2, TurboID, miniTurbo; and engineered ascorbate peroxidase, which includes APEX and APEX2. Other proximity labeling proteins include horse radish peroxidase (HRP).
[0155] In some embodiments, a PL enzyme is a split enzyme (e.g., a split Turbo ID or split APEX2 (see e.g., Han, Yisu, et al., “Directed evolution of split APEX2 peroxidase.’1ACS' chemical biology 14.4 (2019): 619-635).
[0156] In proximity labelling, after the proteins are labelled by the labelling enzyme, cells are lysed and the biotinylated proteins are extracted, and subjected to mass spectrometry (Ummethum, Henning, and Stephan Hamperl. Frontiers in Genetics 11 (2020): 450, which is herein incorporated by reference in its entirety).
[0157] In a proximity labeling system, ascorbate peroxidase (e.g., APEX or APEX2) catalyzes one-electron oxidation of biotin-phenol into a highly reactive and short-lived biotin-phenoxyl radical (in the presence of H2O2), which biotinylates tyrosine predominantly in nearby interacting proteins. Alternatively, a biotin ligase (e.g., BioID or TurboID / miniTurbo) catalyzes the synthesis of biotinoyl-5’-AMP intermediate from biotin and ATP, which intermediate then tags lysine in nearby interacting proteins.
[0158] The biotinylation process is also dependent on the protein of interest and occurs in a proximate dependent manner. Biotinylated proteins are then enriched and purified by streptavidin pulldown assay, further digested into peptides and identified by quantitative LC-MS / MS. (see Xu, Yangfan, Xianqun Fan, and Yang Hu. Cell & Bioscience 11.1 (2021): 1-9, which is herein incorporated by reference in its entirety).
[0159] Fusion proteins, as described herein, may utilize various proximity labeling enzymes including but not limited to biotin ligase, which includes BioID, BioID2, TurboID, miniTurbo; horse radish peroxidase (HRP); and engineered ascorbate peroxidase, which includes APEX and APEX2. In some embodiments, proximity labelling enzymes include a Page 53 of 16513092399vlAttorney Docket No. 2013763-0005variant of BioID, BioID2, TurboID, miniTurbo; horse radish peroxidase (HRP); APEX or APEX2.
[0160] In some embodiments, Turbo-ID proximity labeling enzyme is selected for a fusion protein as described herein in order to achieve long-term interaction of the protein of interest, which can range from hours to days, and because it uses endogenous or natural biotin molecules for the labelling. In some embodiments, an APEX2 proximity labeling enzyme is selected for a fusion protein as described herein to label interacting proteins within a short timeframe such as a minute.(i) Peroxidases
[0161] One example of a peroxidase enzyme used in PL is horseradish peroxidase (HRP). HRP can be fused to a protein of interest (POI). HRP, however, is limited to cell surfaces and compartments with oxidizing conditions (e.g., secretory pathway or the ER, since the enzyme requires disulfide bonds and calcium (Ca2+)-binding to maintain its structural integrity) (see Mair, Andrea, and Dominique C. Bergmann. Plant physiology 188.2 (2022): 756-768, which is herein incorporated by reference in its entirety).
[0162] In some embodiments, a peroxidase used for PL comprises an engineered ascorbate peroxidase such as APEX or APEX2. APEX and its derivatives function in many cellular compartments, unlike HRP. APEX2, which was derived from (soybean) APEX by yeast display -based evolution, has improved activity and sensitivity, allowing use with low expressing bait proteins (Lam et al., 2015). Major advantages of peroxidase-based labeling are an extremely high labeling speed (1 min with APEX / APEX2) and the ability to timely control the labeling process through H2O2 availability. Cellular toxicity of H2O2 and poor membrane permeability of biotin-phenol, however, have limited the use of APEX methods primarily to cultured human cells. APEX enzyme is a 28 kDa monomeric ascorbate peroxidase. In proximity labeling in living cells using APEX, biotin-phenol (BP) is added and catalyzed by APEX in the presence of hydrogen peroxide which produces a biotin-phenoxyl intermediate. Subsequently, the intermediate can covalently react with electronrich amino acids (e.g., tyrosine) in interacting proteins.Page 54 of 16513092399vlAttorney Docket No. 2013763-0005
[0163] One of drawbacks of using APEX in PL is its relatively low cellular activity and sensitivity. Thus, higher amount of total protein extracts are typically required in order to provide sufficient biotinylated proteins for subsequent MS identification. Alternatively, APEX2 has shown higher catalytic activity and sensitivity and has been used to successfully identify interacting proteins in living cells. Additionally, APEX2 is able to label interacting proteins on a one-minute time scale. In some embodiments, instead of a traditional APEX2 substrate biotin-phenol, a alkyne-phenol (Alk-Ph) substrate, may be used. Such a substrate used with APEX2 has been shown to enhance membrane permeability and labeling efficiency (see Xu, Yangfan, Xianqun Fan, and Yang Hu. Cell & Bioscience 11.1 (2021): 1-9, which is herein incorporated by reference in its entirety).
[0164] APEX2C32S is another APEX that has been further engineered from APEX2 which lacks a conserved Cys residue, promises improved stability of APEX2 fusion proteins (Huang et al., 2019).
[0165] In some embodiments, a PL enzyme used in a fusion protein described herein is a peroxidase (e.g., APEX2) that comprises an amino acid sequence as shown in SEQ ID NO: 10 or a fragment or variant thereof, or encoded by a nucleic acid sequence SEQ ID NO: 3 or 9, or a fragment or variant thereof.(ii) Biotin Ligase
[0166] In some embodiments, a fusion protein comprises a proximity labelling enzyme using wild type or mutant biotin ligase.
[0167] Biotin ligase (BirA) or BioID is a 33.5 kDa enzyme that is 321 amino acids in length and was derived from E. coli. BirA catalyzes context-specific conjugation of biotin to a lysine 8-amine in biotin retention and biosynthesis pathways. This reaction is ATP-dependent. As used herein, wild type biotin ligase refers to a naturally occurring bacterial biotin ligase having wild type biotinylation activity. Wild type biotin ligase is represented in the amino acid sequence GenBank Accession No. M10123 and the nucleotide sequence of represented in GenBank Accession No. M10123. Biotin ligase is also known as biotin protein ligase, biotin operon repressor protein, BirA, biotin holoenzyme synthetase and biotin-[acetyl-CoA carboxylase] synthetase.Page 55 of 16513092399vlAttorney Docket No. 2013763-0005
[0168] In some embodiments, a proximity labeling enzyme may comprise a biotin ligase mutant that recognize biotin analogs. Biotin ligase mutants can be generated in any number of ways, including phage display technology. A biotin ligase mutant may have any various mutations, including addition, deletion or substitution of one or more amino acids relative to the wild type sequence. In some embodiments, a mutation will be present in the biotin interaction and activation region, spanning amino acids 83-235 of the wildtype sequence. In some embodiments, a biotin ligase mutant comprises any of the mutants described in US Publication No. 20050233389, which is herein incorporated by reference in its entirety.
[0169] As BirA (i.e., Bio ID) is relatively slow (labeling times of about 24 hours), certain mutants of wild type including BioID2, TurboID, and miniTurbo have been developed and may be utilized in fusion proteins as described herein. BioID2 is an exemplary mutant generated from Aquifex aeolicus BirA. It is smaller and requires less biotin, but has an increased temperature optimum (50°C). Another mutant biotin ligase is Modified Bacillus subtilis BirA (BASU), which has been shown to have improved labeling speed.
[0170] TurboID and miniTurbo are engineered biotin ligases generated through yeast surface display directed evolution. TurboID enables labeling times of approximately 10 minutes in contrast to the approximately 18 hour labeling times required by BioID. In addition these mutants allow for broadening the optimal temperature range and are effective in different model organisms. TurboID, however, has a slightly larger labeling radius of at least >35 nm that increases with time, has low-level activity with endogenous biotin, and can lead to cellular toxicity. miniTurbo is half as active as TurboID, but has lower background when no exogenous biotin is added.
[0171] AirlD is a synthetic BirA (another mutant) that shows greatly enhanced activity (1-6 h labeling time) and a broad temperature range.
[0172] Most recently microID and ultralD (Zhao et al., 2021), which are the smallest mutants, were engineered from BioID2. Both have similar kinetics and activity to TurboID in human cells and have less background labeling from endogenous biotin.Page 56 of 16513092399vlAttorney Docket No. 2013763-0005
[0173] One advantage of using biotin ligase in PL is that labeling only requires the addition of biotin (which is membrane -permeable and nontoxic). Biotin ligases have been demonstrated for us in PL in systems that include live animal and plant cells. However, the labeling times are longer, which can make biotin ligases less suitable PL that involves proteins of interest that must be labeled in a short window of time (see Mair, Andrea, and Dominique C. Bergmann. Plant physiology 188.2 (2022): 756-768, which is herein incorporated by reference in its entirety).
[0174] In some embodiments, a PL enzyme used in a fusion protein described herein is a biotin ligase that comprises an amino acid sequence as shown in SEQ ID NO: 32, or a fragment or variant thereof, or the nucleic acid sequence as shown in SEQ ID NO: 32.
[0175] Additional PL enzymes may include any that are described in Mair et al., 2022, which is herein incorporated by reference, including enzymes such as NEDDylation, a fusion of the human NEDD8 (developmentally down-regulated protein 8) E2 ligase Ubiquitin (Ubq) -conjugating 12 (Ubcl2); enzyme-mediated proximity cell labeling (EXCELL), mini Singlet Oxygen Generator (miniSOG), and split enzyme versions for various PL enzymes such as HRP, APEX2, BioID, TurboID, and miniSOG.
[0176] BirA biotin ligase sequences from a number of bacterial species are listed in the National Center for Biotechnology Information (NCBI) database. See, for example, NCBI entries: Accession Nos. YP_002410237, NP_312927, WP_063115295, WP_063082625, WP_060615925, NP_844010, NP_844010, NP_390125, WP_044306464, WP_011109968, NP_390125, WP_060398894, WP_041117801, WP_041109603,YP 499991, WP_042909036, WP_031903905, NP_359307, WP_061816626, WP_061767634, NP_252970, NP_790457, YP_237632, WP_057960767, WP_057400631, WP_061193045, YP_237632, WP_058975108, WP_052967038, WP_054095365, WP 003292971, WP_046622626, WP_025240331, NP_764699, NP_715854, YP_352592, NP_952984, YP_205808, NP_639277, YP_001034965, YP_003029217, NP_771543, NP_301572, YP_006969295, NP_213397, NP_225061, NP_220244, and YP_001004658. Any of these sequences, or a biologically active fragment thereof, or a variant thereof comprising a sequence having at least about 80-100% sequence identity thereto, including any percent identity within this range, such as 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92,Page 57 of 16513092399vlAttorney Docket No. 2013763-000593, 94, 95, 96, 97, 98, or 99% sequence identity thereto, can be used in fusion proteins as described herein.2. Proteins of Interest
[0177] Fusion proteins, as described herein, comprise a protein of interest (POI) or binding agent(s) that targets a POI. Fusion proteins described herein and used in proximity labelling (PL) can include a wide range of proteins, whose interacting components are to be identified in a cell.
[0178] The present disclosure provides and exemplifies particular embodiments that apply provide technologies to transcription complexes. Those skilled in the art, reading the present disclosure, will now appreciate that technologies described herein may also be applicable to other cellular protein complexes, including in many embodiments to essential cellular protein complexes. In some embodiments, provided technologies are applied to cellular protein complexes involved in pathways of transcription, translation, replication (e.g., involved in DNA damage and / or repair), tumor metastasis and / or protein degradation.
[0179] In some embodiments, a POI includes replication protein (e.g., a protein that participates in a complex involved in, and typically required for, productive nucleic acid replication, such as DNA replication). In some embodiments, a POI is or comprises Claspin (CLSPN), which is an interface between helicase and polymerase.
[0180] In some embodiments, a POI includes a protein involved in DNA damage (e.g., a protein that participates in a complex involved in, and typically required in DNA damage processes). In some embodiments, a POI is or comprises ATM serine / threonine kinase or ataxia-telangiectasia mutated (ATM). In some embodiments, a POI is or comprises serine / threonine -protein kinase (ATR), i.e., ataxia telangiectasia and Rad3-related protein ATR or FRAP-related protein 1 (FRP1) ATR.
[0181] In some embodiments, a POI includes ubiquitination protein (e.g., a protein that participates in a complex involved in, and typically required for, productive ubiquitination, such as protein ubiquitination for ATP-dependent degradation). In some embodiments, a POI is or comprises proteasome 20S subunit beta 5 (PSB5) or proteasomePage 58 of 16513092399vlAttorney Docket No. 2013763-000520S subunit beta 2 (PSB2) or proteasome 20S subunit beta 1 (PSB1), and other 20S and 26S proteasome catalytic subunits.
[0182] In some embodiments, a POI includes translation protein (e.g., protein that participates in a complex involved in, and typically required for, productive translation of RNA into protein). In some embodiments, a POI is or comprises ribosomal protein lateral stalk subunit (P0RPLP0). In some embodiments, a POI is or comprises ribosomal protein lateral stalk subunit Pl (RPLP1). In some embodiments, a POI is or comprises ribosomal protein lateral stalk subunit P2 (RPLP2). In some embodiments, a POI is or comprises eukaryotic translation elongation factor 1 alpha 1 (EFl A).
[0183] In some embodiments, a POI is a transcriptional core component. In some embodiments, a POI is a translational core component. In some embodiments, a POI is a replication core component (e.g., involved in DNA damage and / or repair). In some embodiments, a POI is a component involved in tumorigenesis or metastasis. In some embodiments, a POI is a component involved in a protein degradation pathway.
[0184] Examples of POIs include generally, but are not limited to signal transduction proteins (e.g., cell surface receptors, kinases, adapter proteins, etc.), nuclear proteins (e.g., transcription factors, histones, etc.), mitochondrial proteins (e.g., cytochromes, transcription factors, etc.) and hormone receptors. Other POIs may include a cytosolic protein, a membrane protein, a P-body protein, or a secretory pathway protein.
[0185] In some embodiments, a POI is a protein associated with a specific subcellular compartment or region (e.g., the nucleus, endoplasmic reticulum, golgi, mitochondria, mitochondria outer membrane, mitochondria inner membrane, mitochondria matrix space, chloroplasts, synaptic cleft, presynaptic membrane, postsynaptic membrane, dendritic spines, transport vesicles, regions of contact between mitochondria and endoplasmic reticulum, nuclear membrane, etc.). In some embodiments, a POI is present in a particular cell type (e.g., astrocytes, dendrocytes, myoblasts, cancer cells, stem cells, etc.) and is, for example, present within a specific cell type within a complex tissue, animal, or cell population. In some embodiments, a POI is associated with a particular macromolecular complex (e.g., protein complexes such as ribosomes, replisome, transcription complex, spliceosome, DNAPage 59 of 16513092399vlAtorney Docket No. 2013763-0005repair complex, faty acid synthase, polyketide synthase, non-ribosomal peptide synthase, glutamate receptor signaling complex, neurexin-neuroligin signaling complex, etc.)(i) Transcriptional Machinery Components
[0186] In some embodiments, the present disclosure provides fusion proteins that include, for example, transcriptional machinery components, e.g., components that activate or repress transcription.
[0187] Transcription is the process of copying a DNA sequence into RNA.Transcription is carried out by an enzyme called RNA polymerase and a number of accessory proteins called transcription factors. Transcription factors interact and bind specific DNA sequences (e.g., enhancer and promoter sequences) and recruit RNA polymerase to a transcription start site and form the transcription initiation complex. Upon initiation of transcription, RNA polymerase begins to synthesize messenger RNA (mRNA). Elongation of the mRNA occurs and once the strand is completely synthesized, transcription is terminated. The resulting mRNA copies of the gene encored proteins to be synthesized during the process of translation.
[0188] In some embodiments, a POI is a transcriptional core component involved in the initiation of transcription and the transcription initiation complex or pre-initiation complex (PIC). In order to initiate transcription, RNA polymerase II enzyme (Pol II) must contact the promoter upstream of the gene coding region to be transcribed and released to start transcription. The process of initiation involves a number of transcription factors and cofactors that form a pre-initiation complex (PIC) around Pol II and the promoter region of DNA. The PIC includes three families of proteins: General Transcription Factors (GTFs), Mediator (coactivator), and Pol II (see Sanders, et al., Genomics, Circuits, and Pathways in Clinical Neuropsychiatry. Academic Press, 2016. 3-26). In some embodiments, a POI is a transcriptional core component is a transcriptional repressor.
[0189] In some embodiments, a POI comprises a GTF protein, or a functional portion thereof, and / or a protein that shares a characteristic sequence element therewith. GTF proteins form six key subunits that make up portions of the PIC. In some embodiments, a POI comprises TFIID, or a functional portion thereof, and / or a protein that shares a Page 60 of 16513092399vlAttorney Docket No. 2013763-0005characteristic sequence element therewith. TFIID binds to the promote region through interaction with TATA-box binding protein (TBP). In some embodiments, a POI comprises TFIIB, or a functional portion thereof, and / or a protein that shares a characteristic sequence element therewith. TFIIB also targets B recognition elements near the promoter and selection of the transcription start site (TSS). In some embodiments, a POI comprises TFIIH, or a functional portion thereof, and / or a protein that shares a characteristic sequence element therewith f. TFIIH includes helicases that unwind DNA. In some embodiments, a POI comprises TFIIE, or a functional portion thereof, and / or a protein that shares a characteristic sequence element therewith. TFIIE works to recruit TFIIH to the PIC. In some embodiments, a POI comprises TFIIF, or a functional portion thereof, and / or a protein that shares a characteristic sequence element therewith. TFIIF recruits Pol II to the PIC. In some embodiments, a POI comprises TFIIA, or a functional portion thereof, and / or a protein that shares a characteristic sequence element therewith. TFIIA stabilizes the binding of TFIID to the DNA (see Sanders, et al., 2016). In some embodiments, a POI comprises a general transcription factor present in specific cell types, e.g., general transcription factor Hi (GTF2I). GTF2I forms a complex TFII-I (see Sanders, et al., 2016), or a functional portion thereof, and / or a protein that shares a characteristic sequence element therewith.
[0190] Pol II is an enzyme that transcribes DNA into RNA. The PIC recruits Pol II to the TSS by the PIC. Pol II is bound by TFIIF and the mediator complex. Pol II is composed of 12 subunits named POLR2Ato POLR2L in humans (Sainsbury et al., 2015). In some embodiments a POI comprises Pol II, or a functional portion thereof, and / or a protein that shares a characteristic sequence element therewith (e.g., POLR2A, POLR2B, POLR2C, POLR2D, POLR2E, POLR2F, POLR2G, POLR2H, POLR2I, POLR2J, POLR2K, or POLR2L). In some embodiments, a Pol II comprises a sequence represented in SEQ ID NO: 28, or a functional portion thereof, and / or that shares a characteristic sequence element therewith.
[0191] In some embodiments, a POI comprises a mediator factor (i.e., coactivator) involved in transcription. Mediator works with transcription factors at other DNA sites (e.g., enhancers) and other cofactors to influence transcription initiation. Components of mediator have been shown to vary by cell type and species and different subunits bind to different transcription factors. Mediator modulates PIC binding at the promoter and the stability of the Page 61 of 16513092399vlAttorney Docket No. 2013763-0005PIC (see Sanders, et al., 2016). In some embodiments, a POI comprises a mediator subunit (e.g., mediator complex subunit 23 and mediator complex subunit 12).
[0192] In some embodiments, a POI is TATA-binding protein (TBP), e.g., human TBP (hTBP), or a functional portion thereof, and / or a protein that shares a characteristic sequence element therewith. Protein-coding genes comprises characteristic sequences of nucleotides e.g., T-A-T-A-a / t-A-a / t, also known as TATA box, upstream of the start site of transcription. TBP targets and binds to this sequence upstream of the gene coding region and marks the start site of transcription for Pol II. In some embodiments, a POI is TATA-binding protein 2 (TBP2), or a functional portion thereof, and / or a protein that shares a characteristic sequence element therewith. In some embodiments, a protein of interest comprises an amino acid sequence that has at least 90% identity to SEQ ID NO: 29, or to a functional portion thereof, and / or that shares a characteristic sequence element therewith. In some embodiments, a protein of interest comprises an amino acid sequence as shown in SEQ ID NO: 29, or a functional portion thereof, and / or that shares a characteristic sequence element therewith. In some embodiments, the protein of interest is encoded by a nucleic acid sequence that has at least 90% identity to SEQ ID NO: 5. In some embodiments, a protein of interest is encoded by a nucleic acid sequence as shown in SEQ ID NO: 5.
[0193] In some embodiments, an expressed antibody of the present disclosure may be uniformly purified after being isolated from a host cell. Isolation and / or purification of an antibody agent of the present disclosure may be performed by a conventional method for isolating and purifying a protein. For example, not wishing to be bound by theory, an antibody agent of the present disclosure can be recovered and purified from recombinant cell cultures by well-known methods including, but not limited to, protein A purification, protein G purification, ammonium sulfate or ethanol precipitation, acid extraction, anion or cation exchange chromatography, phosphocellulose chromatography, hydrophobic interaction chromatography, affinity chromatography, hydroxylapatite chromatography and lectin chromatography. High performance liquid chromatography ("HPLC") can also be employed for purification. See, e.g., Colligan, Current Protocols in Immunology, or Current Protocols in Protein Science, John Wiley & Sons, NY, N.Y., (1997-2001), e.g., chapters 1, 4, 6, 8, 9, and 10, each entirely incorporated herein by reference. In some embodiments, an antibody ofPage 62 of 16513092399vlAttorney Docket No. 2013763-0005the present disclosure may be isolated and / or purified by additionally combining filtration, superfiltration, salting out, dialysis, etc.
[0194] Purified antibody agents of the present disclosure can be characterized by, for example, ELISA, ELISPOT, flow cytometry, immunocytology, BIACORE™ analysis, SAPID YNE KINEXA™ kinetic exclusion assay, SDS-PAGE and Western blot, or by HPLC analysis as well as by a number of other functional assays disclosed herein.3. Linkers
[0195] In some embodiments, a fusion protein described herein comprises a linker. A fusion protein, as described herein, may comprise (i) a protein of interest (POI) (e.g., a transcriptional machinery component), (ii) a linker, and (iii) a proximity labeling enzyme, where the linker connects the protein of interest to the proximity labeling enzyme. In some embodiments, a linker is utilized in order to control the labeling radius (i.e., the space allowed between the fusion protein and interacting protein in order for the proximity labeling enzyme to reach and label the interacting protein). Various linkers are contemplated to be used in generating fusion proteins described herein.
[0196] In some embodiments, a linker includes a flexible linker so as to provide flexibility at one or more locations within fusion protein described herein (e.g., between a POI or binding agent and a proximity labeling enzyme). In some embodiments, a flexible linker contains at least 1 flexible amino acid (e.g., Gly).
[0197] Exemplary flexible linkers include glycine polymers (G)n, glycine-serine polymers (including, for example, (GS)n, (GSGGS: SEQ ID NO: ll)n and (GGGS: SEQ ID NO: 12)n, where n is an integer of at least one), glycine -alanine polymers, alanine-serine polymers, and other flexible linkers known in the art. Glycine and glycine-serine polymers are relatively unstructured, and therefore may be able to serve as a neutral tether between components. Glycine accesses significantly more phi-psi space than even alanine, and is much less restricted than residues with longer side chains (see Scheraga, Rev. Computational Chem. 11: 173-142 (1992)). Exemplary flexible linkers include, but are not limited Gly-Gly-Ser-Gly: SEQ ID NO: 13, Gly-Gly-Ser-Gly-Gly: SEQ ID NO: 14, Gly-Ser-Gly-Ser-Gly: SEQ ID NO: 15, Gly-Ser-Gly-Gly-Gly: SEQ ID NO: 16, Gly-Gly-Gly-Ser-Gly: SEQ ID NO: 17,Page 63 of 16513092399vlAttorney Docket No. 2013763-0005Gly-Ser-Ser-Ser-Gly: SEQ ID NO: 18, and the like. Additional exemplary linkers also include the following:
[0198] GGGGSGGGGSGGGGS (SEQ ID NO: 19)
[0199] GGGGSGGGGSGGGGSGGGGS (SEQ ID NO: 20)
[0200] GGGGSGGGGSGGGGSGGGGSSGGGGS (SEQ ID NO: 21)
[0201] In some embodiments, a linker comprises a glycine asparagine linker (GGNN: SEQ ID NO: 22)n or (NNGG: SEQ ID NO: 23)n. In some embodiments, a linker comprises a glycine asparagine linker represented in SEQ ID NO: 24; GGNNGGNNGGNNGG.
[0202] In some embodiments, a linker sequence (e.g., used to link a nanobody to a proximity labeling enzyme to generate a fusion protein described herein) comprises an amino acid sequence SEQ ID NO: 25; DPPVAT. In some embodiments, a linker sequence (e.g., used to link a nanobody to a proximity labeling enzyme to generate a fusion protein described herein) is encoded by a nucleic acid sequence comprising (SEQ ID NO: 6).
[0203] The ordinarily skilled artisan will recognize that design of fusion proteins described herein can include linkers that are all or partially flexible, such that the linker can include a flexible linker as well as one or more portions that confer less flexible structure to provide for a desired molecule structure.
[0204] Additional exemplary suitable linkers can be readily selected and can be of various lengths, such as from 1 amino acid (e.g., Gly) to 20 amino acids, from 2 amino acids to 15 amino acids, from 3 amino acids to 12 amino acids, including 4 amino acids to 10 amino acids, 5 amino acids to 9 amino acids, 6 amino acids to 8 amino acids, or 7 amino acids to 8 amino acids (e.g., at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, or 19 amino acids).
[0205] As described herein, the length of the linker provides a labeling radius for the proximity labeling enzyme. In some embodiments, all proteins present within the vicinity of the proximity labeling enzyme may be tagged. The present disclosure provides fusion proteins with a range of labeling radii, from about 500 nm to less than 10 nm, e.g., tagging Page 64 of 16513092399vlAttorney Docket No. 2013763-0005radii of about 500 nm, about 400 run, about 300 nm, about 250 nm, about 200 nm, about 100 nm, about 90 nm, about 80 nm, about 70 nm, about 60 nm, about 50 nm, about 40 nm, about 30 nm, about 20 nm, about 10 nm, about 5 nm, about 2.5 nm, or about 1 nm. In some embodiments, a linker is configured to be a length such that the labeling radii is between about 500 nm to less than 10 nm, e.g., about 500 nm, about 400 nm, about 300 nm, about 250 nm, about 200 nm, about 100 nm, about 90 nm, about 80 nm, about 70 nm, about 60 nm, about 50 nm, about 40 nm, about 30 nm, about 20 nm, about 10 nm, about 5 nm, about 2.5 nm, or about 1 nm. In some embodiments, a linker is configured to be a length such that the labeling radii is between about lOnm to lOOnm.
[0206] In some embodiments, a linker is connected to the N-terminus of a protein of interest. In some embodiments, a linker is connected to the C-terminus of a protein of interest.4. Fusion Protein Configurations
[0207] Fusion proteins, as described herein, may comprise (i) a protein of interest (POI) (e.g., a transcriptional machinery component), (ii) a linker, and (iii) a proximity labeling enzyme, where the linker connects the protein of interest to the proximity labeling enzyme. In some embodiments, in addition to or instead of containing a POI, a fusion protein may comprise (i) a binding agent that binds to a POI, (ii) a linker, and (iii) a proximity labeling enzyme. A fusion protein, as described herein, may use a proximity labeling enzyme to label proteins that interact with the protein of interest when they are in proximity therewith. Various configurations may be utilized in order to generate a fusion protein described herein. Such configurations may take into account the particular POI, the target cell, and / or the labeling radii.
[0208] The present disclosure recognizes that there are multiple configurations in which fusion proteins described herein can be configured. POIs (and binding agents that target POIs) and proximity labeling enzymes can be joined through linker in various ways.
[0209] A fusion protein can have an N-terminus and a C-terminus. If a first entity (e.g., a POI or binding agent that binds to a POI) is “upstream” of a second entity (e.g., a proximity labeling enzyme), the first entity is closer to the N-terminus than the second entity.Page 65 of 16513092399vlAttorney Docket No. 2013763-0005Conversely, if a first entity is “downstream” of a second entity, the first entity is closer to the C-terminus than the second entity.
[0210] In some embodiments, a POI or binding agent that binds to a POI can be upstream of a proximity labeling enzyme in a fusion protein. In some embodiments, POI or binding agent that binds to a POI can be downstream of a proximity labeling enzyme in a fusion protein.
[0211] In some embodiments, a POI or binding agent that binds to a POI and a proximity labeling enzyme of a fusion protein can be directly linked. In some embodiments, a POI or binding agent that binds to a POI and a proximity labeling enzyme can be linked indirectly, e.g., via a linker. In some embodiments, a linker can comprise a glycine -arginine linker (including, e.g., SEQ ID NO: 24), or any of the linkers described herein.
[0212] In addition to a POI or binding agent that binds to a POI and a proximity labeling enzyme, and a linker, a fusion protein can include additional components including a regulatory moiety or signaling constituent, regulatory elements, signal sequences, and tags.
[0213] In various embodiments, a fusion protein described herein can include a tag. In some embodiment, a tag comprises a fluorescent tag. In some embodiments, a tag comprises an epitope tag. In some embodiments, an epitope tag comprises a V5 epitope tag (e.g., as shown in SEQ ID NO: 2).
[0214] In some embodiments, a linker is connected to the N-terminus of a protein of interest. In some embodiments, a linker is connected to the C-terminus of a protein of interest. In some embodiments, a linker is connected to the N-terminus of a proximity labeling enzyme. In some embodiments, a linker is connected to the C-terminus of a proximity labeling enzyme.
[0215] In some embodiments, a proximity labeling enzyme is connected to a linker that connects to the N-terminus of the POI. In some embodiments, a proximity labeling enzyme is connected to a linker that connects to the C-terminus of the POI.
[0216] In some embodiments, a fusion protein comprises (i) a POI, (ii) a linker, and (iii) a proximity labeling enzyme. In some embodiments, the POI is a core transcriptional Page 66 of 16513092399vlAttorney Docket No. 2013763-0005component such as TBP (e.g., TBP2) or Pol II. In some embodiments, a TBP POI comprises SEQ ID NO: 29, or a functional portion thereof, and / or that shares a characteristic sequence element therewith. In some embodiments, a Pol II POI comprises SEQ ID NO: 29, or a functional portion thereof, and / or that shares a characteristic sequence element therewith. In some embodiments, a proximity labeling enzyme comprises one of APEX (e.g., APEX2) or Turbo-ID. In some embodiments an APEX2 enzyme comprises SEQ ID NO: 31, or a fragment or variant thereof. In some embodiments, a Turbo-ID enzyme comprises SEQ ID NO: 32, or a fragment or variant thereof. In some embodiments, a linker comprises a glycine-arginine linker (e.g., as shown in SEQ ID NO: 24). As described herein, in some embodiments, a fusion protein may comprise one of the following configurations (from N-terminus to C-terminus):
[0217] (i) TBP (SEQ ID NO: 29), Gly-Arg linker (SEQ ID NO: 24), APEX2 (SEQ ID NO: 31)
[0218] (ii) TBP (SEQ ID NO: 29), Gly-Arg linker (SEQ ID NO: 24), APEX
[0219] (iii) TBP (SEQ ID NO: 29), Gly-Arg linker (SEQ ID NO: 24), Turbo-ID
[0220] (iv) TBP (SEQ ID NO: 29), Gly-Arg linker (SEQ ID NO: 24), mini Turbo-ID
[0221] (v) APEX2 (SEQ ID NO: 31), Gly-Arg linker (SEQ ID NO: 24), TBP (SEQ ID NO: 29)
[0222] (vi) APEX, Gly-Arg linker (SEQ ID NO: 24), TBP (SEQ ID NO: 29)
[0223] (vii) Turbo-ID, Gly-Arg linker (SEQ ID NO: 24), TBP (SEQ ID NO: 29)
[0224] (viii) mini-Turbo-ID, Gly-Arg linker (SEQ ID NO: 24), TBP (SEQ ID NO: 29)
[0225] In some embodiments, a fusion protein may comprise, in addition to any of the above components, a tag located e.g., between any one of the components (e.g., included in a nucleic acid sequence between sequences encoding any one of the described components).
[0226] In some embodiments, a fusion protein comprises a binding agent that binds a POI. In some embodiments, a binding agent is an antibody agent as described herein. In Page 67 of 16513092399vlAttorney Docket No. 2013763-0005some embodiments, a fusion protein comprises an antibody agent that binds a protein of interest linked to a proximity labeling enzyme. In some embodiments, an antibody agent is covalently linked to a proximity labeling enzyme. In some embodiments, an antibody agent is covalently linked to a proximity labeling enzyme via a linker. A linker may include any of the linkers described herein or known in the art. An antibody agent may include any of the antibody agents described herein or known in the art.
[0227] In some embodiments, an antibody agent is connected to a linker through its variable domain (e.g., an amino acid within the antibody agent variable domain). In some embodiments, a linker is connected to an antibody agent by cross-linking to an internal ammo acid of an antigen -binding domain of the antibody agent In some embodiments, a linker is connected to the N-terminus of an antibody agent. In some embodiments, a linker is connected to the C-terminus of an antibody agent. In some embodiments, an antibody agent comprises a nanobody. In some embodiments, an antibody agent comprises a mintbody. In some embodiments, an antibody agent comprises a configuration as shown in FIG. 10.
[0228] In some embodiments, a fusion protein may comprise one of the following configurations (from N-terminus to C-terminus):
[0229] (i) VH-VL-Linker-Proximity Labeling Enzyme
[0230] (ii) VL-VH-Linker-Proximity Labeling Enzyme
[0231] (iii) VH- Proximity Labeling Enzyme
[0232] (iv) VHH- Proximity Labeling Enzyme
[0233] (v) Proximity Labeling Enzyme -Linker- VH-VL
[0234] (vi) Proximity Labeling Enzyme-Linker- VL-VH
[0235] (vii) Proximity Labeling Enzyme-Linker- VH
[0236] (viii) Proximity Labeling Enzyme -Linker- VHH,Page 68 of 16513092399vlAttorney Docket No. 2013763-0005
[0237] In some embodiments, such configurations include a proximity labeling enzyme that is APEX, APEX2, Turbo-ID, or mini-Turbo-ID and variable domain(s) that specifically bind Pol II in an activated state (e.g., containing a Ser2 phosphorylation). In some embodiments, a fusion protein comprises a configuration as shown in FIG. 10.
[0238] In some embodiments, a fusion protein is encoded by a nucleic acid sequence that has at least 80%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 93%, 95%, 96%, 97%.98%, or 99% or greater identity to SEQ ID NO: 1, or a fragment of variant thereof. In some embodiments, a fusion protein is encoded by an nucleic acid sequence that has at least 80%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 93%, 95%, 96%, 97%, 98%, 99% or greater identity to SEQ ID NO: 6. In some embodiments, a fusion protein is encoded by an nucleic acid sequence that comprises SEQ ID NO: 1, or a fragment of variant thereof. In some embodiments, a fusion protein is encoded by a nucleic acid sequence that comprises SEQ ID NO: 6, or a fragment of variant thereof.
[0239] In some embodiments, a fusion protein part of a proximity labelling system (e.g., a split proximity labeling system). In some embodiments, a fusion protein comprises a protein of interest (POI) and a protein fragment, wherein the protein fragment comprises a fragment of a functional indicator protein (e.g., a green fluorescent protein (GFP) fragment, e.g., GFP10 or GFP11). In some embodiments, a fusion protein comprises a proximity labelling enzyme and a protein fragment, wherein the protein fragment comprises another fragment of a functional indicator protein. In some embodiments, a system comprises two or more fusion proteins, where each fusion protein contains a fragment of an indicator protein, where the fragments associate when in proximity to one another and form a functional indicator protein (e.g., GFP). In some embodiments, a functional indicator protein is GFP that fluoresces when its fragments are associated and it is fully assembled. Such fusion proteins allow for conditional visualization of complexes that have been associated and where proximity labeling has occurred. In some embodiments, a fusion protein may be part of a proximity labelling system (e.g., a split proximity labeling system), that contains one or more fusion proteins and targets a particular target genomic locus.
[0240] In some embodiments, a fusion protein comprises a POI and a protein fragment, wherein the first protein fragment comprises a fragment of a functional indicator Page 69 of 16513092399vlAttorney Docket No. 2013763-0005protein (e.g., a green fluorescent protein (GFP) fragment, e.g., GFP10 or GFP11). In some embodiments, a fusion protein comprises a DNA-targeting moiety (e.g., dCas9 or LNA) that binds to a target genomic locus and a protein fragment, wherein the protein fragment comprises fragment of a functional indicator protein (e.g., a green fluorescent protein (GFP) fragment, e.g., GFP10 or GFP11). In some embodiments, a fusion protein comprises a proximity labeling enzyme (e.g., APEX2) and a protein fragment, wherein the protein fragment comprises another fragment of a functional indicator protein (e.g., a GFP fragment, e.g., GFPl-9r).
[0241] In some embodiments, a fusion protein comprises one of the following configurations (from N-terminus to C-terminus):
[0242] (i) POI (e.g., TBP) - Linker - GFP10
[0243] (ii) POI (e.g., TBP) - Linker - GFP11
[0244] (in) APEX2 - Linker - GFPl-9r
[0245] (iv) dCas9 - Linker - GFP 10
[0246] (v) dCas9 - Linker - GFP11
[0247] (vi) LNA- Linker- GFP 10
[0248] (vii) LNA- Linker- GFP 11
[0249] (viii) TurboID - Linker - GFPl-9r
[0250] (ix) mini-Turbo-ID - Linker - GFPl-9r
[0251] In some embodiments, a fusion protein described herein comprises a fusion as shown in e.g., FIGs. 12-14. In some embodiments, a fusion protein is encoded by a nucleic acid sequence having at least 80% identity to SEQ ID NO: 40 (e.g., at least 80%, at least 85%, at least 90%, at least 95% or 100% identical to SEQ ID NO: 40). In some embodiments, a fusion protein comprises an amino acid sequence having at least 90% identity to SEQ ID NO: 40 (e.g., at least 90%, 91%, 92%, 93 %, 94%, 95%, 96%, 97%, 98%,Page 70 of 16513092399vlAttorney Docket No. 2013763-000599% or 100% identity to SEQ ID NO: 40). In some embodiments, a fusion protein is encoded by a nucleic acid sequence having at least 80% identity to SEQ ID NO: 42 (e.g., at least 80%, at least 85%, at least 90%, at least 95% or 100% identical to SEQ ID NO: 42). In some embodiments, a fusion protein comprises an amino acid sequence having at least 90% identity to SEQ ID NO: 43 (e.g., at least 90%, 91%, 92%, 93 %, 94%, 95%, 96%, 97%, 98%, 99% or 100% identity to SEQ ID NO: 43). In some embodiments, a fusion protein is encoded by a nucleic acid sequence having at least 80% identity to SEQ ID NO: 52 (e.g., at least 80%, at least 85%, at least 90%, at least 95% or 100% identical to SEQ ID NO: 52). B. Proximity Labeling Systems
[0252] In some embodiments, one or more fusion proteins described herein may be utilized in a proximity labeling system, e.g., a “split proximity labelling system,” e.g., a tripartite proximity labelling system, also referred to herein as “a three-component proximity labelling system”.
[0253] Spit proximity labeling systems (e.g., split GFP proximity labelling systems described herein) can be utilized to study and identify protein-protein and DNA-protein interactions in cells. In some embodiments, a split proximity labelling system includes one or more fusion proteins described herein, e.g., two or more fusion proteins described herein. In some embodiments, a split proximity labelling system comprises a first fusion protein and a second fusion protein, where a first fusion protein comprises a first POI and a first protein fragment and the second fusion protein comprises a proximity labelling enzyme and a second protein fragment, where the first and second protein fragments are fragments of a functional indicator protein, that is only functional when the protein fragments are associated. In some embodiments, a proximity labeling system further comprises a third fusion protein, where the third fusion protein comprises a second POI and third protein fragment. In some embodiments, a first fusion protein, second fusion protein, and third fusion protein of a split proximity labelling system described herein, when in proximity to one another, associate to form a complex via the association of the first protein fragment, second protein fragment, and third protein fragment, thereby forming the functional indicator protein. In some embodiments, a functional indicator protein is only functional when the first, second andPage 71 of 16513092399vlAttorney Docket No. 2013763-0005third protein fragments of the functional indicator protein are associated, such that the functional indicator protein is reconstituted.
[0254] Systems described herein also allow for targeting particular DNA loci, and for pulling down protein-DNA and DNA-DNA complexes (e.g., transcriptional complexes, e.g., of activation and / or repression nature) from cells (e.g., living human cells). In some embodiments, systems described herein may be used in methods relating to pull down of a DNA region of interest (target genomic locus) and labeling the interactome of a protein of interest (POI) or gene of interest (GOI) with great precision. In some embodiments, a proximity labelling system as described herein comprises a split proximity labelling system that includes a DNA-targeting moiety. For example, in some embodiments, a system comprises a first fusion protein, a second fusion protein, and a third fusion protein, where each fusion protein comprises a first, second and third protein fragment of a functional indicator protein, respectively. In some embodiments, a first fusion protein of a system comprises a POI and a first protein fragment of a functional indicator protein. In some embodiments, a second fusion protein of a system comprises a DNA-targeting moiety that binds to a target genomic locus and a second protein fragment of a functional indicator protein. In some embodiments, a third fusion protein of a system comprises a proximity labelling enzyme and a third protein fragment of a functional indicator protein. In some embodiments, a first fusion protein, second fusion protein, and third fusion protein of a split proximity labelling system described herein, including a DNA-targeting moiety, when in proximity to one another, associate to form a complex via the association of the first protein fragment, second protein fragment, and third protein fragment, thereby forming the functional indicator protein. In some embodiments, a functional indicator protein is only functional when the first, second and third protein fragments of the functional indicator protein are associated, such that the functional indicator protein is reconstituted.Indicator Proteins
[0255] In some embodiments, an indicator protein utilized in a proximity labelling system as described herein includes a functional indicator protein that can indicate presence of the fusion protein within a cell. In some embodiments, a functional indicator protein as part of a split proximity labelling system as described herein, is only functional as an Page 72 of 16513092399vlAttorney Docket No. 2013763-0005indicator protein when fragments of the functional indicator protein utilized in fusion proteins of the system are associated and the functional indicator protein is reconstituted.
[0256] In some embodiments, the indication is by visualization under a microscope (e.g., a fluorescence microscope), e.g., a functional indicator protein is a fluorescent protein. In some embodiments, a fluorescent protein comprises a green fluorescent protein (GFP). In some embodiments a GFP is used in a split GFP proximity labelling system as described herein. In such a system, a GFP protein is split into halves including a first fragment GFP Beta 10 (GFP10), and second fragment GFP Beta 11 (GFP 11), e.g., encoded by the nucleic acid sequence of SEQ ID NO: 37. Both fragments are not fluorescent on their own. Only when in proximity to each do they form a complex which can then be recognized by the third major fragment, called GFP l-9r, e.g., as shown in SEQ ID NO: 38. GFP l-9r recognizes and forms a complex with dimerized GFP 10 and GFP 11 fragments, forming the whole fluorescent GFP (FIG. 12). Under normal cellular conditions, the entropy of the GFP fragments is high such that they cannot reconstitute to a functional protein.DN A-Targcting Moiety
[0257] In some embodiments, fusion proteins and system described herein include a DNA-targeting moiety. In some embodiments, In some embodiments, a DNA-targeting moiety comprises a nucleic acid sequence (e.g., an oligonucleotide) that is complementary to a target genomic locus. In some embodiments, a DNA-targeting moiety comprises an antisense oligonucleotide (ASO). In some embodiments, a DNA-targeting moiety comprises a DNA molecule. In some embodiments, a DNA-targeting moiety comprises an RNA molecule.
[0258] In some or any embodiments, a DNA-targeting moiety is an oligomeric base compound or oligonucleotide mimetic that hybridizes to at least a portion of a nucleic acid sequence of a target genomic locus, and e.g., modulates its function. In some embodiments, a DNA-targeting moiety is single stranded or double stranded. A variety of exemplary DNA-targeting moieties are known and described in the art. In some embodiments, a DNA-targeting moiety is an antisense oligonucleotide, locked nucleic acid (LNA) molecule, peptide nucleic acid (PNA) molecule, ribozyme, siRNA, antagomirs, external guide sequence (EGS)Page 73 of 16513092399vlAttorney Docket No. 2013763-0005oligonucleotide, microRNA (miRNA), small, temporal RNA (stRNA), guide RNAs (gRNA), or single- or double-stranded RNA interference (RNAi) compounds.Locked Nucleic Acids (LNAs)
[0259] The present disclosure provides, among other things, methods of targeting and isolating, also referred to as “pull down,” and / or labelling of specific target genomic loci (e.g., transcriptional and / or regulatory regions within a target locus). The present disclosure provides systems that comprise specific targeting mechanism (e.g., DNA-targeting moiety) for locus pull down using systems described herein, including the use of locked nucleic acids (LNAs).
[0260] It is understood that the term “LNA” or “LNA molecule” refers to a molecule that comprises at least one LNA modification; thus LNA molecules may have one or more locked nucleotides (conformationally constrained) and one or more non-locked nucleotides. It is also understood that the term “LNA” includes a nucleotide that comprises any constrained sugar that retains the desired properties of high affinity binding to a complementary target genomic locus, nuclease resistance, lack of immune stimulation, and rapid kinetics. Similarly, it is understood that the term “PNA molecule” refers to a molecule that comprises at least one PNA modification and that such molecules may include unmodified nucleotides or intemucleoside linkages. LNAs are modified RNAs containing a bridge between the 2’ oxygen and 4’ carbon of the ribose, locking the ribose ring in a specific conformation. Such a locked conformation provides strong binding affinity for complementary DNA and RNA (e.g., on a target genomic locus).
[0261] In some or any embodiments, an LNA comprises at least one nucleotide and / or nucleoside modification (e.g., modified bases or with modified sugar moieties), modified intemucleoside linkages, and / or combinations thereof. In some embodiments, LNAs can comprise natural as well as modified nucleosides and linkages. Examples of such chimeric LNAs, including hybrids or gapmers, are described below.
[0262] In some embodiments, an LNA comprises one or more modifications comprising: a modified sugar moiety, and / or a modified intemucleoside linkage, and / or a modified nucleotide and / or combinations thereof. In some embodiments, the modified Page 74 of 16513092399vlAttorney Docket No. 2013763-0005intemucleoside linkage comprises at least one of: alkylphosphonate, phosphorothioate, phosphorodithioate, alkylphosphonothioate, phosphoramidate, carbamate, carbonate, phosphate triester, acetamidate, carboxymethyl ester, or combinations thereof. In some embodiments, the modified sugar moiety comprises a 2'-O-methoxyethyl modified sugar moiety, a 2'-methoxy modified sugar moiety, a 2'-O-alkyl modified sugar moiety, or a bicyclic sugar moiety. Other examples of modifications include locked nucleic acid (LNA), peptide nucleic acid (PNA), arabinonucleic acid (ANA), optionally with 2'-F modification, 2'-fluoro-D-Arabinonucleic acid (FANA), phosphorodiamidate morpholino oligomer (PMO), ethylene-bridged nucleic acid (ENA), optionally with 2'-0,4'-C-ethylene bridge, and bicyclic nucleic acid (BN A).
[0263] In some embodiments, a DNA-targeting moiety a nucleic acid sequence that is 5-40 bases in length (e.g., 12-30, 12-28, 12-25). In some embodiments, a nucleic acid sequence may also be 10-50, or 5-50 bases length. For example, a nucleic acid sequence may be one of any of 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, or 50 bases in length. In some embodiments, a nucleic acid is double stranded and comprises an overhang (optionally 2-6 bases in length) at one or both termini. In other embodiments, a nucleic acid sequence is double stranded and blunt-ended. In some embodiments, a nucleic acid sequence comprises a sequence of bases at least 80% or 90% complementary to, e.g., at least 5, 10, 15, 20, 25 or 30 bases of, or up to 30 or 40 bases of, a sequence within or comprising a target genomic locus, or comprises a sequence of bases with up to 3 mismatches (e.g., up to 1, or up to 2 mismatches) over 10, 15, 20, 25 or 30 bases of a sequence within or comprising a target genomic locus.
[0264] In some embodiments, a nucleic acid sequence of a DNA-targeting moiety as described herein can comprise or consist of a sequence of bases at least 80% complementary to at least 10 contiguous bases of the target RNA, or at least 80% complementary to at least 15, or 15-30, or 15-40 contiguous bases of sequence within or comprising a target genomic locus, or at least 80% complementary to at least 20, or 20-30, or 20-40 contiguous bases of within or comprising a target genomic locus, or at least 80% complementary to at least 25, or 25-30, or 25-40 contiguous bases of the target RNA, or at least 80% complementary to at least 30, or 30-40 contiguous bases of within or comprising a target genomic locus, or at least Page 75 of 16513092399vlAttorney Docket No. 2013763-000580% complementary to at least 40 contiguous bases of within or comprising a target genomic locus. Moreover, in some embodiments, a nucleic acid sequence of a DNA-targeting moiety as described herein can comprise or consist of a sequence of bases at least 90% complementary to at least 10 contiguous bases of a sequence within or comprising a target genomic locus, or at least 90% complementary to at least 15, or 15-30, or 15-40 contiguous bases of a sequence within or comprising a target genomic locus, or at least 90% complementary to at least 20, or 20-30, or 20-40 contiguous bases of a sequence within or comprising a target genomic locus, or at least 90% complementary to at least 25, or 25-30, or 25-40 contiguous bases of a sequence within or comprising a target genomic locus, or at least 90% complementary to at least 30, or 30-40 contiguous bases of a sequence within or comprising a target genomic locus, or at least 90% complementary to at least 40 contiguous bases of a sequence within or comprising a target genomic locus. Similarly, a nucleic acid sequence of a DNA-targeting moiety described herein can comprise or consist of a sequence of bases fully complementary to at least 5, 10, or 15 contiguous bases of a sequence within or comprising a target genomic locus.dCas9 and gRNAs
[0265] In some embodiments, the present disclosure provides systems that comprise specific targeting mechanisms (e.g., DNA-targeting moiety) for locus pull down and / or proximity labelling using systems described herein, including the use of DNA-targeting moiety such as a guide RNA (gRNA), e.g., as part of a Cas9 protein, e.g., a dead Cas9 or an enzyme deficient Cas9.
[0266] In some embodiments, a split proximity labelling system (e.g., a split GFP proximity labelling system) can be used to isolated and / or pull down, and / or label transcriptional complexes (e.g., transcription complexes comprising transcriptional elements that repress or initial transcription of ataiget genomic locus) by using a DNA-targeting moiety comprising a gRNA that directs one or more of the fusion proteins of the system to the target genomic locus. In some embodiments, a gRNA is part of a Cas9 protein, e.g., a dead Cas9 or an enzyme deficient Cas9. In some embodiments, such a system comprises a fusion protein comprising a protein fragment (e.g., a protein fragment of a functional indicator protein, e.g., a GFP fragment, e.g., GFP 10) fused to dCas9. In some embodiments dCas9 protein is fused Page 76 of 16513092399vlAttorney Docket No. 2013763-0005to a protein fragment via a linker. In some embodiments, when a fusion protein comprising a protein fragment and a dCas9 protein is expressed in a population of cells, the fusion protein localizes at the target locus through its specific gRNA and binds to the target locus (or an interacting component at the target locus). In some embodiments, other fusion proteins within a split proximity labelling system described herein may co-localize via their protein fragments to the target locus bound by the dCas9. In some embodiments, a protein complex (including the associated fusion proteins) can be detected by a functional indicator protein and / or via proximity labelling and pull down via by mass spectrometry.
[0267] In some embodiments, a DNA-targeting moiety comprises a gRNA. In some embodiments a gRNA is provided in a system with a “dead Cas9” or “endonuclease deficient Cas9” or “dCas9”. dCas9 is capable of binding to DNA and guide RNA but does not contain a functional endonuclease for cutting targeting DNA. In some embodiments, a gRNA comprises a sequence that is complementary to sequence within or comprising a target genomic locus. In some embodiments, a gRNA comprises any one of the sequences shown in e.g., SEQ ID NOs: 44-47, or a sequence that is at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or greater identity to SEQ ID NOs: 44-47).
[0268] In some embodiments, a gRNA comprises a nucleic acid sequence that is 5-40 bases in length (e.g., 12-30, 12-28, 12-25). In some embodiments, a gRNA comprises a nucleic acid sequence may also be 10-50, or 5-50 bases length. For example, a gRNA comprises a nucleic acid sequence may be one of any of 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, or 50 bases in length. In some embodiments, a gRNA comprises a nucleic acid sequence comprises a sequence of bases at least 80% or 90% complementary to, e.g., at least 5, 10, 15, 20, 25 or 30 bases of, or up to 30 or 40 bases of, a sequence within or comprising a target genomic locus, or comprises a sequence of bases with up to 3 mismatches (e.g., up to 1, or up to 2 mismatches) over 10, 15, 20, 25 or 30 bases of a sequence within or comprising a target genomic locus.
[0269] In some embodiments, a gRNA comprises a nucleic acid sequence can comprise or consist of a sequence of bases at least 80% complementary to at least 10 contiguous bases of a sequence within or comprising a target genomic locus, or at least 80%Page 77 of 16513092399vlAttorney Docket No. 2013763-0005complementary to at least 15, or 15-30, or 15-40 contiguous bases of sequence within or comprising a target genomic locus, or at least 80% complementary to at least 20, or 20-30, or 20-40 contiguous bases of within or comprising a target genomic locus, or at least 80% complementary to at least 25, or 25-30, or 25-40 contiguous bases of a sequence within or comprising a target genomic locus, or at least 80% complementary to at least 30, or 30-40 contiguous bases of within or comprising a target genomic locus, or at least 80% complementary to at least 40 contiguous bases of within or comprising a target genomic locus. Moreover, in some embodiments, a gRNA comprises a nucleic acid sequence can comprise or consist of a sequence of bases at least 90% complementary to at least 10 contiguous bases of a sequence within or comprising a target genomic locus, or at least 90% complementary to at least 15, or 15-30, or 15-40 contiguous bases of a sequence within or comprising a target genomic locus, or at least 90% complementary to at least 20, or 20-30, or 20-40 contiguous bases of a sequence within or comprising a target genomic locus, or at least 90% complementary to at least 25, or 25-30, or 25-40 contiguous bases of a sequence within or comprising a target genomic locus, or at least 90% complementary to at least 30, or 30-40 contiguous bases of a sequence within or comprising a target genomic locus, or at least 90% complementary to at least 40 contiguous bases of a sequence within or comprising a target genomic locus. Similarly, a nucleic acid sequence of a DNA-targeting moiety described herein can comprise or consist of a sequence of bases fully complementary to at least 5, 10, or 15 contiguous bases of a sequence within or comprising a target genomic locus.C. Compositions
[0270] Various embodiments include delivery of compositions comprising fusion proteins, and proximity labelling systems as described herein, to a cell, subject, or other system. Those of skill in the art will appreciate delivery of a composition to a cell, subject, or other system of a composition can depend on the type of composition and that, for various types of compositions, means and techniques for delivery are known in the art.Nucleic Acids
[0271] In some embodiments, a composition includes and is formulated for delivery of a nucleic acid, in other words, a composition comprises a nucleic acid that encodes a fusion protein described herein. Nucleic acids encoding fusions proteins or components Page 78 of 16513092399vlAttorney Docket No. 2013763-0005thereof described herein can be used for delivery of a fusion protein inside cells and whole live organisms to effectuate proximity-dependent labeling of interacting proteins.
[0272] Coding sequences for fusion proteins described herein or components thereof can be isolated and / or synthesized and cloned into any suitable vector or replicon for expression in a suitable host cell or host subject. In some embodiments, a nucleotide sequence encoding a fusion protein described herein is integrated into a vector. Numerous vectors are known in the art including, but not limited to, linear polynucleotides, polynucleotides associated with ionic or amphiphilic compounds, plasmids, and viruses.
[0273] Non-limiting examples of viral vectors include: retrovirus, Harvey murine sarcoma virus, murine mammary tumor virus, Rous sarcoma virus), adenovirus, adeno-associated virus, SV40-type virus, polyomavirus, Epstein-Barr virus, papilloma virus, herpes virus, vaccinia virus, and polio virus. Exemplary classes of viral vectors include retroviral vectors (see, e.g., Axelrod et al., PNAS 87:5173-5177 (1990); Kay et al., Hum. Gene Then 3:641-647 (1992); Van den Driessche et al., PNAS 96:10379-10384 (1999); Xu et al., ASAIO J. 49:407-416 (2003); and Xu et al., PNAS 102:6080-6085 (2005)), lentiviral vectors (see, e.g., McKay et al., Curr. Pharm. Des. 17:2528-2541 (2011); Brown et al., Blood 109:2797-2805 (2007); and Matrai et al., Hepatology 53:1696-1707 (2011)), adeno-associated viral (AAV) vectors (see, e.g., Herzog et al., Blood 91:4600-4607 (1998)), and adenoviral vectors (see, e.g., Brown et al., Blood 103:804-810 (2004) and Ehrhardt et al., Blood 99:3923-3930 (2002)).
[0274] Retroviruses are enveloped viruses that belong to the viral family Retroviridae. Once in a host’s cell, the virus replicates by using a viral reverse transcriptase enzyme to transcribe its RNA into DNA. The retroviral DNA replicates as part of the host genome, and is referred to as a provirus. A selected nucleic acid can be inserted into a vector and packaged in retroviral particles using techniques known in the art. Protocols for the production of replication-deficient retroviruses are known in the art (see, e.g., Kriegler, M., Gene Transfer and Expression, A Laboratory Manual, W.H. Freeman Co., New York (1990) and Murry, E. J., Methods in Molecular Biology, Vol. 7, Humana Press, Inc., Cliffion, N.J. (1991)). The recombinant virus can then be isolated and delivered to cells of the subject either in vivo or ex vivo. A number of retroviral systems are known in the art, for example Page 79 of 16513092399vlAttorney Docket No. 2013763-0005See U.S. Pat Nos. 5,994,136, 6,165,782, and 6,428,953. Retroviruses include the genus of Alpharetrovirus (e.g., avian leukosis virus), the genus of Betaretrovirus; (e.g., mouse mammary tumor virus) the genus of Deltaretrovirus (e.g., bovine leukemia virus and human T-lymphotropic virus), the genus of Epsilonretrovirus (e.g., Walleye dermal sarcoma virus), and the genus of Lenti virus.
[0275] In some embodiments, the retrovirus is a lentivirus of the Retroviridae family. Lentiviral vectors can transduce non-proliferating cells and show low immunogenicity. In some examples, the lentivirus is, but is not limited to, human immunodeficiency viruses (HIV-1 and HIV-2), simian immunodeficiency virus (S1V), feline immunodeficiency virus (FIV), equine infections anemia (EIA), and visna virus. Vectors derived from lentiviruses can achieve significant levels of nucleic acid transfer in vivo.
[0276] In some embodiments, the vector is an adenovirus vector. Adenoviruses are a large family of viruses containing double stranded DNA. They replicate within the nucleus of a host cell, using the host’s cell machinery to synthesize viral RNA, DNA and proteins. Adenoviruses are known in the art to affect both replicating and non-replicating cells, to accommodate large transgenes, and to code for proteins without integrating into the host cell genome.
[0277] In some embodiments, the viral vector is an adeno-associated virus (AAV) vector. AAV systems are generally well known in the art (see, e.g., Kelleher and Vos, Biotechniques, 17(6): 1110-17 (1994); Cotten etal., P.N.A.S. U.S.A., 89(13):6094-98 (1992); Curiel, Nat Immun, 13(2-3): 141-64 (1994); Muzyczka, Curr Top Microbiol Immunol, 158:97-129 (1992); and Asokan A, et al., Mol. Then, 20(4):699-708 (2012)). Methods for generating and using recombinant AAV (rAAV) vectors are described, for example, in U.S. Pat. Nos. 5,139,941 and 4,797,368.
[0278] In addition to the major elements identified above for an viral vector, a vector can also include conventional control elements operably linked to the transgene in a manner that permits its transcription, translation and / or expression in a cell transfected with the vector or infected with the virus produced by the disclosure. Expression control sequences include appropriate transcription initiation, termination, promoter and enhancer sequences; efficient RNA processing signals such as splicing and polyadenylation (poly A) signals; sequences that Page 80 of 16513092399vlAttorney Docket No. 2013763-0005stabilize cytoplasmic mRNA; sequences that enhance translation efficiency (i.e., Kozak consensus sequence); sequences that enhance protein stability; and when desired, sequences that enhance secretion of the encoded product. A number of expression control sequences, including promoters that are native, constitutive, inducible and / or tissue-specific, are known in the art and may be included in a vector described herein. In some embodiments, operably linked coding sequences yield a functional RNA and protein.
[0279] Examples of constitutive promoters include, without limitation, the retroviral Rous sarcoma virus (RSV) LTR promoter (optionally with the RSV enhancer), the cytomegalovirus (CMV) promoter (optionally with the CMV enhancer), the SV40 promoter, and the dihydrofolate reductase promoter. Inducible promoters allow regulation of gene expression and can be regulated by exogenously supplied compounds, environmental factors such as temperature, or the presence of a specific physiological state, e.g, acute phase, a particular differentiation state of the cell, or in replicating cells only. Inducible promoters and inducible systems are available from a variety of commercial sources, including, without limitation, Invitrogen, Clontech and Ariad. Many other systems have been described and can be readily selected by one of skill in the art. Examples of inducible promoters regulated by exogenously supplied promoters include the zinc-inducible sheep metallothionine (MT) promoter, the dexamethasone (Dex) -inducible mouse mammary tumor virus (MMTV) promoter, the T7 polymerase promoter system, the ecdysone insect promoter, the tetracycline-repressible system, the tetracycline-inducible system, the RU486-inducible system and the rapamycin-inducible system. Still other types of inducible promoters which may be useful in this context are those which are regulated by a specific physiological state, e.g., temperature, acute phase, a particular differentiation state of the cell, or in replicating cells only. In another embodiment, a native promoter, or fragment thereof, for a transgene will be used. In a further embodiment, other native expression control elements, such as enhancer elements, polyadenylation sites or Kozak consensus sequences may also be used to mimic the native expression.
[0280] In some embodiments, regulatory sequences impart tissue-specific gene expression capabilities. In some cases, the tissue-specific regulatory sequences bind tissuespecific transcription factors that induce transcription in a tissue specific manner. Such tissue-specific regulatory sequences (e.g., promoters, enhancers, etc.) are well known in the Page 81 of 16513092399vlAttorney Docket No. 2013763-0005art. In some embodiments, the promoter is a chicken [3-actin promoter, a pol II promoter, or a pol III promoter.
[0281] In some embodiments, a viral vector (e.g., a lentiviral vector) comprises a DNA sequence encoding a fusion protein described herein (e.g., fusion proteins represented by SEQ ID NOs: 1 and 6).1. Cells
[0282] Among other things, the present disclosure provides a cell comprising a nucleic acid sequence described herein or a viral particle described herein. A variety of technologies are known to those skilled in the art for engineering any of a variety of cells to contain and / or express a nucleic acid as described herein (see, for example, Green & Sambrook Molecular Cloning, Cold Spring Harbor Laboratory Press).
[0283] A vector (e.g., an expression vector) can be transfected, transformed or transduced into a host cell. As used herein, the terms “transfection,” “transformation” and “transduction” all refer to the introduction of an exogenous nucleic acid sequence into a host cell. Host cells can be transfected, transformed, or transduced with nucleic acid coding sequences controlled by appropriate expression control elements (e.g., promoter, enhancer, sequences, transcription terminators, polyadenylation sites, etc.), and / or in conjunction with a selectable marker. In some embodiments, vectors including nucleic acid sequences encoding part of all of a fusion protein described herein are transfected, transformed or transduced into a host cell simultaneously and / or sequentially. Examples of transformation, transfection and transduction methods, which are well known in the art, include liposome delivery, i.e., lipofectamine™ (Gibco BRL) Method of Hawley-Nelson, Focus 15:73 (1193), electroporation, CaPO4 delivery method of Graham and van der Erb, Virology, 52:456-457 (1978), DEAE-Dextran medicated delivery, microinjection, biolistic particle delivery, polybrene mediated delivery, cationic mediated lipid delivery, transduction, and viral infection, such as, e.g., retrovirus, lentivirus, adenovirus adeno-associated virus and Baculovirus (Insect cells). General aspects of cell host transformations have been described in the art, such as by Axel in U.S. Pat. No. 4,399,216; Sambrook, supra, Chapters 1-4 and 16-18; Ausubel, supra, chapters 1, 9, 13, 15, and 16. For various techniques for transforming mammalian cells, see Keown et al., Methods in Enzymology (1989), Keown et al., Methods in Page 82 of 16513092399vlAttorney Docket No. 2013763-0005Enzymology, 185:527-537 (1990), and Mansour et al.. Nature, 336:348-352 (1988).Technologies for introducing nucleic acids into mammalian cells include transfection (e.g., mediated by cationic lipid reagents, by calcium phosphate, by DEAE-Dextran, by DOTMA / DOGS, by electroporation, and / or by combinations thereof) and use of viral vectors (e.g., adenoviral vectors, retroviral vectors, lentiviral vectors, and / or combinations thereof).
[0284] In some embodiments, a provided cell may transiently contain and / or express a nucleic acid and / or polypeptide disclosed herein.
[0285] In some embodiments, a provided cell may contain and / or express multiple copies or instances (e.g., 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23,24, 25, or more copies or instances) of a nucleic acid disclosed herein. In some embodiments, a provided cell may contain and / or express only a single copy of a nucleic acid disclosed herein.
[0286] In some embodiments, a cell provided herein may be designed, engineered and / or utilized for production and / or secretion of a polypeptide described herein. A variety of host-expression vector systems can be utilized to express a polypeptide described herein. Such host-expression systems represent, among other things, means by which a polypeptide of interest can be produced, after which a produced polypeptide can be subsequently purified. Host cells can include but are not limited to microorganisms such as prokaryotic bacteria (e.g., attenuated Bacillus anthracis strains, E. coli, B. subtilis) transformed with recombinant bacteriophage DNA, plasmid DNA or cosmid DNA expression vectors; yeast (e.g., Saccharomyces, Pichid) transformed with recombinant yeast expression vectors, insect cell systems infected with recombinant virus expression vectors (e.g., baculovirus); plant cell systems infected with recombinant virus expression vectors (e.g., cauliflower mosaic virus, CaMV; tobacco mosaic virus, TMV) or transformed with recombinant plasmid expression vectors (e.g., Ti plasmid); and / or mammalian cell systems (e.g., COS, CHO, BHK, NS0, 293, or 3T3 cells) harboring recombinant expression constructs containing promoters derived from the genome of mammalian cells (e.g., metallothionein promoter) or from mammalian viruses (e.g., the adenovirus late promoter; the vaccinia virus 7.5K promoter). In some embodiments, a cell may be or comprise a human cell.Page 83 of 16513092399vlAttorney Docket No. 2013763-0005
[0287] A host cell can be chosen that modulates the expression of an inserted sequence, or modifies and processes a gene product in the specific fashion desired. In some embodiments, modifications (e.g., glycosylation) and processing (e.g., cleavage) of polypeptide products can be important for the function of a polypeptide. Those of skill in the art will appreciate that different host cells have characteristic and specific mechanisms for the post-translational processing and modification of proteins and gene products. Appropriate cell lines or host systems can be chosen to ensure the correct modification and processing of the foreign protein expressed.
[0288] In some embodiments, mammalian host cells can include, but are not limited to, CHO, VERY, BHK, HeLa, COS, MDCK, 293, 3T3, W138, BT483, Hs578T, HTB2, BT20 and T47D, NSO (a murine myeloma cell line that does not endogenously produce any immunoglobulin chains), CRL7O3O and HsS78Bst cells. In mammalian host cells, a number of viral-based expression systems can be utilized. In some cases in which an adenovirus is used as an expression vector, a nucleic acid encoding a polypeptide can be ligated to an adenovirus transcription / translation control complex, e.g., the late promoter and tripartite leader sequence. This chimeric gene can then be inserted into the adenovirus genome by in vitro or in vivo recombination. Insertion in a non-essential region of the viral genome (e.g., region El or E3) can result in a recombinant virus that is viable and capable of expressing a trigger-responsive immune-modulating signaling polypeptide in infected hosts. (See, e.g., Logan & Shenk, Proc. Natl. Acad. Sci. USA, 81:355-359 (1984)).
[0289] Specific initiation signals can also be required for efficient translation of inserted trigger-responsive immune-modulating signaling polypeptide coding sequences. These signals include the ATG initiation codon and adjacent sequences. Furthermore, the initiation codon must be in phase with the reading frame of the desired coding sequence to ensure translation of the entire insert. These exogenous translational control signals and initiation codons can be of a variety of origins, both natural and synthetic. The efficiency of expression can be enhanced by the inclusion of appropriate transcription enhancer elements, transcription terminators, etc. (see Bittner, et al., Methods in Enzymol., 153:51-544 (1987)).
[0290] In some embodiments a host cell is from a cell line. In some embodiments, a cell line comprises HEK293T cells. In some embodiments, a cell line is a cancer cell line. In Page 84 of 16513092399vlAttorney Docket No. 2013763-0005some embodiments, a host cell can be a cell of the immune system (e.g., a monocyte, eosinophil, neutrophil, basophil, macrophage, dendritic cell, natural killer cell, T cell (e.g., helper T cell and cytotoxic T cell), T regulatory cell, or B cell). In some embodiments, a host cell can be a T cell, e.g., primary T cell or an immortal T cell line. An immortal T cell line can be a Jurkat cell line, for example, Neo Jurkat cells, BCL2 Jurkat cells, Jurkat E6.1 cells, J.RT3-T3.5 cells, Daudi cells, HuT78 cells, 19.2 cells, or Loucy cells. In some embodiments, a T cell can be a wild-type T cell. In some embodiments, a T cell can be an engineered T cell, e.g., a CAR-T cell. In some embodiments, a host cell comprises a cell from any one of the following cells lines HEK-293T, SH5Y5, HeLa, C2C12, BT549, HS, 578T, MCF7, MDA-MB-231, MDA-MB-468, T-47D, SF268, SF295, SF539, SNB-19, SNB-75, U251, Colo205, HCC, 2998, HCT-116, HCT-15, HT29, KM12, SW620, 786-0, A498, ACHN, CAKI, RXF, 393, SN12C, TK-10, UO-31, CCRF-CEM, HL-60, K562, MOLT-4, RPMI-8226, SR, A549, EKVX, HOP-62, HOP-92, NCI-H226, NCI-H23, NCI-H322M, NCI-H460, NCI-H522, LOX, IMVI, M14, MALME-3M, MDA-MB-435, SK-MEL-2, SK-MEL-28, SK-MEL-5, UACC-257, UACC-62, IGR0V1, OVCAR-3, OVCAR-4, OVCAR-5, OVCAR-8, SK-OV-3, NCI-ADR-RES, DU145, or PC-3 (or a combination thereof). In some embodiments, a cell line is a human neuroblastoma cell line. In some embodiments, a cell population comprises a human neuroblast cell. In some embodiments a cancer cell comprises a cell isolated from a tumor tissue. In some embodiments, a cell comprises a non-diving cell such as a myocyte, adipocyte, skin cell or neuron. In some embodiments, a non-dividing cell comprises an oocyte, (e.g., a human, mouse, or xenopus oocyte).
[0291] Methods commonly known in the art of recombinant DNA technology can be routinely applied to select desired recombinant clones, and such methods are described, for example, in Current Protocols in Molecular Biology, Ausubel, et al., eds. (John Wiley & Sons, NY 1993); Kriegler, Gene Transfer and Expression, A Laboratory Manual (Stockton Press, NY 1990); and Current Protocols in Human Genetics, Dracopoli, et al., eds. (John Wiley & Sons, NY 1994), Chapters 12 and 13; Colberre-Garapin, et al., J. Mol. Biol., 150:1 (1981). A number of selection systems can be used, including, to provide just a few examples, selection for herpes simplex virus thymidine kinase (Wigler, et al., Cell, 11:223 (1977)), hypoxanthine-guanine phosphoribosyltransferase (Szybalska & Szybalski, Proc. Natl. Acad. Sci. USA, 48:202 (1992)), and / or adenine phosphoribosyltransferase (Lowy, et Page 85 of 16513092399vlAttorney Docket No. 2013763-0005al., Cell, 22:817 (1980)) genes in tk-, hgprt-, and / or aprt- cells, respectively. In other examples, anti-metabolite resistance can be used as the basis of selection, e.g., for dhfir, which confers resistance to methotrexate (Wigler, et al., Proc. Natl. Acad. Sci. USA, 77:357 (1980); O'Hare, et al., Proc. Natl. Acad. Sci. USA, 78:1527 (1981)); gpt, which confers resistance to mycophenolic acid (Mulligan & Berg, Proc. Natl. Acad. Sci. USA, 78:2072 (1981)); neo, which confers resistance to the aminoglycoside G-418; Wu and Wu, Biotherapy, 3:87-95 (1991); Tolstoshev, Ann. Rev. Pharmacol. Toxicol., 32:573-596 (1993); Mulligan, Science, 260:926-932 (1993); and Morgan and Anderson, Ann. Rev. Biochem., 62:191-217 (1993); Can, 1993, TIB TECH 11(5): 155-215); and hygro, which confers resistance to hygromycin (Santerre et al., Gene, 30:147 (1984)).
[0292] Expression levels of can be increased by vector amplification (for a review, see Bebbington and Hentschel, “The use of vectors based on gene amplification for the expression of cloned genes in mammalian cells in DNA cloning,” Vol. 3. (Academic Press, New York (1987)). If an amplified sequence is associated with a nucleic acid sequence encoding a polypeptide, production of polypeptide can also increase (Crouse, et al., Mol. Cell. Biol., 3:257 (1983)).2. Delivery
[0293] Constructs encoding fusion proteins described herein can be administered to a cell or cell population, subject or other system using standard gene delivery protocols.Methods for gene delivery are known in the art. See, e.g., U.S. Pat. Nos. 5,399,346, 5,580,859, 5,589,466. A number of viral based systems have been developed for gene transfer into cells. These include adenoviruses, retroviruses (y-retro viruses and lentiviruses), poxviruses, adeno-associated viruses, baculoviruses, and herpes simplex viruses (see e.g., Warnock et al. (2011) Methods Mol. Biol. 737:1-25; Walther et al. (2000) Drugs 60(2):249-271; and Lundstrom (2003) Trends Biotechnol. 21 (3) : 117- 122; herein incorporated by reference).
[0294] A synthetic expression cassette of interest can also be delivered without a viral vector. For example, the synthetic expression cassette can be packaged as DNA or RNA in liposomes prior to delivery to the subject or to cells derived therefrom. Lipid encapsulation is generally accomplished using liposomes which are able to stably bind or entrap and retain Page 86 of 16513092399vlAttorney Docket No. 2013763-0005nucleic acid. The ratio of condensed DNAto lipid preparation can vary but will generally be around 1:1 (mg DNA:micromoles lipid), or more of lipid. For a review of the use of liposomes as carriers for delivery of nucleic acids, see, Hug and Sleight, Biochim. Biophys. Acta. (1991.) 1097:1-17; Straubinger et al., in Methods of Enzymology (1983), Vol. 101, pp.512-527.
[0295] Recombinant vectors carrying a synthetic expression cassette of the present invention are formulated into compositions for delivery to a host cell, subject and / or system. The compositions will comprise an “effective amount” of the nucleic acid of interest such that a sufficient amount of the modified biotin ligase can be produced for detectable proximity-dependent biotinylation of proteins in the host cell, subject and / or system to which it is administered. An appropriate effective amount can be readily determined by one of skill in the art.
[0296] The compositions may include one or more “pharmaceutically acceptable excipients or vehicles” such as water, saline, glycerol, polyethyleneglycol (PEG), hyaluronic acid, ethanol, etc. Additionally, auxiliary substances, such as wetting or emulsifying agents, pH buffering substances, surfactants and the like, may be present in such vehicles. Certain facilitators of nucleic acid uptake and / or expression can also be included in the compositions or co-administered.
[0297] Once formulated, compositions can be administered directly to a cell, subject and / or system, e.g., delivered ex vivo, to cells derived from the subject, using methods such as those described above, or in vitro. For example, methods for delivery are known in the art and can include, e.g., dextran-mediated transfection, calcium phosphate precipitation, polybrene mediated transfection, lipofectamine and LT-1 mediated transfection, protoplast fusion, electroporation, encapsulation of the polynucleotide (s) in liposomes, and direct microinjection of the DNA into nuclei.
[0298] Direct delivery of synthetic expression cassette compositions in vivo will generally be accomplished with or without viral vectors, as described above, by injection using either a conventional syringe, needless devices such as Bioject or a gene gun, such as the Accell gene delivery system (PowderMed Ltd, Oxford, England).Page 87 of 16513092399vlAttorney Docket No. 2013763-0005
[0299] In some embodiments, a composition includes and is formulated for delivery to a cell, e.g., where the cell is a cell is from a cell line. In some embodiments, a cell line comprises HEK293T cells. In some embodiments, a cell line is a cancer cell line. In some embodiments, a composition is formulated for delivery to a cell of the immune system (e.g., a monocyte, eosinophil, neutrophil, basophil, macrophage, dendritic cell, natural killer cell, T cell (e.g., helper T cell and cytotoxic T cell), T regulatory cell, or B cell). In some embodiments, a composition is formulated for delivery to a cell from a cell line comprising HEK-293T, SH5Y5, HeLa, C2C12, BT549, HS, 578T, MCF7, MDA-MB-231, MDA-MB-468, T-47D, SF268, SF295, SF539, SNB-19, SNB-75, U251, Colo205, HCC, 2998, HCT-116, HCT-15, HT29, KM12, SW620, 786-0, A498, ACHN, CAKI, RXF, 393, SN12C, TK-10, UO-31, CCRF-CEM, HL-60, K562, MOLT-4, RPMI-8226, SR, A549, EKVX, HOP-62, HOP-92, NCI-H226, NCI-H23, NCI-H322M, NCI-H460, NCI-H522, LOX, IMVI, M14, MALME-3M, MDA-MB-435, SK-MEL-2, SK-MEL-28, SK-MEL-5, UACC-257, UACC-62, IGR0V1, OVCAR-3, OVCAR-4, OVCAR-5, OVCAR-8, SK-OV-3, NCI-ADR-RES, DU 145, or PC-3 (or a combination thereof). In some embodiments, a cell line is a human neuroblastoma cell line. In some embodiments, a composition is formulated for delivery to a human neuroblast cell. In some embodiments, a composition is formulated for delivery to a cancer cell. In some embodiments, a cancer cell comprises a cell isolated from a tumor tissue. In some embodiments, a composition is formulated for delivery to a non-diving cell such as a myocyte, adipocyte, skin cell or neuron. In some embodiments, a non-dividing cell comprises an oocyte, (e.g., a human, mouse, or xenopus oocyte).
[0300] In some embodiments, a composition includes and is formulated for delivery of a viral particle.D. Uses
[0301] A fusion protein, or one or more fusion proteins as part of proximity labelling systems as described herein, may be formulated for delivery to a cell, subject and / or system, and once delivered, the fusion protein may utilize its proximity labeling enzyme to label proteins that interact with the protein of interest when they are in proximity therewith. Such methods allow for identification of interacting components that interact with a protein of interest (POI). In addition, proximity labelling systems, when expressed in cells, allow for proximity labelling of POIs or genes of interest (GDIs) (e.g., a target genomic locus) and Page 88 of 16513092399vlAttorney Docket No. 2013763-0005their interacting components, and in turn isolation or pull down of the particular target genomic locus or POI and its interacting components. In some embodiments, a POI is a core transcriptional component, and interacting proteins identified are active transcriptional components within a cell. In some embodiments, a particular genomic locus is targeted through a DNA-targeting moiety, and a DNA-targeting moiety may specifically bind to a region within or comprising the target DNA locus associated with transcriptional regulation.
[0302] Proximity labeling enzymes used in fusion proteins described herein can be used generally for biotinylation of proteins. Contacting a proximity labeling enzyme with its substrates (e.g., biotin or a biotin derivative such as desthiobiotin) and ATP, results in biotinylation of proteins in proximity to the fusion protein. For example, a fusion protein can be used for in vitro biotinylation of proteins (e.g., individual purified proteins in a test tube or unpurified proteins such as in a cell lysate). In addition, fusion proteins can be used for proximity labeling of proteins in a cell or live organism. For example, a fusion protein can be introduced into a cell or live organism, and contacted with a substrate (e.g., biotin or a biotin derivative such as desthiobiotin) and ATP, wherein proteins in proximity to the fusion protein are biotinylated.
[0303] Biotinylated proteins can be isolated with a biotin-binding protein, such as streptavidin or avidin. In some embodiments, a biotin-binding protein is an antibody. The biotin-binding protein may be immobilized on a solid support, such as, but not limited to, a magnetic bead, non-magnetic bead, microtiter plate well, glass plate, nylon, agarose, or acrylamide to facilitate removal of biotinylated proteins from a liquid. Isolated biotinylated proteins can then be analyzed by any appropriate method for protein identification such as, but not limited to, mass spectrometry, liquid chromatography-mass spectrometry (LC / MS), immunoassay (e.g., enzyme-linked immunosorbent assay (ELISA), immunoprecipitation), Western blot, immunoelectrophoresis, immunostaining, high-performance liquid chromatography (HPLC), protein sequencing, and peptide mass fingerprinting (see e.g., the schematic in FIG. 1).
[0304] Methods described herein may be applied to cell samples comprising a single cell or a population of cells of interest and can be performed on any type of cell, including any cell from a prokaryotic, eukaryotic, or archaeon organism, including bacteria, archaea,Page 89 of 16513092399vlAttorney Docket No. 2013763-0005fungi, protists, plants, and animals. Cells from tissues, organs, and biopsies, as well as recombinant cells, cells from cell lines cultured in vitro, and artificial cells (e.g., nanoparticles, liposomes, polymersomes, or microcapsules encapsulating nucleic acids) may all be used in the methods described herein. Methods described herein are also applicable for investigating protein and / or nucleic acid localization in cellular fragments, cell components, or organelles comprising nucleic acids.
[0305] In some embodiments, proximity-dependent biotinylation is performed on an intact, naturally occurring and / or modified cell. A cell may be isolated from other cells, mixed with other cells in a culture, or within a tissue (partial or intact), or a whole live organism. Methods for proximity labeling and the related reagents, materials and compositions described herein are well suited for use in live cells or whole live organisms, and fixed cells and tissues, for example, fixed cells and tissues obtained from a subject, e.g., in a clinical setting as well as lysed cells.
[0306] In some embodiments, methods and strategies for proximity-dependent labelling involve the use of a biotin ligase labelling enzyme (e.g., Turbo-ID). In such methods, the biotin ligase catalyzes a reaction with biotin and ATP that generates a reactive unstable biotinoyl-5'-AMP reaction intermediate that is capable of covalently labeling nearby proteins. The half-life of the reaction intermediate generated by the biotin ligase determines how far the reagent can travel from its point of generation before reacting with a molecule. Accordingly, the half-life of biotinoyl-5'-AMP determines its labeling radius. Because the enzyme generated reaction intermediate has a short half-life in cells, only proteins in proximity to the biotin ligase and the reaction intermediate generated by the biotin ligase (typically a few tens to hundreds of nanometers) are sufficiently close to be covalently modified (i.e., biotinylated).
[0307] A biotin ligase can be introduced into a cell (e.g., within a fusion protein) and contacted with the biotin and ATP substrates under conditions suitable for the biotin ligase to produce the reactive biotinoyl-5'-AMP intermediate, which biotinylates proteins in the vicinity of the enzyme. A biotin ligase may be delivered to the cell interior or exterior, depending on which region of the cell is being analyzed. In some embodiments, a biotin ligase is delivered to the interior of the cell, and in some instances, to specific subcellular Page 90 of 16513092399vlAttorney Docket No. 2013763-0005compartments. In some embodiments, a biotin ligase is delivered to a tissue. A biotin ligase may also be introduced into a cell by transfecting the cell with a recombinant polynucleotide comprising a promoter operably linked to a polynucleotide encoding the biotin ligase within a fusion protein. A recombinant polynucleotide may comprise an expression vector, for example, a bacterial plasmid vector or a viral expression vector, such as, but not limited to, an adenovirus, retrovirus (e.g., y-retrovirus and lentivirus), poxvirus, adeno-associated virus, baculovirus, or herpes simplex virus vector.
[0308] In some embodiments, a proximity labeling enzyme (e.g., a biotin ligase or peroxidase) within a fusion protein described herein, is introduced into a whole live organism. For example, a fusion protein can be introduced into bacteria, archaea, fungi, protists, plants, and animals (both vertebrates and invertebrates), including, without limitation, plants such as flowering plants, conifers and other gymnosperms, fems, clubmosses, homworts, liverworts, mosses, and green algae; fungi such as molds and yeasts; protists such as amoebae, flagellates, and ciliates; worms; insects such as beetles, ants, bees, moths, butterflies, and flies; amphibians such as frogs and salamanders (e.g., axolotls); fish; reptiles; mammals, including human and non-human mammals such as non-human primates, including chimpanzees and other apes and monkey species; laboratory animals such as mice, rats, rabbits, hamsters, guinea pigs, and chinchillas; domestic animals such as dogs and cats; farm animals such as sheep, goats, pigs, horses and cows; and birds such as domestic, wild and game birds, including chickens, turkeys and other gallinaceous birds, ducks, geese; and transgenic animals.
[0309] In some embodiments, a fusion protein is introduced into a model organism, such as an animal model or test subject for use in scientific or biomedical research or drug screening. Model organisms include, but are not limited to, prokaryotic model organisms such as bacteria (e.g., Escherichia coli) and eukaryotic model organisms such as yeasts (e.g., Saccharomyces cerevisiae and Schizosaccharomyces pomhe). plants, including flowering plants (e.g., Arahidopsis thaliana), mosses (e.g., Physcomitrella patens)', and unicellular green alga (e.g., Chlamydomonas reinhardtii),' protists (e.g., Tetrahymena thermophila),' invertebrates such as worms (e.g., Caenorhahditis elegans) and flies (e.g., Drosophila melanogaster),' amphibians such as frogs (e.g., Xenopus tropicalis, Xenopus laevis) and salamanders (e.g., axolotls); fish (e.g., Danio rerio, Fundulus heteroclitus,Page 91 of 16513092399vlAttorney Docket No. 2013763-0005Nothobranchius furzeri), mammals such as rodents, including guinea pigs (e.g., Cavia porcellns). mice (e.g., Mus musculus), and rats (e.g., Rattus norvegicus). and non-human primates such as the rhesus macaque and chimpanzee. Model organisms can be used, for example, to study disease pathology, development, toxicology, aging, gene function, signaling pathways, intracellular processes, and physiological systems, and in production and screening of therapeutics and vaccines.
[0310] In some embodiments, a proximity labeling enzyme is modified to improve its capability in proximity labeling at a particular subcellular location. For example, a proximity labeling enzyme can be engineered to be expressed and / or active only within a subcellular compartment or structure of interest. A proximity labeling enzyme may also be engineered to comprise one or more mutations that enhance its catalytic activity in a subcellular compartment or structure of interest. Examples of such modifications is provided in US 2021 / 0214708 Al, which is herein incorporated by reference in its entirety.
[0311] Fusion proteins, as described herein, may direct a proximity labeling enzyme to a POI in a number of ways. For example, a fusion protein may include a targeting sequence that directs the proximity labeling enzyme to the subcellular region of interest. Targeting sequences that can be used include, but are not limited to, a secretory protein signal sequence, a membrane protein signal sequence, a nuclear localization sequence, a mitochondrial localization sequence, an outer mitochondrial membrane sequence, an endoplasmic reticulum localization sequence, an endoplasmic reticulum membrane targeting sequence, a nucleolar localization signal sequence, a nuclear export signal sequence, a peroxisome localization sequence, and a protein binding motif sequence.
[0312] In some embodiments, a fusion protein contains a proximity labeling enzyme that is covalently linked to a peptide or protein that directs the fusion protein to a subcellular region of interest, such as a cytosolic protein, a nuclear protein, a membrane protein, a mitochondrial protein, a P-body protein, or a secretory pathway protein, any other protein of interest described herein. Attachment to the protein of interest results in proximity labeling of proteins surrounding the protein of interest in the locations where it resides in the cell. In some embodiments, a fusion protein comprises a proximity labeling enzyme covalently linked to a binding agent that specifically binds a particular epitope found on certain proteins Page 92 of 16513092399vlAttorney Docket No. 2013763-0005in a subcellular region of interest, which similarly allows proximity labeling of surrounding nearby proteins.
[0313] In addition, proximity-dependent biotinylation of proteins can be combined with crosslinking of nucleic acids to the labeled proteins to identify nucleic acids within or near a particular subcellular compartment in vivo and for mapping protein-nucleic acid interactions within a cell. Crosslinking of nucleic acids to the biotinylated cellular proteins allows identification of nucleic acids (e.g., RNA or DNA) in the vicinity of the biotinylated proteins. Furthermore, such crosslinking allows nucleic acids to be mapped to particular organelles, including subcompartments of organelles without subcellular fractionation.
[0314] Crosslinking agents that can be used for crosslinking proteins and nucleic acids include, but are not limited to, dimethyl suberimidate, N-hydroxysuccinimide, formaldehyde, and glutaraldehyde. In addition, carboxyl -reactive chemical groups such as diazomethane, diazoacetyl, and carbodiimide can be included for crosslinking carboxylic acids to primary amines. In particular, the carbodiimide compounds, l-ethyl-3-(-3-dimethylaminopropyl) carbodiimide hydrochloride (EDC) and N',N'-dicyclohexyl carbodiimide (DCC) can be used for conjugation with carboxylic acids. In order to improve the efficiency of crosslinking reactions, N-hydroxysuccinimide (NHS) or a water-soluble analog (e.g., Sulfo-NHS) may be used in combination with a carbodiimide compound. The carbodiimide compound (e.g., EDC or DCC) couples NHS to carboxyl groups to form an NHS ester intermediate, which readily reacts with primary amines at physiological pH. In addition, ultraviolet light can be used for crosslinking proteins to nucleic acids. For a description of various crosslinking agents and techniques, see, e.g., Wong and Jameson Chemistry of Protein and Nucleic Acid Cross-Linking and Conjugation (CRC Press, 2nd edition, 2011), Hermanson Bioconjugate Techniques (Academic Press, 3rd edition, 2013), herein incorporated by reference in their entireties.
[0315] In certain embodiments, crosslinking of proteins and nucleic acids is performed using click chemistry. Crosslinking of proteins and nucleic acids using click chemistry can be performed with suitable crosslinking agents comprising reactive azide or alkyne functional groups. See, e.g., Kolb et al., 2004, Angew Chem Int Ed 40:3004-31; Evans, 2007, Aust J Chem 60:384-95; Millward et al. (2013) Integr Biol (Camb) 5(1): 87-95),Page 93 of 16513092399vlAttorney Docket No. 2013763-0005Lallana et al. (2012) Pharm Res 29(1): 1-34, Gregoritza et al. (2015) Eur J Pharm Biopharm.97(Pt B):438-453, Musumeci et al. (2015) Curr Med Chem. 22(17):2022-2050, McKay et al. (2014) Chem Biol21 (9): 1075-1101, Ulrich et al. (2014) Chemistry 20(l):34-41, Pasini (2013) Molecules 18(8):9512-9530, and Wangler et al. (2010) Curr Med Chem. 17(11): 1092-1116; herein incorporated by reference in their entireties.
[0316] In particular, crosslinking can be performed using strain-promoted azidealkyne cycloaddition (SPAAC) click chemistry, a Cu-free variation of click chemistry that is generally biocompatible with cells. SPAAC utilizes a substituted cyclooctyne having an internal alkyne in a strained ring system. Ring strain together with electron- withdrawing substituents in the cyclooctyne promote a [3+2] dipolar cycloaddition with an azide functional group. SPAAC can be used for bioconjugation and crosslinking by attaching azide and cyclooctyne moieties to molecules. For a description of SPAAC, see, e.g., Baskin et al. (2007) Proc Natl Acad Sci USA 104(43): 16793-16797, Agard et al. (2006) ACS Chem. Biol.1: 644-648, Codelli et al. (2008) J. Am. Chem. Soc. 130:11486-11493, Gordon et al. (2012) J. Am. Chem. Soc. 134:9199-9208, Jiang et al. (2015) Soft Matter 11(30): 6029-6036, Jang et al. (2012) Bioconjug Chem. 23(11):2256-2261, Ornelas et al. (2010) J Am Chem Soc.132(11):3923-3931; herein incorporated by reference in their entireties.
[0317] Biotinylated interacting proteins or nucleic acids can be isolated with a biotinbinding protein, such as streptavidin or avidin. In some embodiments, biotinylated interacting proteins or nucleic acids can be isolated with anti-biotin antibody. The biotinbinding protein may be immobilized on a solid support (e.g., streptavidin beads or magnetic beads) as described above to facilitate removal from a liquid. The isolated proteins can then be analyzed to identify nucleic acids and / or proteins by any appropriate method (e.g., mass spectrometry or immunoassays for identification of proteins and sequencing or polymerase chain reaction (PCR) with suitable primers for identification of nucleic acids). RNA may be reverse transcribed into cDNA with a reverse transcriptase prior to performing PCR (i.e., RT-PCR) and / or sequencing.
[0318] Any high-throughput technique for sequencing the nucleic acids can be used in the practice of the invention. Deep sequencing of nucleic acids can be used, for example, to improve sequence accuracy and for determining the frequency of RNA molecules in Page 94 of 16513092399vlAttorney Docket No. 2013763-0005particular subcellular compartments or regions. DNA sequencing techniques include dideoxy sequencing reactions (Sanger method) using labeled terminators or primers and gel separation in slab or capillary, sequencing by synthesis using reversibly terminated labeled nucleotides, pyrosequencing, 454 sequencing, sequencing by synthesis using allele specific hybridization to a library of labeled clones followed by ligation, real time monitoring of the incorporation of labeled nucleotides during a polymerization step, polony sequencing, SOLID sequencing, and the like.
[0319] As discussed above, fusion proteins described herein comprise a proximity labeling enzyme that can be genetically targeted to a cellular region of interest to identify proteins and / or nucleic acids in the vicinity of tagged proteins within a specific subcellular compartment or region (e.g., the nucleus, endoplasmic reticulum, Golgi, mitochondria, mitochondria outer membrane, mitochondria inner membrane, mitochondria matrix space, chloroplasts, synaptic cleft, presynaptic membrane, postsynaptic membrane, dendritic spines, transport vesicles, regions of contact between mitochondria and endoplasmic reticulum, nuclear membrane, etc.) can be specifically tagged. In some embodiments, proteins within particular cell types (e.g., astrocytes, dendrocytes, stem cells, etc.) can be specifically tagged, for example, proteins within a specific cell type within a complex tissue, animal, or cell population. In some embodiments, proteins within particular macromolecular complexes (e.g., protein complexes such as ribosomes, replisome, transcription complex, spliceosome, DNA repair complex, fatty acid synthase, polyketide synthase, non-ribosomal peptide synthase, glutamate receptor signaling complex, neurexin-neuroligin signaling complex, etc.) can be tagged. In each context, the tagged proteins or protein-nucleic acid fusions can be analyzed (e.g., isolated and identified) to map protein and / or nucleic acid localization in specific cells, cellular compartments or regions, or macromolecular complexes of interest. This information can be used for research, diagnostic, therapeutic, and other applications.
[0320] For example, cells may be isolated from a patient, amplified or differentiated using IPS cell technology (induced pluripotent stem cell), contacted with a vector (e.g., a viral vector) that expresses a biotin ligase, for example, a fusion protein fused to a localization signal effecting localization of the biotin ligase in a specific subcellular compartment. Labeling and / or crosslinking can be performed in the living cells, as described herein, and the resulting tagged proteins or protein-nucleic acid fusions can be analyzed, for Page 95 of 16513092399vlAttorney Docket No. 2013763-0005example, to identify patient specific information that can be useful to assist in diagnostic, prognostic, and / or therapeutic decisions, and in drug screening assays.
[0321] In some embodiments, the reactive intermediate, once created, biotinylates (i.e., labels) proteins that are within the vicinity of the proximity labeling enzyme. The term “within the vicinity” refers to the spatial location around the enzyme and / or substrate that is labeled or within the “labeling radii”. Proteins that are further from a proximity labeled enzyme are generally labeled to a lesser extent than proteins that are closer to a proximity labeling enzyme. Proteins that are not within the vicinity of a proximity labeling enzyme (i.e., within the “labeling radii”) are not exposed to the reactive intermediate and hence not labeled. Some proteins in the vicinity of a proximity labeling enzyme may fail to get labeled, e.g. if they are sterically buried or do not have any exposed residues capable of being biotinylated.
[0322] In some embodiments, in vivo protein tagging is performed with a fusion protein that can be genetically targeted to any part of a live cell. In some embodiments, a fusion protein is present and / or active in all regions of the cell. In some embodiments, a fusion protein is present and / or active only in a subcellular compartment of the cell. In some embodiments, biotin substrate can be added or uncaged for the desired window of time, to permit precise temporal control of labeling. In some embodiments, it is preferable for the reactive species not to cross cell membranes, to allow mapping of membrane -bounded structures.
[0323] In some embodiments, a fusion protein is engineered to be expressed and / or targeted in vivo or in situ to specific cells, cellular compartments (e.g., endoplasmic reticulum, Golgi apparatus, mitochondria, nucleus, the synaptic cleft, transport vesicles, etc.), and / or macromolecular complexes (e.g., protein complexes such as ribosomes, nuclear pore complex, fatty acid synthases) of interest. In some embodiments, a fusion protein is engineered to tag proteins that are located within a limited distance of the biotin ligase. As a result, in some embodiments, proteins that are located within the targeted cell, cellular compartment, and / or macromolecular complex (e.g., protein complex) are specifically tagged relative to other proteins that are not located near the biotin ligase. It should be appreciated that the tagging process itself does not need to be protein specific. For example, in some Page 96 of 16513092399vlAttorney Docket No. 2013763-0005embodiments, it is the specific localization of a fusion protein that results in the specific tagging of a subset of proteins of interest. In some embodiments, proteins that are present within the vicinity of a fusion protein may be tagged for further analysis. In some embodiments, all proteins present within the vicinity of the biotin ligase may be tagged. Various versions of the methodology offer a range of labeling radii, from about 500 nm to less than 10 nm, e.g., tagging radii of about 500 nm, about 400 nm, about 300 nm, about 250 nm, about 200 nm, about 100 nm, about 90 nm, about 80 nm, about 70 nm, about 60 nm, about 50 nm, about 40 nm, about 30 nm, about 20 nm, about 10 nm, about 5 nm, about 2.5 nm, or about 1 nm. In some embodiments, a labeling radii of a fusion protein described herein is between 10 and lOOnm.
[0324] In some embodiments, the reactive intermediate produced by the biotin ligase is inactivated by contacting it with a quenching agent (e.g., water for an unstable reaction intermediate such as produced by biotin ligase). As a result, the reactive moiety can have a short half-life and only modify proteins that are located within a short distance of the site of production (the biotin ligase) before being inactivated. Accordingly, the zone of tagging can be limited by the diffusion rate and half-life of the reactive reaction intermediate.
[0325] The methods provided herein can also be used to map protein and / or nucleic acid localization in specific cell types within complex tissues or heterogeneous cell populations, or of specific subcellular structures or organelles within specific cells in complex tissues or populations. The methods are particularly useful for mapping subcellular localization of proteins and / or nucleic acids in rare cells within complex cell populations.
[0326] Maps of subcellular localization of proteins and / or nucleic acids can be developed not only for different cells, subcellular compartments, tissues, or organisms but also for cells, tissues, or organisms exposed to different conditions or environments. For example, cells or organisms exposed to different therapeutic agents, different concentrations of therapeutic agents, and / or combinations of therapeutic agents may be mapped and analyzed independently or compared against one another to examine changes occurring within a cell, tissue, or organism. Additionally, changes in protein and / or nucleic acid localization in cells, tissues, or organisms over time associated with diseased states can bePage 97 of 16513092399vlAttorney Docket No. 2013763-0005monitored by comparison of mapped nucleic acid localization in cells, tissues, or organisms in diseased and normal (i.e., healthy control, not having the disease) states.
[0327] Maps of subcellular localization of proteins and / or nucleic acids can also be developed for cells, subcellular compartments, tissues, or organisms at different developmental stages. For example, a map of the subcellular localization of proteins and / or nucleic acids can be compared to reference maps for cells, subcellular compartments, tissues, or organisms at the same or different developmental stages.
[0328] In some embodiments, methods described herein may be used to characterize a protein that interacts with a target protein of interest. In some embodiments, characterizing a protein that interacts with a target protein of interest comprises applying a fusion protein as described herein to a system that comprises a target protein of interest under conditions under which the protein of interest may interact with the interacting proteins, and in which the proximity labeling enzyme of the fusion protein may label the interacting component in proximity with the target protein of interest. Labeled interacting components may be isolated from the system.
[0329] Methods provided herein allow for proximity labeling of interacting proteins that interact with a protein of interest in various systems. In some embodiments a fusion protein is applied to two systems. In some embodiments, a first system comprises one or more cancer cells. In some embodiments, a second system comprises one or more comparable non-cancer cells. In some embodiments labeled interacting proteins from a first system may be compared to labeled interacting components in a second system.
[0330] In some embodiments, a protein of interest is a core transcriptional component and a fusion protein comprises a proximity labeling enzyme linked to a core transcriptional component (e.g., TBP). In some embodiments, a fusion protein is applied to a first system and a second system, where the first system comprises one or more cancer cells and the second system comprises one or more comparable non-cancer cells. In some embodiments, upon contacting each system with the fusion protein, interacting components in each system are labeled by the proximity labeling enzyme. The interacting components may be isolated from the system using methods described herein and compared. Interacting componentsPage 98 of 16513092399vlAttorney Docket No. 2013763-0005present in a first system but not present in the second system may be characterized as core transcriptional machinery associated with cancer.
[0331] In some embodiments a cell population within a system described herein comprises one or more cancer cells, e.g., isolated from a cancerous tissue or tumor or from a cancer cell. In some embodiments, a cancer can be carcinoma, sarcoma, melanoma, lymphoma, leukemia, or blastoma. In some embodiments, a carcinoma can be a basal cell carcinoma, squamous cell carcinoma, renal cell carcinoma, ductal carcinoma in situ, invasive ductal carcinoma, and / or adenocarcinoma. In some embodiments, a carcinoma can be a prostate cancer, an ovarian cancer, a uterine cancer, a cervical cancer, a colorectal cancer, a breast cancer, a bladder cancer, a pancreatic cancer, an esophageal cancer, a gastrointestinal cancer, a hepatocellular cancer, a thyroid cancer, or a lung cancer. In some embodiments, a sarcoma can be a angiosarcoma, a chondrosarcoma, an Ewing’s sarcoma, fibrosarcoma, a gastrointestinal stromal cancer, a Leiomyosarcoma, a liposarcoma, an osteosarcoma, a pleomorphic sarcoma, a rhabdomyosarcoma, or a synovial sarcoma. In some embodiments, a melanoma can be a superficial spreading melanoma, a nodular melanoma, a lentigo maligna melanoma, or an acral melanoma. In some embodiments, a lymphoma can be a B-cell lymphoma, a T cell lymphoma, or an NK-cell lymphoma. In some embodiments, a leukemia can be an acute myeloid leukemia, a chronic myeloid leukemia, acute lymphocytic leukemia, or a chronic lymphocytic leukemia. In some embodiments, a blastoma can be a hepatoblastoma, a medulloblastoma, a nephroblastoma, a neuroblastoma, a pancreatoblastoma, a pleuropulmonary blastoma, retinoblastoma, or a glioblastoma. In some embodiments, a cancer can be a Stage I, Stage II, Stage III, or Stage IV cancer. In some embodiments, a cancer can be metastatic. In some embodiments, a tumor is a neuroendocrine tumor (e.g., a tumor resulting from adrenal cancer, a carcinoid tumor, a merkel cell carcinoma, a pancreatic neuroendocrine tumor, a paraganglioma, or a pheochromocytoma).
[0332] In some embodiments, a cancer cell line is a human brain, breast, colon, head and neck, kidney, leukemia, liver, lung, metastatic lines, lymphoma, and / or prostate cancer cell line. In some embodiments, a cell population includes or comprise any one of the following cells lines HEK-293T, SH5Y5, HeLa, C2C12, BT549, HS, 578T, MCF7, MDA-MB-231, MDA-MB-468, T-47D, SF268, SF295, SF539, SNB-19, SNB-75, U251, Colo205, HCC, 2998, HCT-116, HCT-15, HT29, KM12, SW620, 786-0, A498, ACHN, CAKI, RXF,Page 99 of 16513092399vlAttorney Docket No. 2013763-0005393, SN12C, TK-10, UO-31, CCRF-CEM, HL-60, K562, MOLT-4, RPMI-8226, SR, A549, EKVX, HOP-62, HOP-92, NCI-H226, NCI-H23, NCI-H322M, NCI-H460, NCI-H522, LOX, IMVI, M14, MALME-3M, MDA-MB-435, SK-MEL-2, SK-MEL-28, SK-MEL-5, UACC-257, UACC-62, IGROV1, OVCAR-3, OVCAR-4, OVCAR-5, OVCAR-8, SK-OV-3, NCI-ADR-RES, DU145, or PC-3 (or a combination thereof). In some embodiments, a cell line is a human neuroblastoma cell line. In some embodiments, a cell population comprises human neuroblast cells. In some embodiments a cancer cell comprises a cell isolated from a tumor tissue.
[0333] Such methods of comparing core transcriptional components within various populations of cells associated with cancer allows for insights into potential transcription factors associated cancer metastasis that could be new targets for cancer therapies.
[0334] In some embodiments a cell population within a system described herein comprises a non-diving cell population. Non-diving tissues and organs of mammals consist of various cell types, including those that are dividing and those that are non-dividing. Cell types such as myocytes, adipocytes, skin cells and neurons, are typically in a non-dividing state, i.e., terminally-differentiated. Terminal differentiation is the process by which cells develop and mature, and take on their specialized function a specific structural, functional, and biochemical properties and roles. In some embodiments, a non-dividing cell or cell population comprises one or more oocytes, (e.g., human, mouse, or xenopus oocytes). In some embodiments, a non-dividing cell or cell population comprises one or more cardiomyocytes. In some embodiments, a non-dividing cell or cell population comprises one or more neurons.
[0335] In some embodiments, a fusion protein is applied to a first system and a second system, where the first system comprises one or more cancer cells and the second system comprises one or more non-dividing cells or terminally differentiated cells. In some embodiments, upon contacting each system with the fusion protein, interacting components in each system are labeled by the proximity labeling enzyme. The interacting components may be isolated from the system using methods described herein and compared. Interacting components present in a first system but not present in the second system may be characterized as core transcriptional machinery associated with cancer. Interacting Page 100 of 16513092399vlAttorney Docket No. 2013763-0005components present in the second system but not present in the first system may be characterized as core transcriptional machinery associated with non-dividing cells.
[0336] In some embodiments, a fusion protein is applied to a first system and a second system, where the first system comprises one or more non-dividing cells or terminally differentiated cells and the second system comprises one or more comparable non-dividing cells or terminally differentiated cells. In some embodiments, upon contacting each system with the fusion protein, interacting components in each system are labeled by the proximity labeling enzyme. The interacting components may be isolated from the system using methods described herein and compared. Interacting components present in the first system but not present in the second system may be characterized as core transcriptional machinery components associated with initiation or progression of terminal differentiation in cells.
[0337] Such methods of comparing core transcriptional components within various populations of cells associated with cancer cells or dividing cells and terminally differentiated cells or non-dividing cells provides insights into potential transcription factors, proteins, etc., that could be involved in terminal differentiation. In some embodiments, these identified factors and proteins may be utilized and upregulated in cancer cells that are rapidly dividing in order to push the cells towards a differentiated state.
[0338] In some embodiments, a comparable non-cancer cell is a HEK293 cell. In some embodiments, a comparable non-cancer cell comprises a terminally differentiated cell. In some embodiments, a comparable non-cancer cell comprises a non-dividing cell (e.g., an oocyte).
[0339] In some embodiments, a system, after contacting with a fusion protein, is further exposed hydrogen peroxide. In some embodiments, biotin is added to a system, after contacting with a fusion protein, such that labeled interacting proteins are biotinylated proteins. In some embodiments, separating labeled interacting proteins from the system comprises contacting the extracted protein with an anti -biotin antibody or streptavidin beads. In some embodiments, labeled interacting proteins are characterized using LC MS / MS analysis. Proteins that have been characterized using LC MS can be further characterized by based on cell atlas location. Cell atlas location refers to location of an identified protein within a cell. In some embodiments, this can be determined by identifying signaling,Page 101 of 16513092399vlAttorney Docket No. 2013763-0005regulatory, or other characteristic sequences included in the identified protein sequence, such as a nuclear localization sequence (NLS), which indicates an identified protein is localized in the nucleus of the cell. In some embodiments, an identified protein is characterized as having a cell atlas location within the mitochondria of the cell when the identified protein contains a mitochondrial transport signal sequence. One of skill in the art will understand that other characteristic sequences may be used to characterize an identified protein cell atlas location.
[0340] In some embodiments, a core transcriptional components identified as being specific to non-dividing or terminally differentiated cells may be upregulated in a cancer cell or cell population, as described herein. In some embodiments, a core transcriptional component identified as being specific to non-dividing or terminally differentiated cells may be upregulated in an actively dividing cell population, as described herein. In some embodiments, expression of such a transcriptional component in a cell promotes, increases, or induces differentiation of a cell and / or reduces cellular replication.
[0341] In some embodiments, methods described herein include administering to a subject suffering from a cancer or tumor a therapy that targets a cancer-associated interacting component (e.g., a core transcriptional component in cancer cells), identified according to methods described herein.
[0342] Methods described herein for identifying core transcriptional components associated with specific cell populations (e.g., cancers cells) may also be applied to other cellular pathways and processes. For example, in some embodiments, methods are applied to identify core translational components using fusion proteins described herein. In such embodiments, a protein of interest (or binding agent specific for the protein of interest) is a core translational component. Other cellular processes and pathways to which methods described herein may apply include processes such as of transcription, translation, replication (e.g., involved in DNA damage and / or repair), tumor metastasis and protein degradation. In some embodiments, interacting proteins are identified and / or characterized based on interactions with fusion proteins comprising a protein of interest known to be active in one or more of transcription, translation, replication (e.g., involved in DNA damage and / or repair), tumor metastasis and protein degradation.Page 102 of 16513092399vlAttorney Docket No. 2013763-0005
[0343] All publications, patent applications, patents, and other references mentioned herein are incorporated by reference in their entirety. In addition, the materials, methods, and examples are illustrative only and not intended to be limiting. Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention belongs. Although methods and materials similar or equivalent to those described herein can be used in the practice or testing of the present invention, suitable methods and materials are described herein.
[0344] The disclosure is further illustrated by the following examples. The examples are provided for illustrative purposes only. They are not to be construed as limiting the scope or content of the disclosure in any way.EXEMPLIFICATIONExample 1: Fusion Protein for Identification of Transcriptional Components
[0345] The present Example demonstrates methods and compositions used to identify protein complexes interacting with transcriptional machinery in cancer cells. In this Example, a fusion protein was generated that comprises a proximity labeling enzyme fused with a transcriptional machinery component. In this Example, TATA-binding protein (TBP) was used as the transcriptional machinery component. The fusion protein described in this Example was contacted with various cell populations, including both normal and cancer cells. Comparing the labelled components identified from each cell type allows for the identification core transcriptional components specific to a target cell population (e.g., cancer cells).
[0346] As such, the method described in this Example provides a new way of understanding transcriptional patterns and compositions of transcriptional apparatus in cancer and other target cell populations. This can be used to understand metastasis and acquisition of drug resistance in cancer cells and can increase rate of drug discovery.Exemplary Fusion Protein DesignPage 103 of 16513092399vlAttorney Docket No. 2013763-0005
[0347] Proximity labelling enzymes include Biotin ligases and Peroxidases. Biotin ligases include a range of Bio-ID enzymes and their engineered counterparts such as Turbo-ID and mini -Turbo ID. Turbo-ID is used to achieve the long-term interaction of the protein of interest, which can range from hours to days, and it uses endogenous or natural biotin molecules for the labelling. Exemplary Peroxidases include the enzymes APEX and an engineered version called APEX2, which has high specificity and temporal resolution.APEX2 is able to map the interactions of a Protein of Interest (POI) within a timeframe of about a minute, and it uses a toxic oxidizing agent such as hydrogen peroxide and a Biotinphenol.
[0348] The function and interaction of a PL enzyme with a POI, e.g., a transcriptional component, varies between particular enzymes. Proximity labelling enzymes such as Turbo-ID and APEX2 are examples of enzymes used in the methods described herein.
[0349] In this Example, a labelling radius of the Proximity labelling (PL) enzyme was determined by a flexible linker with a length ranging from lOnm-100 nm. Labelling radius and linker length was designed specifically for its use with TBP. Linkers containing short sequences (e.g., short than a length of 5-10 amino acids) may prohibit or alter the function of the protein, such as its interaction with the other transcription factors. In this example, various DNA binding assays were used to optimize the linker sequence used with TBP. The appropriate linker was determined based on whether the fusion protein resulted in activity / response comparable to the un-fused protein.
[0350] The PL enzyme was coupled with a transcriptional machinery component, which enables the mapping of the protein interaction with active transcriptional components. Methods of proximity labeling described in this Example for biotinylated protein pull down have been improved by using an anti -biotin antibody instead of streptavidin magnetic beads.Production of an Exemplary Pusion Protein:1. A cDNA sequence of protein of interest (TBP) was taken from genomic database for human genome (CR456776.1).2. The cDNA was synthesized using a gene synthesis service.Page 104 of 16513092399vlAttorney Docket No. 2013763-00053. The amplified DNA sequence was confirmed and cloned into a vector with the PL enzyme Turbo-ID or APEX2 linked together through a flexible linker as shown in Figure 1. Vectors were designed to allow for the fusion of APEX2 and Turbo-ID at both C and N terminals. A control fusion protein was designed that includes TBP with a deleted DNA binding domain to show the non-specific effect of a TBP-Fusion tag.4. After the fusion protein sequence was confirmed, the resultant APEX2 -TBP and Turbo-TBP was transfected into mammalian cells (e.g., HEK-293T) cells to confirm the protein expression through western blot and the correct co-localization by immunofluorescence .5. After the confirmation of fusion protein’s localization and its expression, ChlP-qPCR is performed to check the functional perturbation through the distribution of TBP fusion protein to genomic loci. This information shows where the TBP- APEX2 fusion protein has bound in the DNA.6. After functional tests of the fusion protein were performed, the TBP-APEX2 (TBP fused at the N terminus of APEX2) was selected and cloned into lentivirus vectors for infection of the tumor cells. The fusion protein was under the control of a promoter such as Tet-On system (Takara).
[0351] The nucleic acid sequence encoding the exemplified fusion protein is shown in SEQ ID NO: 1. Individual components of the construct are as follows in 5’ to 3’ order: V5 epitope tag (SEQ ID NO: 2), APEX2 PL enzyme (SEQ ID NO: 3), Glycine Asparagine Linker (SEQ ID NO: 4), and TBP (SEQ ID NO: 5).Proximity Labelling in the Primary Cells and Primary Tumors
[0352] These methods include performing proximity labeling in a neuroblastoma cell line SH5Y5 using the above-described fusion protein to identify the transcriptional regulators responsible for tumorigenesis.1. HEK293T cells were plated and infected with a control (TBP ADNA binding domain) and the experimental construct comprising a vector encoding TBP-APEX2 and a packaging complex (either a 3rd or 4th generation Lentivirus packaging system).2. The virus was harvested and the viral titer is calculated by the PCR or Go-Pro sticks.Page 105 of 16513092399vlAttorney Docket No. 2013763-00053. Neuroblastoma cells were plated in a 6 well or 10 cm dish and infected with virus for 24 hours in the presence of polybrene. After 24 hours, the media was replaced with normal growth medium.4. After 48-72 hours of infection, cells were incubated with a selection medium containing 2pg / ml of puromycin.5. Cells were continually expanded in the selection media until there were enough to freeze.6. 1 plate of control and TBP-APEX2 cells were induced by Ipg / ml of Doxycycline. In experiments where cellular response to a stimulus is measured, cells were grown in the presence of SILAC amino acids to label nascent proteins which might have been expressed in an external response. The amount of SILAC was determined empirically for each cell type.7. Biotin tyramide phenol (Iris Biotech) in DMSO (stock concentration 500 mM) was added directly to cell culture media to a final concentration of 500 pM, and the media was swirled until the precipitate is dissolved.8. After 30 minutes of incubation at 37 °C, hydrogen peroxide was added to the cell culture media to a final concentration of 1 mM to induce biotinylation. First, the hydrogen peroxide was diluted in media to 100 mM before being added to the cell culture media. After 60 seconds of very gentle swirling, the media was decanted as quickly as possible, and the cells were washed three times with 15 ml of ice-cold PBS containing 100 mM sodium azide, 100 mM sodium ascorbate, and 50 mM TROLOX (6-hydroxy-2,5,7,8- tetramethylchroman-2 -carboxylic acid).9. Cells were scraped and transferred to 15-ml Falcon tubes with ice-cold PBS, spun at 500g for 3 min, flash-frozen in liquid nitrogen, and stored at -80 °C.10. Labelled whole-cell pellets were lysed with RIPA buffer (50 mM Tris, pH 8.0, 150 mM NaCl, 1% NP-40, 0.5% sodium deoxycholate, 0.1% sodium dodecyl sulfate) with protease inhibitors (Roche) and probe-sonicated to shear genomic DNA. Wholecell lysates were clarified by centrifugation at 14,000g for 30 min at 4 °C, and protein concentration was determined by Bradford assay. 500 pL of SA magnetic bead slurry (Thermo) was used for each experimental condition to pull down the biotinlylated beads and the beads were washed with washing buffer.
[0353] Figure 1 shows a schematic of these exemplary methods for identification of global transcriptional interactome through core transcriptional machinery proximity labeling using the fusion protein described above. Green (normal) and Red (Tumor cells) are derived Page 106 of 16513092399vlAttorney Docket No. 2013763-0005from patient specific samples and then infected with a Lentivirus encoding hTBP-APEX and hTBP-TurboID enzymes. Stable clones are then selected through FACS or Puromycin to further be used for the interactome analysis. The cells are generally grown in their growth media, however in some cases to discover the drug resistance mechanisms, the cells expressing the TBP-APEX in that case can be grown in the presence of SILAC amino acids, which can provide a high temporal resolution of transcriptional response to a drug stimulus. The resultant proteins are identified using the mass spectrometry (LC MS-MS).Results:
[0354] Results from these proximity labelling experiments show that the TBP-APEX2 fusion protein was able to identify and profile transcriptional complexes in normal cells, cancer and terminally differentiated cell types of multiple species. The validation of the TBP-APEX fusion construct is described below.TBP-APEX fusion construct encodes a functional protein in Human Embryonic Kidney (HEK293) cells.
[0355] Tagging a protein of interest with another moiety can alter its function, its folding, and its expression in many ways. After engineering a specific linker (SEQ ID NO: 4) sequence and generating a fusion of TBP with APEX2 (described above), the controlled expression in the human cells was confirmed.
[0356] HEK293T cells (10cm2plate containing IxlO7cells) were transfected with plasmid encoding for APEX2 only, TBP2 only, and the TBP-APEX2 fusion protein. The constructs were induced with 2pg / ml of Doxycline to induce the expression of TBP-APEX2 protein. TBP-APEX2 protein was translated and accumulated in the nucleus of HEK293 cells. After 48 hours of the induction, cells from control and experiment conditions were scrapped from the culture dish and analyzed by the western blot using an anti-TBP antibody.Figure 2 shows that TBP-APEX2 construct encodes a fusion protein in a dose-dependent manner. Figure 2A is a schematic diagram of the experimental strategy described above for testing the expression of TBP-APEX2 fusion in HEK293T cells. Figure 2B shows the expression of TBP-APEX2 fusion in HEK293T cells in a dose dependent manner.Specifically, Figure 2B shows the results of an anti-TBP western blot showing the dose Page 107 of 16513092399vlAttorney Docket No. 2013763-0005dependent expression of TBP-APEX2 fusion protein with Ipg / ml and 2pg / ml of Doxycline. Lane 1 and 2 show results from transfecting with TBP only construct; Lane 2-3 and 4-5 show results from transfecting with the TBP-APEX2 fusion with 2pg / ml and Ipg / ml Doxycline, respectively. Images were taken with a Licor imaging system with Alexa 688 and Alexa 800 secondary antibodies again rabbit anti-TBP.TBP-APEX2 construct encodes a functional protein.
[0357] Attachment of APEX2 and other proximity labelling enzymes have been shown to hamper the protein function in some cases. To investigate whether the fusion of TBP-APEX2 is encoding a functional protein, we performed localization analysis of the TBP-APEX2 fusion protein and Chromatin Immunoprecipitation followed by qPCR assays (ChlP-qPCR). Lor localisation analysis, HEK293 cells were transfected with TBP-APEX2 and induced through 2pg / ml of Dox for 48 hours. Cells expressing TBP-APEX2 were subjected to immunofluorescence. Figure 3A shows schematic diagram showing the experimental setup for performing immunostaining of HEK293T cells. Anti-TBP antibody and Anti-V5 antibody shows the localisation of TBP-APEX2 fusion proteins in the nucleus within 10 hours after the induction (Figure 3B). Specifically, Figure 3B shows the immunostaining DAPI (blue)m Anti-V5 (Green), and Anti-TBP (Red) after 48 hours of Dox induction. Cells were fixed with Paraformaldehyde and subjected to standard immunostaining protocols. There were 2 biological replicates and each experiment was done using 2x106cells.
[0358] ChlP-qPCR assays were performed to map the binding of TBP-APEX2 protein in different active promoters in the genome. TBP knockout is a lethal phenotype. Competition ChIP procedure was performed to determine if the binding of TBP-APEX2 fusion protein changes the expression of the genes associated with its binding.
[0359] Figure 4 shows that the TBP-APEX2 fusion protein binds to genomic binding sites and induces the transcription of TATA box dependent genes. Figure 4A shows a schematic diagram of the experimental strategy which involves induction of cells with 2ug / ml of Dox for 48 hours, after which cells were subjected to lysis and ChlP-qPCR. Figures 4B and 4C shows by Anti-TBP Chromatin immunoprecipitation the relative percentage of DNA as compared to input material. Immunoprecipitated DNA was subjected to qPCR to specifically amplify TATA boxes in RPS9 and RAB5B in both induced (Figure 4B) and Page 108 of 16513092399vlAttorney Docket No. 2013763-0005uninduced (Figure 4C) conditions. Each experiment contained 3 replicates with 2xl06cells. Error bars shows SEM of n=3. *SEM p<0.05, **SEM p< 0.05 and *** p<0.02.TBP-APEX2 can map the transcriptional complexes in Neuroblastoma cell line SH- SY5Y.
[0360] In addition to determining that the construct encoding an APEX2-TBP fusion protein yields a functional protein when transfected into cells, this Example also demonstrates that the construct encoding an APEX2 fusion with TBP yields a functional protein with retained abilities for proximity labelling.
[0361] Mutant TBP-APEX2 construct (i.e., fusion protein which lacks the DNA binding domain) and wildtype (WT) TBP-APEX2 construct, as described earlier in this examples, was expressed in a Human Neuroblastoma cell line (SH-5YSY) through lentiviral transduction and cells were selected using Puromuycin selection. Cells were screened for positive clones which express the TBP-APEX2 fusion protein in a functional manner. Cells expressing the TBP-APEX2 fusion protein were further expanded and split into 10cm2plates for proximity labelling experiments.
[0362] Proximity labelling procedure was carried out as previously described in the methods section of this Example. After the proximity labelling, the cells were lysed and an anti -biotin antibody pull down was performed to isolate the biotinylated proteins. An antibody was used instead of the streptavidin beads in this method, as on beads, protein capture can be troublesome due to the steric hindrance. Figure 5A shows a schematic diagram showing the experimental strategy for transcriptional profding of the SH-5YSY cell line and normal human neuroblast cells. SILAC based proximity labelling was performed in the cells and subjected to Mass spectrometry. A silver stain gel (10% Tris glycine gel) shows the biotinylated protein pull down in neuroblastoma cells (Figure 6A). The proteins were then subjected to on beads digestions before submitting them for the LC MS / MS analysis.
[0363] Figure 6B shows a scattered plot showing the differentially expressed genes among Neuroblastoma SH-5YSY cells (Red) and normal human neuroblast cells (Green). Data was normalized and fdtered with respect to the APEX only and -Dox controls. Genes in black are expressed in both normal and diseased cells, while genes in Red and Green are Page 109 of 16513092399vlAttorney Docket No. 2013763-0005differentially expressed or silenced in diseased and normal cells respectively. Plot shows an average value plotted for each gene from n=9. Figure 6B shows that some of the labelled proteins are differentially expressed and localized in the nucleus of neuroblastoma cells.
[0364] Among them, several identified proteins are suggested to play an important role in transcriptional regulation and many of them are involved in the chromatin organization. Two of these proteins include Ggtl and ZFP238. To confirm whether these proteins are of therapeutic value, a knockout cell line for Ggtl, and for ZFP238 was generated. Surprisingly, the survival curve shows that both of these target proteins were lethal in neuroblastoma cells, and leads the apoptosis-related death of these cells within few days.
[0365] RNA Pol-II was also identified in both of the proximity labelling experiments as an active transcriptional component. RNA Pol-II was labelled in both of the proximity labelling experiments.Identification of a Mutant transcription factor which can differentiate the neuroblastoma cells into terminally differentiated neurons.
[0366] In addition to Ggtl and ZFP238 proteins identified in the proximity labeling experiments in SH-5YSY cells, one neurogenic transcription factor Stab-001 which was associated with Huwel proteosome was also identified. It was determined experimentally that the transcription factor (Stab-001) if mutated at S / A residues can have a longer half-life and not subject to an earlier degradation in the cells. Upon expression into the neuroblastoma cells, terminal differentiation of the cells into post mitotic neurons was observed. However, there was a very little differentiation was induced in the neuroblast cells.Example 2: TBP-APEX2 Fusion Protein Applied to Non-dividing Cells
[0367] Some cells such as neurons and cardiomyocytes rarely become cancerous. Transcriptional strategies in these cells are of a remarkable importance. Therefore, in this Example, the nature of transcriptional complexes in one type of non-dividing specialized cell was explored.Page 110 of 16513092399vlAttorney Docket No. 2013763-0005
[0368] Non-dividing cells present a challenge to study in that the material (e.g., number of cells) required to perform proteomics experiment is limited. To overcome this challenge, Xenopus oocyte was used as the non-dividing specialized cell type for this Example. A Xenopus oocyte is an immature egg that develops into a mature egg (i.e., a nondividing cell) of Xenopus Laevis, and has the ability to develop into a functional organism upon activation through progesterone and fertilization through a sperm. Xenopus oocytes have been shown to have stable transcriptional complexes when injected with the DNA template containing cis element / promoters for gene activation (see Gurdon et al., 2020 and Javed et al., 2022, which are herein incorporated by reference in their entirety).
[0369] In this experiment, mRNA encoding a TBP-TurboID fusion protein was injected into Xenopus oocytes and the oocytes were incubated them overnight at 18°C. The next day DNA encoding a GFP reporter driven by a CMV promoter and TATA box was injected into the Xenopus oocytes. 24 hours after DNA injection into nucleus of the oocytes, the medium was supplemented with 50um of biotin and incubated for another 6 hours.Oocytes were then washed with 1X-MBS supplemented with antibiotics and subsequently lysed in oocytes lysis medium. The lysate was centrifuged at 13000 r.p.m. and the pellet was discarded. The supernatant from this fraction was transferred to another tube and incubated with 50 ul of Streptavidin beads for overnight capturing of the biotinylated proteins. The next day the beads were washed twice with the lysis buffer and then with the washing buffer according to the manufacturer protocol (thermofisher cat #88816). The washed beads were then subjected to on-beads digestions followed by LC MS / MS analysis.
[0370] The result obtained by mass spectrometry yielded information about the nature of transcriptional complexes found in the oocytes (Figure 7). The proteins found in complex in the Germinal vesicle (GV) of the oocyte were characterized by their class. The grey plot (top) shown in Figure 7A shows proteins in non-transcription factor families, i.e., that are not directly involved in transcription (i.e., by comparing to controls where cells were incubated with streptavidin beads without performing a PL reaction). In the orange plot (bottom) of Figure 7A, all of the transcription factors and their abundance in the GV of the oocytes is listed. For the second phase of characterization, transcription factors according to their cell atlas location was plotted (Figure 7B) determined by analyzing characteristic sequences (e.g., nuclear localization sequences and other organelle specific sequences). For determining cell Page 111 of 16513092399vlAttorney Docket No. 2013763-0005atlas location, experiments were done in triplicate with each sample containing further 3 replicates, each containing 30 individual oocytes. Data was normalized using controls injected with water and / or without Biotin.
[0371] In this example, proteins identified in the nucleus were analyzed separately from those identified in the cytoplasm. Many of the known important transcription factors such as YBX2, TCF25, PURA and DNAJC2 were found to remain outside the nucleus, while most of the transcription factors which are related to maintenance of transcription such as E2E5, SMARC1, HMGA2, and EZH2 were found to reside exclusively in the nucleus. The ratio of all transcription factors, outside and inside the nucleus was also determined in subsequent analysis.TBP-002 induces rapid terminal differentiation dividing cells
[0372] From the Xenopus oocyte experiment, a component of general transcription machinery called factor “TBP-002” was identified for being present only in eggs or oocytes or a few terminally differentiated cells in human. Presence of TBP-002 stabilizes the transcription in non-dividing cells, to express certain genes in a highly efficient and sustained way.
[0373] In this experiment, TBP-002 was expressed in a cancer cell line (i.e., rapidly dividing cells) to test whether TBP-002 would cause a rapid differentiation toward a specialized state. Lentiviruses were engineered to express TBP2 in mouse C2C12 cells. Normally, when serum is withdrawn from the media and C2C12 cells are allowed to differentiate into terminally differentiated myotubes, the process takes about 7-10 days. When C2C12 cells were induced to express TBP2 (using doxycycline induction), C2C12 cells underwent differentiation in about 48 hours (see Figure 8) without any serum starvation or without any addition of signaling molecules like insulin. Expression of differentiation markers using qPCR shows the early upregulation of the genes involved in cell cycle exit and for the stable expression of myogenic factors (see Figure 9).
[0374] These results suggest that TBP2 may be useful in inducing the differentiation of malignant cancer cells.Page 112 of 16513092399vlAttorney Docket No. 2013763-0005
[0375] Single chain variable domain nanobodies are an effective tool for disease treatment and in the field molecular biology. Nanobodies that recognize Pol-II have previously been developed (see Shibuta et al., 2021). In this Example, an RNA-Pol-II mintbody has been engineered such that it is specific not only for Pol-II, but more specifically, it targets active Pol-II complexes in a living cell. The nucleic acid sequence encoding the RNA-Pol-II mintbody sequence is shown in SEQ ID NO: 7. Specifically, the RNA-Pol-II mintbody was engineered such that it recognizes and binds to the Ser2 modification on active Pol-II and does not recognize and bind to Pol-II when it does not contain the Ser2 phosphorylation, and instead contains a Ser5 phosphorylation, which is a Pol-II enzyme that is “paused” or not actively transcribing (see Figure 10).
[0376] A fusion protein as described herein was then generated using the selective nanobody (“RNA-Pol-II-Ser2”) by fusing the mintbody with APEX2 or Turbo-ID for selective targeting and labelling of active transcriptional components. The fusion protein sequence comprises a RNA-Pol-II-Ser2 mintbody sequence (SEQ ID NO: 7), a linker sequence (SEQ ID NO: 8), APEX2 enzyme (SEQ ID NO: 9). The full sequence is shown in SEQ ID NO: 6. Exemplary fusion proteins and engineering strategy for the same are shown in Figure 10. The RNA-Pol II-Ser2 mintbody described in this Example has been used for visualizing the distribution of RNA-Pol-II-Ser2 active subunits on the chromatin in human cells. Exploitation of the RNA-Pol-II-Ser2 mintbody specificity with APEX mediated labelling chemistry provides mechanistic insights in transcriptional regulation of diseased cells (e.g., how tumor cells maintain their transcriptional control during metastasis).ResultsPol-II serine mintbody fused with APEX2 reveals differentially interacting proteins among neuroblastoma and neuroblast cells
[0377] In this experiment, proximity labelling was performed to validate the specificity of the RNA-Pol-II-Ser2 mintbody.Page 113 of 16513092399vlAttorney Docket No. 2013763-0005
[0378] Human neuroblastoma and normal human neuroblast cells were engineered with lentivirus encoding a fusion protein comprising APEX2 fused with the RNA-Pol-II-Ser2 mintbody (encoded by SEQ ID NO: 7). The cells were treated with mitmyosin and then subjected to proximity labelling with the APEX2-RNA-Pol-II-Ser2 mintbody fusion protein according to proximity labeling methods described in earlier Examples.
[0379] Briefly, the proximity labelling performed for 1 minute with biotin phenol-supplemented human neuroblastoma and neuroblastoma cells, followed by cold cell lysis.
[0380] Figure 11 shows the differential expression pattern of the proteins labeled from neuroblastoma and neuroblast cells. LC MS / MS data was analyzed initially through max quant with the peptide count threshold of 25, and false discovery rate to 1%.Differentially expressed protein interactors were plotted on the fold change detection of X axis and -logP value on the Y axis. The data shown here represent 3 independent experiments (n=3).
[0381] Mass spectrometry results show a significant difference in interaction of the transcriptional machinery components among neuroblastoma and neuroblast cells. The data was filtered by deleting the abundant proteins, like nuclear pore and cytoskeleton-related proteins. As a positive control, RNA-Pol II C Terminal domain was one of the most abundant proteins to be biotinylated by the reaction.
[0382] This result demonstrates that APEX2 fused with the RNA-Pol-II-Ser2 mintbody can label the active transcriptional complexes in both primary cells (neuroblasts) and transformed tumors (neuroblastoma).Example 4: Targeted Pull Down of DNA locus for Identification of Transcriptional Activation and Repression Complexes using Split GFP proximity labeling system
[0383] Drug discovery relies on identification of targets from e.g., genetic screens, proteomics and DNA sequencing. However, in cells, proteins, such as those interacting with transcriptional machinery, perform their function by interacting or recruiting other proteins. For example, core transcriptional machinery including TBP (TATA binding protein) binds to TATA box on DNA, and recruits the transcriptional initiation complex. Mechanistic Page 114 of 16513092399vlAttorney Docket No. 2013763-0005insights on how DNA binding proteins (e.g., TBP) interact with specific sequences combinatorically will provide valuable target identification which are currently unknown.
[0384] The present Examples illustrates use of various systems described herein for targeting particular DNA loci, and for pulling down protein-DNA and DNA-DNA complexes (e.g., transcriptional complexes, e.g., of activation and / or repression nature) from cells (e.g., living human cells). The present Example also describes exemplary methods and procedures for using the described systems. Specifically, the following describes experiments and methods relating to pull down of a DNA region of interest (target genomic locus) and labeling the interactome of a protein of interest (POI) or gene of interest (GOI) with great precision.Split Green Fluorescent Protein (GFP) system for Proximity Labelling.
[0385] The present Example demonstrates the use of a split GFP protein “tripartite” proximity labelling system that can be used for proximity labeling, e.g., of transcriptional machinery. Split GFP is a technology in which a GFP protein is split into halves including a first fragment GFP Beta 10 and second fragment GFP 11. Both fragments are not fluorescent on their own. Only when in proximity to each do they form a complex which can then be recognized by the third major fragment, called GFP l-9r. GFP l-9r recognizes and forms a complex with dimerized GFP 10 and GFP 11 fragments, forming the whole fluorescent GFP (FIG. 12). Under normal cellular conditions, the entropy of the GFP fragments is high such that they cannot reconstitute to a functional protein.
[0386] The present Example utilizes a split GFP system for proximity labeling of interacting proteins. For example, FIG. 12 shows a schematic of an exemplary tripartite split GFP system (Figure adapted from S Castillo et al., 2023). In this exemplary system, Protein A and Protein B (exemplary interacting POIs) are fused with GFP Beta 10 and GFP Beta 11, respectively. GFP Beta 11 (“GFP 11”) and GFP Beta 10 (“GFP 10”) are 20 amino acid tags. Using this system, when the two POIs (A and B) are in close proximity or are interacting with one another, the GFP Beta 10 and GFP Beta 11 proteins dimerize, and when the system is supplemented with GFP l-9r, the GFP fragments complex to make a functional reconstituted GFP that can be visualized.Page 115 of 16513092399vlAttorney Docket No. 2013763-0005Using split GFP systems for proximity labelling of two interacting proteins and their complexes
[0387] Spit systems (e.g., split GFP systems described herein) can be utilized to study and identify protein-protein and DNA-protein interactions both in live and fixed cells / tissues. Specifically, in this Example, a GFP 10 fragment was fused with a protein of interest (POI) “protein A” and a 20 amino acid tag and a GFP 11 fragment was fused with POI “protein B” and a 20 amino acid tag, and both were expressed in a population of cells. A fusion protein containing GFP 1-9 fused with APEX2 proximity labeling enzyme was then expressed in the population of cells. The recognition of reconstituted GFP will subsequently allow APEX2 to label surrounding proteins.
[0388] This method can reduce and / or eliminate off target labelling, as the interaction of the proteins canfirst be visualized prior to proximity labelling (FIG. 13). FIG. 13 shows a schematic diagram of interaction-induced proximity labelling coupled with fluorescence visualization. FIG. 13A shows Proteins A and B fused with GFP 10 and GFP 11 fragments, respectively, and a GFPl-9r (GFP detection fragment) fused with APEX2. Upon interaction of all three fragments, GFP is reconstituted and can be visualized (FIG. 13B). Proximity labelling of interacting components may be achieved by addition biotin phenol and hydrogen peroxide into the cultured cells or tissues, as described in the present disclosure.Site specific pull down of repressive or transcriptional activation complex by split- APEX2 GFP system
[0389] In this experiment, the above exemplary split GFP protein system was utilized to pull down transcriptional complexes that are bound to a specific DNA sequence (e.g., a “target genomic locus”). This may be achieved by utilizing (1) a fusion protein that includes a GFP 10 sequence fused to dCas9 via a 20 amino acid linker (“dead Cas9” or “endonuclease deficient Cas9” is capable of binding to DNA and guide RNA) (“GFP10-dCas9 fusion protein”); (2) a fusion protein that includes GFP11 is fused with a POI (“GFP 11 -Protein A”); and (3) a fusion protein that includes GFPl-9r fused to an APEX2 proximity labelling enzyme (“GFPl-9r-APEX2”) (see fusion proteins of FIG. 14A). When expressed in a population of cells, the GFP10-dCas9 fusion protein localizes at the target locus through its specific gRNA, the POI (“Protein A”) binds to the target locus (or an interacting component Page 116 of 16513092399vlAttorney Docket No. 2013763-0005at the target locus), and the GFPl-9r-APEX2 fusion protein will reconstitute with the GFP10 and GFP11 fragments to fluoresce, which allows for the visualization. Interacting components are subsequently biotinylated upon addition of the hydrogen peroxide and biotin phenol to the system. The protein complex (including the associated fusion proteins) can be detected by mass spectrometry following cell lysis (FIG. 14B). This method provides high spatiotemporal control and specificity compared to other locus pull down techniques and other global proximity labelling methods.DNA Sequences and Maps
[0390] Exemplary sequences for the fusion proteins and their components and target loci described in this Example are provided below in Table 1.Table 1:Page 117 of 16513092399vlAtorney Docket No. 2013763-0005Page 118 of 16513092399vlAtorney Docket No. 2013763-0005Page 119 of 16513092399vlAtorney Docket No. 2013763-0005Page 120 of 16513092399vlAtorney Docket No. 2013763-0005Page 121 of 16513092399vlAttorney Docket No. 2013763-0005Experimental strategy for Locus Pull Down
[0391] The present experiment, among other things, confirms reproducibility and functionality of the split GFP protein “tripartite” proximity labelling system described herein. In this experiment, TetOn and TetOff systems are utilized to confirm the functionality of the systems described in this Example used for locus pull down.Use of Tetracycline (TetOn system) to validate locus pull down technology
[0392] TetOn and TetOff systems are inducible gene expression systems that allow for precise temporal and spatial control of gene expression. Utilizing a TetOn system, gene expression is activated when tetracycline or its analog, doxycycline, is present. The reverse tetracycline-controlled transactivator (rtTA), a component of the TetOn system, binds to tetracycline and interacts with the tetracycline-responsive element (TRE) in the promoter region, activating transcription of the target gene. Conversely, in the TetOff system, gene expression is repressed when tetracycline is present. Here, the tetracycline-controlled transactivator (tTA) binds to the TRE and activates transcription in the absence of tetracycline.Page 122 of 16513092399vlAttorney Docket No. 2013763-0005When tetracycline is added, it binds to tTA, preventing its interaction with the TRE and thus inhibiting gene expression. These systems are particularly useful for studying genes essential for development or those that may be lethal when continuously expressed, as they allow researchers to turn gene expression on or off at specific times during an experiment.
[0393] FIG. 15 shows a schematic of the components and functionality of a TetOn system. In the present experiment, a construct which contains the Tet-element at the 5’ end followed by a mCherry reporter is utilized (FIG. 15A). When the gene (exemplary target genomic locus) is in active form, it will turn the cells red (mCherry will express / fluoresce), and when the gene is in silent form, cells will not turn red (no mCherry expression / fluorescence).Table 2: Components of the exemplary experimental system used in the present experiment.
[0394] The present experiment is designed to detect when the system components (as shown in Table 3) are co-localized (by GFP fluorescence), by using a TetOn system constructed to constitutively express a gene (target genomic locus of interest).Table 3: Exemplary experimental conditions / components and rationale.Page 123 of 16513092399vlAttorney Docket No. 2013763-0005
[0395] In subsequent experiments, the exemplary system described in this experiment will utilize tetracycline to block the binding of the TetR to its binding sequence, and the GFP fluorescence will subsequently disappear. Such a system provides a for validation of and functionality of the split GFP protein “tripartite” proximity labelling system described herein.Example 5: Isolation of Chromatin loci by RNA guide dCAS9 for proteomics study
[0396] Transcription factor binding dynamics depends on chromatin structure, its cognate binding site, and host cell co-factor pools. Change in co-factor pools can lead to a differential gene expression pathways in a cell (see Donaghey et al., 2018). There are several approaches used to isolate a target genomic loci and associated transcription factors. Such approaches are useful in the study the temporal dynamics of transcription. The present Example provides an approach to isolating genomic loci and associated transcription factors in Xenopus oocytes.
[0397] The present Example provides a method that includes isolating (or “pull down”) of fragments of injected DNA representing a target genomic loci (DNA-FF) in oocytes using a split GFP protein “tripartite” proximity labelling system described herein including a specific guide RNA and an associated dCas9 protein targeted to the target genomic loci.Experimental procedures:
[0398] On day 0, 7ng / oocyte of mRNA encoding dCas9 and Ascl-1 was injected into the cytoplasm of a population of oocytes. On the following day, the population of oocytes were injected with reporter DNA and guide RNAs which direct dCas9 to the ta...
Claims
Attorney Docket No. 2013763-0005CLAIMSWe claim:
1. A system comprising :a first fusion protein comprising a first protein of interest (POI) and a first protein fragment of an indicator protein; anda second fusion protein comprising a proximity labeling enzyme and a second protein fragment of the indicator protein.
2. The system of claim 1, where the system further comprises a third fusion protein comprising:a second POI and a third protein fragment of the indicator protein.
3. The system of claim 2, wherein the first fusion protein, the second fusion protein, and the third fusion protein, when in proximity to one another, associate to form a complex via the association of the first protein fragment, second protein fragment, and third protein fragment, thereby forming a functional indicator protein.
4. The system of any one of claims 1-3, wherein:(i) the first POI and the first protein fragment are connected via a linker;(ii) the proximity labelling enzyme and the second protein fragment are connected via a linker; and / or(iii) the second POI and the third protein fragment are connected via a linker.
5. The system any one of claims 2-4, wherein the first POI and the second POI are the same.
6. The system of any one of claims 2-5, wherein the first POI and the second POI are different.
7. The system of any one of claims 2-6, wherein the proximity labeling enzyme is characterized by its ability to label proteins that interact with the first POI and / or second POI when they are in proximity therewith.Attorney Docket No. 2013763-00058. The system of any one of claims 1-7, wherein the first POI and / or the second POI is a transcriptional machinery component.
9. The system of claim 8, wherein the transcriptional machinery component is TATA-binding protein (TBP).
10. The system of any one of claims 3-9, wherein when the system is delivered to a population of cells, the system is capable of labeling proteins that interact with a protein of interest.
11. A system comprising :(i) a first fusion protein comprising a POI and a first protein fragment of an indicator protein;(ii) a second fusion protein comprising a DNA-targeting moiety that binds to a target genomic locus and a second protein fragment of the indicator protein; and(iii) a third fusion protein comprising a proximity labeling enzyme and a third protein fragment of the indicator protein.
12. The system of claim 11, wherein:(i) the POI and the first protein fragment are connected via a linker;(ii) the DNA-targeting moiety and the second protein fragment are connected via a linker; and / or(iii) the proximity labeling enzyme and the third protein fragment are connected via a linker.
13. The system of claim 11 or 12, wherein the proximity labeling enzyme is characterized by its ability to label proteins that interact with the POI and / or target genomic locus when they are in proximity therewith.
14. The system of any one of claims 11-13, wherein the first POI is a transcriptional machinery component.Attorney Docket No. 2013763-000515. The system of claim 14, wherein the transcriptional machinery component is TATA-binding protein (TBP).
16. The complex of any one of claims 11-15, wherein the DNA-targeting moiety comprises a dCas9 protein or a locked nucleic acid (LNA).
17. The fusion protein of any one of claims 12-23, wherein the target genomic locus comprises a TATA box sequence.
18. The system of any one of claims 1-17, wherein the proximity labelling enzyme comprises a biotin ligase or a peroxidase.
19. The system of claim 18, wherein the proximity labelling enzyme comprises an enzyme selected from a group consisting of: APEX, APEX2, Bio-ID, Turbo-ID, and mini Turbo-ID.
20. The system of claim 19, wherein the proximity labelling enzyme comprises APEX2.
21. The system of any one of claims 11-20, wherein when the system is delivered to a population of cells, the system is capable of labeling proteins that interact with the protein of interest and / or the target genomic locus.
22. The system of claim 21, wherein the first fusion protein and the second fusion protein dimerize via the first protein fragment and the second protein fragment.
23. The system of claim 2, wherein the first fusion protein, the second fusion protein, and the third fusion protein, when in proximity to one another (e.g., at the target genomic locus or in proximity to the POI), associate to form a complex via the association of the first protein fragment, second protein fragment, and third protein fragment, thereby forming a functional indicator protein.
24. The system of claim 23, wherein the first, second, and third protein fragments are not a functional indicator protein unless they are associated in a complex.Attorney Docket No. 2013763-000525. The system of any one of claims 1-24, wherein the indicator protein comprises a fluorescent protein.
26. The system of claim 25, wherein the fluorescent protein comprises a green fluorescent protein (GFP).
27. The system of claim 26, wherein the first protein fragment comprises a GFP 10, a GFP11, or a GFPl-9r protein fragment.
28. The system of claim 26 or 27, wherein the second protein fragment comprises a GFP 10, a GFP 11 , or a GFP 1 -9r protein fragment.
29. The system of any one of claims 26-28, wherein the third protein fragment comprises a GFP 10, a GFP11, or a GFPl-9r protein fragment.
30. The system of any one of claims 26-29, wherein the first protein fragment comprises GFP11.
31. The system of any one of claims 26-30, wherein the second protein fragment comprises GFP 10.
32. The system of any one of claims 26-31, wherein the third protein fragment comprises GFPl-9r.
33. A fusion protein comprising :a proximity labeling enzyme APEX2 and a GFP ( 1 -9r) protein, wherein the proximity labeling enzyme is connected to the GFP ( 1 -9r) protein via a linker.
34. A fusion protein comprising:a DNA-targeting moiety fused to a protein fragment GFPl-9r.Attorney Docket No. 2013763-000535. The fusion protein of claim 34, wherein the DNA-targeting moiety comprises a dCas9 protein or a locked nucleic acid (LN A).
36. The fusion protein of claim 34 or 35, wherein the DNA-targeting moiety and the first protein fragment are connected via a linker.
37. The fusion protein of any one of claims 34-36, wherein the DNA-targeting moiety targets a target genomic loci.
38. A method of isolating a repressive or transcriptional activation complex associated with a target genomic locus, the method comprising:delivering the system of any one of claims 11-32 to a cell,maintaining the system under conditions and for a time sufficient such that one or more of the fusion proteins interacts with the target genomic locus:exposing the system to conditions under which the proximity labeling enzyme labels components in proximity with the target genomic locus;isolating the complex comprising the labeled component(s) comprising the repressive or transcriptional activation complex interacting with the target genomic locus from the system.
39. A method of characterizing a protein that interacts with a target protein of interest or a target genomic locus of interest, the method comprising steps of:delivering the system of any one of claims 11-32 to a cell;maintaining the system under conditions and for a time sufficient such that one or more of the fusion proteins interacts with the target protein of interest or target genomic locus of interest;exposing the system to conditions under which the proximity labeling enzyme labels components in proximity with the target protein of interest or target genomic locus;separating labeled component(s) from the system; andcharacterizing the labeled component(s).
40. A method comprising a step of:characterizing a labeled interacting component that was generated by:Attorney Docket No. 2013763-0005delivering the system of any one of claims 11-32 to a cell;maintaining the system under conditions and for a time sufficient such that one or more of the fusion proteins interacts with the target genomic locus:exposing the system to conditions under which the proximity labeling enzyme labels components in proximity with the target protein of interest;separating labeled component(s) from the system; andcharacterizing the labeled component(s).
41. The method of any one of claims 38-40, wherein the first fusion protein, the second fusion protein, and the third fusion protein co-localize at the target genomic locus.
42. The method of any one of claims 38-41, wherein the labelled cell can be visualized by fluorescence microscopy.
43. The method of any one of claims 38-42, wherein the cell comprises a cancer cell.
44. The method of any one of claims 38-42, wherein the cell is a HEK293 cell.
45. The method of any one of claims 38-42, wherein the cell comprises a terminally differentiated cell.
46. The method of any one of claims 38-42, wherein the cell comprises a non-dividing cell.
47. The method of claim 46, wherein the non-dividing cells is an oocyte.
48. The method of any one of claims 38-47, wherein exposing the system to conditions under which the proximity labeling enzyme labels components in proximity with the target protein of interest or target genomic locus comprises adding biotin to the first and second systems.
49. The method of claim 48, wherein the labeled component(s) comprise biotinylated proteins.Attorney Docket No. 2013763-000550. The method of any one of claims 39-49, wherein separating the labeled components comprises extracting protein from the system and contacting the extracted protein with an anti -biotin antibody or streptavidin beads.
51. The method of any one of claims 39-50, wherein characterizing the labeled component(s) comprises analyzing the component(s) using LC MS / MS analysis.