Means and methods for targeting endogenous condensates
A system using IDR polypeptides and effector domains allows for the modification and manipulation of intracellular condensates, addressing the limitations of current tools by enabling detailed analysis of these structures.
Patent Information
- Application Number
- JP2025535963
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-05-04
- Filing Date
- 2023-12-18
- Publication Date
- 2026-01-06
AI Technical Summary
Current tools for studying intracellular condensates, which are membraneless and small, are insufficient for obtaining genome-wide and proteome-wide information, limiting our understanding of their role in cell types like stem cells and cancer cells.
A system comprising an IDR polypeptide sequence tract and an effector polypeptide domain that can covalently modify target biomolecules associated with endogenous intracellular condensates, allowing for manipulation and modification of these condensates.
Enables the modification of intracellular condensates, facilitating the introduction of polypeptides and altering their chemical or physical properties, thereby providing a means to study and manipulate these condensates without disrupting their key characteristics.
Smart Images

Figure 2026500369000013 
Figure 2026500369000014 
Figure 2026500369000015
Abstract
Description
[Technical Field]
[0001] This application claims priority to U.S. Provisional Application No. 63 / 387 / 982, filed December 19, 2022, and European Application (EP) No. 23171660.6, filed May 4, 2023, both of which are incorporated herein in their entireties.
[0002] The present invention relates to multimodular polypeptide constructs that facilitate the manipulation of endogenous intracellular condensates (e.g., condensates generated by liquid-liquid phase separation or biological phase transitions) of proteins containing intrinsically disordered regions (IDRs). The invention further relates to methods of using the multimodular polypeptides in analyzing the components of the condensates. [Background technology]
[0003] The emerging science of liquid-liquid phase separation and biomolecular condensates is impacting diverse fields, including cell biology, neuroscience, and cancer biology. Transcriptional condensates, in particular, have attracted attention because their role in transcription may be crucial for understanding gene expression in various cell types, particularly stem cells and cancer cells. However, available tools for studying these phenomena are limited because condensates are membraneless, liquid, and typically small (<1 μm). In this context, condensation studies primarily rely on microscopy, which allows visualization of condensates without manipulating cells. However, this approach is insufficient to obtain genome-wide and proteome-wide information about the role of condensates. Therefore, new systems and tools for studying liquid condensates while preserving their key characteristics during experiments are desirable.
[0004] Liu et al. (FRONTIERS IN ONCOLOGY 12, 21 March 2022) report on post-translational modifications of BRD4.
[0005] Chiang (DRUG DISCOVERY TODAY:TECHNOLOGIES 19, (2016),17-22) discusses BRD4 phosphorylation and the interaction of BRD4 with drugs that target BRD4.
[0006] Vershininz et al. (SCIENCE ADVANCES 22, 26 May 2021) report that methylation of BRD4 by SETD6 regulates selective transcription and controls mRNA translation.
[0007] Dzuricky et al. (Nature Chemistry (2020) 12(9) 814-825) disclose a system that uses intrinsically disordered proteins to transport dye molecules or alkaline phosphatase into condensates. [Prior art documents] [Non-patent literature]
[0008] [Non-Patent Document 1] Liu et al.,FRONTIERS IN ONCOLOGY 12, 21 March 2022 [Non-patent document 2] Chiang,DRUG DISCOVERY TODAY:TECHNOLOGIES 19, (2016),17-22 [Non-patent document 3] Vershininz et al.,SCIENCE ADVANCES 22,26 May 2021 [Non-patent document 4] Dzuricky et al.,Nature Chemistry (2020)12(9)814-825 Summary of the Invention [Problem to be solved by the invention]
[0009] Based on the above-mentioned state of the art, it is an object of the present invention to provide means and methods for studying intracellular condensates. [Means for solving the problem]
[0010] This object is achieved by the subject matter of the independent claims herein, with further advantageous embodiments described in the dependent claims herein, the examples, the figures and the general description.
[0011] Summary of the Invention One aspect of the present invention relates to a system for modifying a target biomolecule associated with endogenous intracellular condensates. Endogenous intracellular condensates can be formed by intracellular proteins containing intrinsically disordered regions (IDRs). The system of the present invention comprises, as its minimum components, an IDR polypeptide sequence tract containing an intrinsically disordered region of an intracellular protein that forms or associates with intracellular condensates, and an effector polypeptide domain capable of covalently modifying the target biomolecule.
[0012] An alternative invention relates to a system for modifying the chemical or physical properties of endogenous intracellular condensates.
[0013] Another aspect of the invention relates to nucleic acid sequences encoding the systems or encoding components of the systems according to the above disclosed aspects of the invention.
[0014] The present invention further relates to a method for modifying a target, in particular a target protein, associated with endogenous intracellular condensates formed by an intracellular protein comprising an intrinsically disordered region, the method comprising the steps of providing a cell comprising an intracellular protein comprising an intrinsically disordered region, and expressing in said cell a system as defined according to the first aspect of the invention in any of its embodiments.
[0015] Terms and Definitions For the purposes of interpreting this specification, the following definitions shall apply, and where appropriate, terms used in the singular shall also include the plural and vice versa. In the event that a definition set forth below conflicts with any document incorporated herein by reference, the definition set forth herein shall control.
[0016] As used herein, the terms "comprising," "having," "containing," "including," and other similar forms, and their grammatical equivalents, are intended to be equivalent in meaning and to be open-ended in that the listing of one or more items following any one of these words does not imply an exhaustive listing of such one or more items, or that it is limited to only the listed item or items. For example, an item "comprising" components A, B, and C can consist of components A, B, and C (i.e., contain only components A, B, and C), or it can include not only components A, B, and C, but also one or more other ingredients. Thus, "comprising" and its similar forms, and its grammatical equivalents, are intended and understood to include disclosure of "consisting essentially of" or "consisting of" embodiments.
[0017] Where a range of values is provided, unless the context clearly dictates otherwise, it is understood that each intervening value, to the tenth of the unit of the lower limit, between the upper and lower limit of that range, and any other stated or intervening value within that stated range, is encompassed within the disclosure, subject to any specifically excluded limit in the stated range. Where one or both of the limits are included in the stated range, ranges excluding either or both of those included limits are also included in the disclosure.
[0018] As used herein, reference to "about" a value or parameter includes (and describes) a variation on the value or parameter itself. For example, a statement referring to "about X" also includes the statement "X."
[0019] As used in this specification, including the appended claims, the singular forms "a," "or," and "the" include plural references unless the context clearly dictates otherwise.
[0020] As used herein, "and / or" is considered to specifically describe each of the two specified features or components with or without the other features or components. Thus, the term "and / or" used in phrases such as "A and / or B" is intended to include "A and B," "A or B," "A alone," and "B alone." Similarly, the term "and / or" used in phrases such as "A, B, and / or C" is intended to encompass each of the following aspects: A, B, and C; A, B, or C; A or C; A or B; B or C; A and C; A and B; B and C; A alone; B alone; and C alone.
[0021] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art (e.g., cell culture, molecular genetics, nucleic acid chemistry, hybridization techniques and biochemistry, organic synthesis). Standard procedures are used for molecular, genetic, and biochemical procedures (see generally, Sambrook et al., Molecular Cloning: A Laboratory Manual, 4th ed. (2012) Cold Spring Harbor Laboratory Press, Cold Spring Harbor, NY, and Ausubel et al., Short Protocols in Molecular Biology (2002) 5th ed., John Wiley & Sons, Inc.) and chemical procedures.
[0022] The term "condensate" as used herein refers to membrane-free components within a cell characterized by a concentration of biomolecules that differs from the surrounding cell. The term "condensate" is used synonymously with the term biomolecular condensate and refers to membrane-free components composed of selectively concentrated biomolecules with liquid-like properties. See Hyman et al., Liquid-Liquid Phase Separation in Biology, Annual Review of Cell and Developmental Biology Vol. 30 (2014) pp 39-58; Cho WK, Spille J. et al., Science 2018 Jul 27; 361 (6400): 412-415; Narayanan et al., Elife 2019, https: / / doi.org / 10.7554 / eLife.39695, which are incorporated herein by reference.
[0023] Certain embodiments of "endogenous condensates" referred to herein relate to protein condensates. Certain embodiments of "endogenous condensates" referred to herein relate to small, sub-diffractive (<750 nm diameter) condensates. Certain embodiments of "endogenous condensates" referred to herein relate to transcriptional condensates formed by RNA, transcription factors, RNA polymerase, and genomic DNA.
[0024] The term "endogenous condensates" herein relates to condensates that are present in cells prior to the induction or application of the system according to the present invention. In other words, rather than providing the polypeptide components of the condensate via expression of the fusion polypeptides disclosed herein, the methods of the present invention target condensates that are pre-existing under physiological conditions.
[0025] The term "effector polypeptide" as used herein refers to a polypeptide inserted into endogenous intracellular condensates that fulfils the chemical or physical properties of endogenous intracellular condensates. In certain embodiments, the effector polypeptide directly or indirectly modifies the target protein by inducing the formation or cleavage of a covalent bond.
[0026] The term "system" as used herein relates to a combination of functional elements embodied in the form of a polypeptide and comprising at least one or several IDR tracts and an effector polypeptide. A system according to the invention can be embodied by a single polypeptide chain or by two, three or more separate polypeptide chains that assemble upon expression or in response to a stimulus.
[0027] array Sequences similar or homologous (e.g., at least about 70% sequence identity) to the sequences disclosed herein are also part of the present invention. In some embodiments, sequence identity at the amino acid level can be about 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% or more. At the nucleic acid level, sequence identity can be about 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% or more.
[0028] In the present context, the terms "sequence identity" and "percentage of sequence identity" refer to a quantitative parameter that represents the results of sequence comparison, determined by comparing two aligned sequences position by position. Methods for aligning sequences for comparison are well known in the art. Sequence alignment for comparison can be performed by the local homology algorithm of Smith and Waterman, Adv. Appl. Math. 2:482 (1981), the global alignment algorithm of Needleman and Wunsch, J. Mol. Biol. 48:443 (1970), the similarity search method of Pearson and Lipman, Proc. Nat. Acad. Sci. 85:2444 (1988), or computerized implementations of these algorithms, including, but not limited to, CLUSTAL, GAP, BESTFIT, BLAST, FASTA, and TFASTA. Software for performing BLAST analyses is publicly available through, for example, the National Center for Biotechnology Information (http: / / blast.ncbi.nlm.nih.gov / ).
[0029] An example of a comparison of amino acid sequences is the BLASTP algorithm using default settings: Expect threshold: 10; Word size: 3; Max matches in a query range: 0; Matrix: BLOSUM62; Gap Costs: Existence 11, Extension 1; Compositional adjustments: Conditional compositional score matrix adjustment. One such example for comparison of nucleic acid sequences is the BLASTN algorithm using default settings: Expect threshold: 10; Word size: 28; Max matches in a query range: 0; Match / Mismatch Scores: 1.-2; Gap costs: Linear. Unless otherwise specified, sequence identity values provided herein refer to values obtained using the BLAST family of programs using the above-specified default parameters for protein and nucleic acid comparisons, respectively (Altschul, J. Mol. Biol. 215:403-410 (1990)).
[0030] Reference to identical sequences without specifying a percentage includes the meaning of 100% identical sequences (ie, the same sequence).
[0031] The term "having substantially the same biological activity" as used herein relates to the function described for a specific sequence in the present invention, such as the ability of an IDR tract to promote the integration of the construct described in claim 1 herein into endogenous intracellular condensates formed by proteins that have the same IDR tract as part of their native sequence.
[0032] General Biochemistry: Peptides, Amino Acid Sequences The term "polypeptide" as used herein refers to a molecule of 50 or more amino acids forming a linear chain in which the amino acids are connected by peptide bonds. The amino acid sequence of a polypeptide may refer to the amino acid sequence of an entire protein (as found physiologically) or a fragment thereof. The terms "polypeptide" and "protein" are used interchangeably herein and include proteins and fragments thereof. Polypeptides are disclosed herein as amino acid residue sequences.
[0033] The term "peptide" as used herein relates to a molecule consisting of up to 50 amino acids, in particular 8 to 30 amino acids, more in particular 8 to 15 amino acids, which form a linear chain in which the amino acids are joined by peptide bonds.
[0034] The sequence of amino acid residues is written from the amino terminus to the carboxyl terminus. The capital letters at the sequence positions refer to the L-amino acid in single-letter code (Stryer, Biochemistry, Vol. 3, p. 21). The lowercase letters at the amino acid sequence positions refer to the corresponding D- or (2R)-amino acid. The sequence is written from left to right from the amino terminus to the carboxyl terminus. Following standard nomenclature, the sequence of amino acid residues is represented by either the three-letter or single-letter code as follows: Alanine (Ala, A), Arginine (Arg, R), Asparagine (Asn, N), Aspartic Acid (Asp, D), Cysteine (Cys, C), Glutamine (Gln, Q), Glutamic Acid (Glu, E), Glycine (Gly, G), Histidine (His, H), Isoleucine (Ile, I), Leucine (Leu, L), Lysine (Lys, K), Methionine (Met, M), Phenylalanine (Phe, F), Proline (Pro, P), Serine (Ser, S), Threonine (Thr, T), Tryptophan (Trp, W), Tyrosine (Tyr, Y), and Valine (Val, V).
[0035] The term "variant" refers to a polypeptide that differs from a reference polypeptide but retains essential properties. A typical variant of a polypeptide differs from another, reference polypeptide in its primary amino acid sequence. Generally, differences are limited so that the sequences of the reference and variant are similar overall and, in many regions, identical. A variant and reference polypeptide may differ in amino acid sequence by one or more modifications (e.g., substitutions, additions, and / or deletions). A substituted or inserted amino acid residue may or may not be one encoded by the genetic code. A polypeptide variant may be naturally occurring, such as an allelic variant, or it may be a variant that is not known to occur naturally.
[0036] As used herein, the term "amino acid linker" refers to a polypeptide of variable length used to connect two polypeptides to produce a single polypeptide chain. Exemplary embodiments of linkers useful in practicing the invention defined herein are oligopeptide chains of 1, 2, 3, 4, 5, 10, 20, 30, 40, or 50 amino acids. A non-limiting example of an amino acid linker is a monomer or a di-, tri-, or tetramer of a tetraglycine-serine peptide linker.
[0037] General molecular biology: nucleic acid sequence, expression The term "gene" refers to a polynucleotide containing at least one open reading frame (ORF) that is capable of encoding a particular polypeptide or protein after being transcribed and translated. A polynucleotide sequence can be used to identify larger fragments or the full-length coding sequence of the gene with which it is associated. Methods for isolating larger fragment sequences are known to those of skill in the art.
[0038] The term "transgene" as used herein refers to a gene or genetic material introduced from one organism into another. As used herein, the term may also refer to the introduction of a naturally occurring or physiologically intact variant of a gene sequence into a patient's tissue that is deficient in that gene. Furthermore, it may refer to the introduction of a native coding sequence whose expression is driven by a promoter that is absent or silenced in the target tissue.
[0039] The term "recombinant" as used herein relates to a nucleic acid that is the product of one or more cloning, restriction, and / or ligation steps and that differs from naturally occurring nucleic acid. Recombinant viral particles contain recombinant nucleic acids.
[0040] The terms "gene expression" or "expression," or alternatively, "gene product," can refer to either or both the process—and product—of producing nucleic acids (RNA) or peptides or polypeptides, also called transcription and translation, respectively, or any of the intermediate processes that regulate the processing of genetic information to result in a polypeptide product. The term "gene expression" can also apply to the transcription and processing of RNA gene products, such as regulatory RNAs or structural (e.g., ribosomal) RNAs. When the expressed polynucleotide is derived from genomic DNA, expression can include splicing of mRNA in eukaryotic cells. Expression can be assessed at both the level of transcription and translation, i.e., the mRNA and / or protein product.
[0041] The term "nucleic acid expression vector" as used herein refers to a plasmid, viral genome, or RNA that is used to transfect (in the case of a plasmid or RNA) or transduce (in the case of a viral genome) a specific gene of interest into target cells, or—in the case of a transfected RNA construct—to translate the corresponding protein of interest from the transfected mRNA. In vectors that operate at the level of transcription and subsequent translation, the gene of interest is under the control of a promoter sequence that is operable in the target cell such that the gene of interest is transcribed constitutively, in response to a stimulus, or depending on the state of the cell. In certain embodiments, a viral genome is encapsidated, resulting in a viral vector that can transduce target cells.
[0042] Binding; binders, ligands, antibodies: Unless more narrowly defined in the detailed description of the invention, references to binding agents and ligands include antibodies, antibody-like molecules and aptamers as defined in the following paragraphs.
[0043] The term "specific binding" in the present invention refers to the property of a ligand that binds to its target with a particular affinity and target specificity. The affinity of such a ligand is indicated by the dissociation constant of the ligand. A specifically reactive ligand has a dissociation constant of 10 or greater when bound to a target. -8 mol / L or less (especially 10 -9 mol / L or less), but when interacting with a molecule that has nearly the same chemical composition as the target but a different conformation, the dissociation constant is at least three orders of magnitude higher.
[0044] As used herein, an "antibody-like molecule" refers to an antibody that binds to another molecule or target with high affinity / Kd≦10 -7 mol / L (especially ≦10 -9"Antibody-like molecules" refers to molecules that can specifically bind to targets at a specific binding level (mol / L). Antibody-like molecules bind to their targets in a manner similar to the specific binding of antibodies. The term "antibody-like molecule" encompasses repeat proteins such as designed ankyrin repeat proteins (Molecular Partners, Zurich) and engineered antibody-mimetic proteins that exhibit highly specific and high-affinity target protein binding (see US2012142611, US2016250341, US2016075767, and US2015368302). The term antibody-like molecule further encompasses, but is not limited to, polypeptides derived from armadillo repeat proteins, polypeptides derived from leucine-rich repeat proteins, and polypeptides derived from tetratricopeptide repeat proteins. The term "antibody-like molecule" also encompasses specific binding polypeptides derived from protein A domains, fibronectin domain FN3, consensus fibronectin domains, lipocalins (see Skerra, Biochim. Biophys. Acta 2000, 1482(1-2):337-50), polypeptides derived from zinc finger proteins (see Kwan et al., Structure 2003, 11(7):803-813), Src homology domain 2 (SH2) or Src homology domain 3 (SH3), PDZ domains, gamma-crystallin, ubiquitin, cysteine knot polypeptides or knottins, cystatins, Sac7d, triple-helical coiled-coils (also known as alpha bodies), Kunitz domains or Kunitz-type protease inhibitors, and carbohydrate-binding module 32-2. The term "antibody-like molecule" also encompasses humanized camelid antibodies. The term "antibody-like molecule" also encompasses scFv fragments.
[0045] The term "protein A domain-derived polypeptide" refers to a molecule that is a derivative of protein A and is capable of specifically binding the Fc region and Fab region of an immunoglobulin.
[0046] The term "armadillo repeat protein" refers to a polypeptide that includes at least one armadillo repeat, which is characterized by a pair of alpha helices that form a hairpin structure.
[0047] As used herein, the term "fragment crystallizable (Fc) region" is used in the sense known in the art of cell biology and immunology; when applied to IgG, it refers to the C region covalently linked by disulfide bonds. H 2 and C H Refers to the portion of an antibody that contains two identical heavy chain fragments consisting of three domains.
[0048] As used herein, the term "single-chain variable fragment (scFv)" refers to a fusion protein of the variable regions of immunoglobulin heavy (VH) and light (VL) chains, which confer antibody-like high affinity for a target from a single polypeptide chain. The VH and VL chains of an scFv are linked via a short linker peptide of 10 to approximately 25 amino acids [Huston et al. (1988). PNAS 85(16):5879-5883]. This linker can either connect the N-terminus of VH to the C-terminus of VL (VL-VH) or in the reverse configuration (VH-VL). DETAILED DESCRIPTION OF THE INVENTION
[0049] Detailed Description of the Invention system A first aspect of the present invention relates to a system for modifying the chemical or physical properties of endogenous intracellular condensates.
[0050] A specific example of such a system is a system for modifying a target biomolecule, where the target biomolecule associates with endogenous intracellular condensates formed by intracellular proteins containing intrinsically disordered regions (IDRs).
[0051] The system comprises at least two components, an IDR tract and an effector, which are associated on a single polypeptide chain, or at least two components, each located on a separate polypeptide chain, which associate into a molecular complex either spontaneously or in response to an external stimulus, such as light irradiation or cessation of light irradiation.
[0052] The components of the system can also be present in at least a three-part system, where the adaptor polypeptide promotes binding of different ratios of IDR moieties to effector moieties.
[0053] The system can be configured with multiple IDRs and / or multiple effectors, facilitating varying the ratio of IDRs to effectors, either more IDRs per effector or more effectors per IDR.
[0054] The two minimum components of the system are: a. an IDR polypeptide sequence tract containing the intrinsically disordered region; and b. an effector polypeptide domain capable of covalently modifying said target is.
[0055] It may be that the entire IDR region of a protein associated with a condensate does not need to be contained within the IDR tract in order for the IDR tract to fit the condensate.
[0056] In certain embodiments, components of the system are expressed as transgenes within cells, for example from artificial expression plasmids encoding the components.
[0057] In certain embodiments, the target is a target biomolecule.
[0058] In certain embodiments, the system comprises an effector domain and two or more IDR sequence tracts, where multiple IDRs may be required to allow large or complex protein "payloads" to be placed into the condensate.
[0059] In certain embodiments, the system comprises one effector domain and two, three, four, five or more IDR sequence tracts.
[0060] In certain embodiments, the system comprises one effector domain and one IDR sequence tract.
[0061] In certain embodiments, the system comprises one IDR sequence tract and two or more effector domains.
[0062] Naturally occurring disorders According to the definition cited by Madan Babu (Biochem Soc Trans. 2016 Oct 15;44(5):1185-1200), an intrinsically disordered region (IDR) is a polypeptide segment that does not contain enough hydrophobic amino acids to mediate cooperative folding. Instead, it typically contains a high proportion of polar or charged amino acids [Uversky et al. Proteins: Struct., Funct., Bioinf. 41,415-427]. Thus, IDRs in the native state do not have a unique three-dimensional structure, either in whole or in part. They generally adopt various conformations that are in dynamic equilibrium under physiological conditions [Forman-Kay and Mittag (2013), Structure 21,1492-1499,32-34].
[0063] In certain embodiments, the IDR-containing protein is a protein located in the nucleus of a eukaryotic cell.
[0064] The embodiments shown in the examples illustrate experiments performed on transcription-associated condensates in eukaryotic cells, in nuclei, but the invention is not limited to application in nuclear condensates or even in eukaryotic systems.
[0065] Numerous physiologically relevant proteins have been identified that contain IDRs. For purposes of defining IDRs in the present invention, the IDRs assigned to the sequences of the proteins included in the following list are considered to be embodiments of IDRs useful in practicing the present invention. However, it is not intended that the present invention be limited to the IDRs of the proteins included in this list.
[0066] In certain embodiments, the IDR-containing protein is selected from the group consisting of the following proteins: BRD4, NELFA, NELFB, CDK9, P-TEFb, Mediator complex; RNA polymerase (RPB1; Uniprot ID P24928) is selected from the group consisting of:
[0067] In certain embodiments, the IDR-containing protein is selected from the group consisting of BRD4, NELFA, NELFB, CDK9, P-TEFb, and RPB1.
[0068] In certain more particular embodiments, the IDR-containing protein is more preferably selected from BRD4 and NELFA;
[0069] In a particular embodiment, the IDR tract is derived from the IDR of BRD4 (Uniprot ID O60885).
[0070] In certain embodiments, the IDR sequence tract is selected from SEQ ID NO: 001 and SEQ ID NO: 002 (ΔN-BDR4 IDR; NELFA IDR), or a sequence variant thereof characterized by at least 85% sequence identity with any of the sequences described above, said variant having a biological function of promoting the association of the system according to the invention with endogenous condensates.
[0071] In certain embodiments, the IDR tract is derived from the IDR of NELFA (Negative Elongation Factor Complex Member A; Uniprot ID Q9H3P2).
[0072] In a particular embodiment, the IDR tract is derived from the IDR of NELFB (negative elongation factor complex member B; Uniprot ID Q8WX92).
[0073] In a particular embodiment, the IDR tract is derived from the IDR of CDK9 (cyclin-dependent kinase 9; Uniprot ID P50750).
[0074] In a particular embodiment, the IDR tract is derived from the IDR of cyclin T1 (CCNT1, Uniprot O60563).
[0075] In a particular embodiment, the IDR tract is derived from the IDRs of human MED1 CTD (Uniprot ID Q15648).
[0076] In a particular embodiment, the IDR tract is derived from the IDR of the C-terminal domain (CTD) of the Pol II catalytic subunit RPB1 (RPB1 IDR; Uniprot ID P24928).
[0077] In certain embodiments, the IDR-containing protein is selected from the following list: [Table 1] JPEG2026500369000002.jpg218159JPEG2026500369000003.jpg136159
[0078] Those skilled in the art will be able to identify additional proteins of interest for practicing the present invention from the following reviews, which are incorporated herein by reference: - Tong et al.,Liquid-liquid phase separation in tumor biology;Nature Signal Transduction and Targeted Therapy 7,Art.No 221(2022); - Boja et al.2021;Biomolecular Condensates and Cancer,Cancer Cell 39,174-192; - Cai et al.2021;Biomolecular Condensates and Their Links to Cancer Progression;Trends in Biochemical Sciences 46,535-549; - Zbinden et al.2020;Phase Separation and Neurodegenerative Diseases:A Disturbance in the Force,Developmental Cell 55 45-68
[0079] Predicting IDR: Based on protein sequence information, disorders can be predicted. We used PONDR (Predictor of Natural Disordered Regions) software, available at pondr.com, and the derived VSL2 score (Peng et al., (2006) BMC Bioinformatics 7:208, incorporated herein by reference) as the output determinant. A sequence tract of 50 or more consecutive amino acid (AA) positions with a VSL2 score of (≥) 0.5 or more is considered to be an IDR tract suitable for implementing the present invention.
[0080] In certain embodiments, an IDR tract is a stretch of 50 amino acids (AA) characterized by a VSL2 score of 0.7 or greater.
[0081] In certain more particular embodiments, the IDR tract is a stretch of 50 AA characterized by a VSL2 score of 0.8 or greater.
[0082] In certain embodiments, an IDR tract is a stretch of 100 AA characterized by a VSL2 score of 0.5 or greater.
[0083] In certain embodiments, an IDR tract is a stretch of 100 AA characterized by a VSL2 score of 0.7 or greater.
[0084] In a particularly more particular embodiment, the IDR tract is a stretch of 100 AA characterized by a VSL2 score of 0.8 or greater.
[0085] In an even more particular embodiment, the IDR tract is a stretch of 50 AA characterized by a VSL2 score of 0.9 or greater.
[0086] In the embodiment shown in the examples, the "native" IDR and the probe (FP1) IDR overlap 100%. Without wishing to be bound by theory, the inventors assume that IDR tracts with similar "chemical grammar" will be assembled into the same condensate. The inventors have not determined in detail what percentage of overlap between the "native" IDR and the sequence tract of the truncated IDR is required to allow the sequence tract of the truncated IDR to be incorporated into the condensate.
[0087] In certain embodiments, the IDR sequence tract comprises at least 80% (or more) of the IDRs of a naturally occurring condensate-forming (intrinsically disordered) protein. In certain embodiments, the IDR sequence tract comprises 85% or more of the IDRs of a naturally occurring condensate-forming protein. In more particular embodiments, this percentage is 90% or more. In certain even more particular embodiments, this percentage is 95% or more, or even 98% or more.
[0088] target In general, the systems of the invention target condensates in the sense that they facilitate manipulation of condensates by allowing condensate-specific addition of polypeptides, either by controlled expression of the systems of the invention or by stimulus-dependent association of the IDRs of the systems with effector moieties, whose effect may consist simply of perturbing the chemical or physical balance within the condensate, thereby resulting in a change in its physiological behavior.
[0089] In practical terms, a straightforward application exemplified by the examples contained herein is the introduction of enzyme effectors that allow for the covalent modification of proteins associated with condensates (either actively forming condensates, or contained in condensates, or functionally associated with condensates).
[0090] In certain embodiments, the target biomolecule is a protein. In certain embodiments, the target biomolecule is an intracellular protein. In more specific embodiments, the target biomolecule is an intracellular nuclear protein. In even more specific embodiments, the target biomolecule is a nuclear protein associated with RNA polymerase II activity.
[0091] Although the target type addressed in the examples is a protein, targets for manipulation or modification are not limited to proteins and include DNA and RNA present within biomolecular condensates. As a non-limiting example, when DamID is used as an effector modality, it modifies (methylates) DNA.
[0092] effector In parallel with what was said about targets in the most general concept in the previous paragraph, effector polypeptides can be considered as functional entities that help to disrupt the chemical or physical balance within the condensate, thereby altering its physiological behavior.
[0093] In certain embodiments, the effector is functionalized to facilitate the covalent modification of biomolecules associated with the condensates by inducing the formation or cleavage of covalent bonds.
[0094] In certain embodiments, the effector polypeptide is a polypeptide capable of biotinylating a target, in particular a target protein.
[0095] In certain embodiments, the effector polypeptide is a polypeptide capable of ubiquitinating a target, in particular a target protein.
[0096] In certain embodiments, the effector polypeptide is a polypeptide capable of methylating a target, in particular a target protein.
[0097] In certain embodiments, the effector polypeptide is a polypeptide capable of demethylating a target, in particular a target protein.
[0098] In certain embodiments, the effector polypeptide is a polypeptide capable of acetylating a target, in particular a target protein.
[0099] In certain embodiments, the effector polypeptide is a polypeptide capable of deacetylating a target, in particular a target protein.
[0100] In certain embodiments, the effector polypeptide is a polypeptide capable of phosphorylating a target, in particular a target protein.
[0101] In certain embodiments, the effector polypeptide is a polypeptide capable of dephosphorylating a target, in particular a target protein.
[0102] In certain embodiments, the effector polypeptide is a biotin activator characterized by SEQ ID NO: 007, or a variant thereof having at least 50% of its activity. The activity of this construct is disclosed in Branon et al., (2018) Nat. Biotechnol., 36(9):880-887, which is incorporated herein by reference.
[0103] Various enzymatic activities can be used to modify target moieties in endogenous condensates. We developed the so-called BioID proximity-based labeling system, which utilizes an inducible E. coli biotin ligase (BirA), characterized by a catalytic site mutation (R118G) that destabilizes the retention of an activated biotin molecule (biotinoyl-5'-AMP). * ) was employed. This activated biotin molecule dissociates from the ligase and reacts with the primary amine of an exposed lysine residue on a neighboring protein, resulting in the covalent attachment of biotin to the target (see Trinkle-Mulcahy, F1000 Research 2019, 8(F1000 Faculty Rev):135, and references therein). Variants of this system are commercially available under the names TurboID, split TurboID, BioID2, BASU, and APEX. Another activity for modifying / tagging proteins is ubiquitin ligase.
[0104] Another activity for modifying DNA is DNA adenine methyltransferase, commercially available as DamID.
[0105] One advantage of the method of the present invention and the components for using the method is its versatility. Biotinylation of target structures is just one possible application. Different IDRs can target different types of endogenous condensates, and different "cargos," i.e., modifying effectors, can modify condensates differently.
[0106] Separation of IDRs and effectors into distinct polypeptide constructs In certain embodiments, the IDR sequence tract and the effector domain are located on separate polypeptide molecules that can be induced to associate in response to a stimulus. In certain embodiments, the stimulus causes the association of separate polypeptide molecules that each contain the IDR sequence tract and the effector domain. These separate polypeptide molecules are present before the stimulus causes the association.
[0107] In certain embodiments, the stimulus is light.
[0108] In certain embodiments, the association of a first fusion polypeptide containing an IDR tract with another fusion peptide containing an effector domain is facilitated by a light-inducible binding partner pair. The light-inducible binding partner pair can consist of a first binding partner and a second binding partner, which associate (or dissociate) in the presence of light. To enable efficient control of the production of the system of the present invention, one of the binding partners is a portion of the fusion polypeptide containing an IDR, and the other binding partner is associated directly or indirectly with the effector domain.
[0109] In certain embodiments, the light-inducible binding partner pair is selected from the group consisting of: - SspB-iLID, ΔPhyA-FHY-1, and ΔPhyA-FHL (association: 660 nm / dissociation: dark or 740 nm; Zhou et al., Nature Biotechnology 40, 262-272 (2022); - PhyB / PIF3 and PhyB / PIF6 (association: 660 nm / dissociation: in the dark or 740 nm; Toettcher et al., Nature Methods 8, 837-839 (2011)); - UVR8 / COP1(300nm;Crefcoeur et al.,Nature Communications 4,Article:1779(2013)); - CRY2 / CIB1(450nm;Konermann et al.Nature 500,472-476(2013)); - FKF1 / GI (450 nm; Yazawa et al., Nature Biotechnology 27, 941-945 (2009)); - A VVD variant of Neurospora crassa called Magnet (450 nm; Kawano et al. Nature Communications 6, Article: 6256 (2015)); - AsLOV2-ePDZ("Tulip";Strickland et al.,Nature Methods 9,379-384(2012)); - cpLOV2-SspB(He et al.,Nature Chemical Biology 17,915-923(2021)); - BphP1 / PpsR2 (association: 760 nm / dissociation: dark or 640 nm; Kaberniuk et al., Nature Methods 13, 591-597 (2016)); - BphP1 / Q-PAS1 (association: 760 nm / dissociation: dark or 640 nm, Redchuk et al., Nature Chemical Biology 13, 633-639 (2017)); - MagRed (DrBphP / Aff6_V17FΔN, association: 660 nm / dissociation: in the dark or 780 nm, Kuwasaki et al., Nature Biotechnology 40, 1672-1679 (2022))
[0110] The use of "optogenetic light switches" that allow protein units to associate as a result of irradiation with light of a specific wavelength is well established in the art, as evidenced by the many examples of such light-induced binding partner pairs described above. Furthermore, binding pairs that undergo dissociation upon exposure to light are also known, as reported, for example, by Karapinar et al., Nature Communications 12, Article: 4488 (2021). Thus, while the examples included herein focus on the establishment of association as a result of light, the reverse mechanism can also be used.
[0111] In one particular embodiment, the light-inducible binding partner pair is SspB and iLID (SEQ ID NOs: 003 and 004), derived from the light-oxygen-voltage 2 (LOV2) domain from oat (Avena sativa) (Guntas et al. (2015) PNAS 112, 112-117). This is the system used in the Examples.
[0112] A three-component system, referred to herein as LITEC In one particular embodiment, the system according to the invention comprises: a. a plurality of first fusion polypeptides comprising an IDR sequence tract containing an intrinsically disordered region (at least 80%), optionally a first fluorescent marker polypeptide, and a first member of a light-inducible binding partner pair; b. a plurality of second fusion polypeptides comprising a second member of a light-inducible binding partner pair, optionally a second fluorescent marker polypeptide, and a binding domain capable of specifically binding to a non-endogenous peptide epitope; c. A third fusion polypeptide comprising a plurality of said non-endogenous peptide epitopes and an effector polypeptide domain capable of covalently modifying said target.
[0113] First fusion polypeptide The first fusion polypeptide (FP1) is a distinct sequence from the second and third fusion polypeptides, and it interacts with the second fusion polypeptide. Association and incorporation into endogenous condensates is mediated by an IDR sequence tract, which is also a "natural" part of the condensate, contained in its constituent proteins.
[0114] FP1 contains an IDR tract and one partner of a light-inducible binding pair. In an embodiment, the IDR is located at the N-terminal portion of the FP1 polypeptide, while the binding partner is located at the C-terminal portion. In some embodiments, this polarity can be reversed, i.e., the IDR can be at the C-terminus and the binding partner can be at the N-terminus.
[0115] The examples show FP1 constructs that include a first fluorescent marker polypeptide, specifically the "mCherry" red fluorescent protein derived from the DsRed protein of Discosoma. Other fluorescent proteins may also be used. The fluorescence or fluorescent marker polypeptide is an optional component that facilitates visualization of the assembly and condensate incorporation process, but is not, in the inventors' view, essential to the practice of the invention.
[0116] In certain embodiments, the first fusion polypeptide does not include a fluorescent marker polypeptide.
[0117] In certain other specific embodiments, the first and / or second polypeptide comprises a fluorescent marker polypeptide selected from mCherry and eGFP.
[0118] In certain embodiments, the first fusion polypeptide is characterized by SEQ ID NO: 008. In certain embodiments, the first fusion polypeptide is characterized by SEQ ID NO: 009. In certain embodiments, the first fusion polypeptide is characterized by SEQ ID NO: 010. In certain embodiments, the first fusion polypeptide is characterized by SEQ ID NO: 011.
[0119] Second Fusion Polypeptide The second fusion polypeptide contains a second binding partner of a light-inducible binding pair, the first partner of which is part of the first fusion polypeptide.
[0120] The second fusion polypeptide (FP2) also contains a binding domain capable of specifically binding to a non-endogenous peptide epitope. This property of binding to a non-physiological epitope allows FP2 to be specifically assembled onto a set of epitopes introduced into cells without being perturbed by endogenous binding partners and without interfering with any physiological functions of these binding partners.
[0121] In certain embodiments, the binding domain capable of specifically binding to a non-endogenous peptide epitope is an scFv (single chain variable) antibody fragment.
[0122] A particular example of such a binding domain is an scFv fragment as used in the examples, but other antibody-like constructs as known in the art can also be used.
[0123] The second binding partner of the light-inducible binding partner pair and the binding domain capable of specifically binding to a non-endogenous peptide epitope are advantageously located at either end of the FP2 construct. In one embodiment, the second binding partner of the light-inducible binding partner pair is at the N-terminus and the binding domain capable of specifically binding to a non-endogenous peptide epitope is at the C-terminus. This order can also be reversed.
[0124] In certain embodiments, the non-endogenous peptide epitope is SEQ ID NO:005.
[0125] In certain embodiments, the non-endogenous peptide epitope is included as multiple copies in the construct exemplified by SEQ ID NO:006.
[0126] The examples demonstrate FP2 constructs that include a second fluorescent marker polypeptide, specifically a green fluorescent protein derived from Aqueora victoria. Other fluorescent proteins may also be used. As noted above, the fluorescence or fluorescent marker polypeptide included on the first and / or second constructs is an optional component that facilitates visualization of the assembly and condensate incorporation process, but is not, in the inventors' view, required to practice the invention.
[0127] In certain embodiments, the second fusion polypeptide does not include a fluorescent marker polypeptide.
[0128] In certain embodiments, the second fusion polypeptide is characterized by SEQ ID NO: 012 (sequence of scFv-GFP-iLID and, where feasible, scFv-iLID).
[0129] Third fusion polypeptide The third fusion polypeptide, FP3, contains an enzymatic function for modifying target moieties in the condensate and multiple attachment / binding sites ("epitopes") for the binding domain contained in the second fusion polypeptide.
[0130] The attachment site for the binding domain of the FP2 construct acts as a multiplier of the signal transmitted by the IDR region, promoting the integration of large entities into the condensate.
[0131] In addition to the scFv / SunTag system employed in the examples, many tagging systems exist, many of which have been developed to facilitate affinity purification or signal amplification of fluorescent proteins, including, but not limited to, MoonTag, FLAG tag, TAP tag, etc.
[0132] In certain embodiments, the FP3 construct comprises 3-40 copies of the epitope, particularly 4-30 copies, more particularly 10-20 copies of the same peptide epitope 6-15 amino acids in length. In certain embodiments, to avoid cross-reactivity of the epitope binding site with cellular structures, the epitope sequence is selected so that it is not part of the expressed genome of the cell in which the condensate is present (i.e., the cell in which the system is used).
[0133] In certain embodiments, the third fusion polypeptide is characterized by SEQ ID NO: 13 (sequence of bioID-SunTag).
[0134] Nucleic acids and vectors Another aspect of the invention relates to one or more nucleic acid sequences encoding a system as defined in any of the above aspects of the invention. Embodiments characterizing distinct properties of the system may be encoded by the nucleic acid sequences.
[0135] One particular embodiment of this aspect relates to a nucleic acid sequence encoding a third fusion polypeptide as described above.
[0136] The nucleic acid according to this embodiment may be contained within an expression vector, such as a DNA plasmid or viral vector. Alternatively, the nucleic acid may be one or more mRNA molecules encoding the system. Portions of the system may be encoded on a plasmid vector, for example, to facilitate the insertion of different effector components or IDRs that facilitate fast and easy insertion of protein components employed as part of the system.
[0137] method Another aspect of the present invention relates to methods for modifying or manipulating the chemical or physical properties of endogenous intracellular condensates.
[0138] In certain embodiments of this aspect, the method is employed to modify a target, particularly a target protein, associated with endogenous intracellular condensates formed by intracellular proteins containing intrinsically disordered regions. The method comprises the steps of: a) providing a cell containing an intracellular protein that includes an intrinsically disordered region; b) expressing in said cell a system as defined in any of the above aspects of the invention.
[0139] In certain embodiments, the method further comprises the steps of: c) recovering or isolating a preparation containing the target biomolecule from said cells; d) isolating the target biomolecule modified by said effector polypeptide capable of covalently modifying the target biomolecule.
[0140] In certain embodiments, the effector polypeptide is capable of biotinylated a target biomolecule, and isolation of the biotinylated target biomolecule is achieved by binding the biotinylated protein to a matrix from a preparation containing the biomolecule.
[0141] The isolated target biomolecules can then be identified, and one particular identification method that has proven useful for this purpose is mass spectrometry.
[0142] In a more particular embodiment, the target biomolecule is a protein.
[0143] For example, when alternative forms of a single separable feature, such as an IDR tract, a light-inducible binding partner pair, or an effector polypeptide, are described herein as "embodiments," it is understood that such alternative forms can be freely combined to form individual embodiments of the invention disclosed herein. Thus, any of the alternative embodiments of an IDR tract can be combined with any of the alternative embodiments of an effector polypeptide, and these combinations can be combined with any of the fluorescent markers described herein.
[0144] The present invention further encompasses the following: Item 1. A system for modifying a target, comprising: the system, wherein the target is associated with endogenous intracellular condensates formed by intracellular proteins containing intrinsically disordered regions (IDRs); or 1. A system for modifying the chemical or physical properties of endogenous intracellular condensates, comprising: The system comprises: a. an IDR sequence tract containing the intrinsically disordered region; and b. an effector domain capable of modifying said target The system comprising:
[0145] Item 2. The system of item 1, wherein the target is modified by a covalent modification.
[0146] Item 3. The system according to Item 1 or 2, wherein the IDR-containing protein is a protein located in the nucleus of a eukaryotic cell.
[0147] Item 4. IDR-containing proteins include the following: BRD4, NELFA, NELFB, CDK9, P-TEFb, the Mediator complex, and RNA polymerase (RPB1). is selected from the group consisting of In particular, wherein the IDR-containing protein is selected from the group consisting of BRD4, NELFA, NELFB, CDK9, P-TEFb, and RPB1; More particularly, the IDR-containing protein is BRD4 or NELFA; Even more particularly, the IDR-containing protein is BRD4. The system according to any one of items 1 to 3.
[0148] Item 5. The system according to any one of Items 1 to 4, wherein the IDR sequence tract is selected from SEQ ID NO: 001 and SEQ ID NO: 002 (ΔN-BDR4 IDR; NELFA IDR), or a sequence variant thereof characterized by having at least 85% sequence identity with either SEQ ID NO: 001 or SEQ ID NO: 002.
[0149] Item 6. The system includes one effector domain and two or more IDR sequence tracts, In particular, the system includes two, three, four, or five or more IDR tracts; The system according to any one of items 1 to 5.
[0150] Item 7. The system according to any one of Items 1 to 5, wherein the system comprises one effector domain and one IDR sequence tract.
[0151] Item 8. The system according to any one of Items 1 to 5, wherein the system comprises one IDR sequence tract and two or more effector domains.
[0152] Item 9. The system of any one of Items 1 to 8, wherein the target is a protein, particularly an intracellular protein, more particularly an intracellular nuclear protein, and even more particularly a nuclear protein associated with RNA polymerase II activity.
[0153] Item 10. The effector polypeptide capable of covalently modifying the target comprises: a. Biotinylation, b. Ubiquitination c. methylating; d. demethylating; e. acetylating; f. deacetylating; g. phosphorylating; h. Dephosphorylate 10. The system according to any one of items 1 to 9, wherein the polypeptide is selected from polypeptides capable of:
[0154] Item 11. The system of any one of Items 1 to 10, wherein the IDR sequence tract and the effector domain are located on separate polypeptide molecules that can be induced to associate in response to a stimulus.
[0155] Item 12. The system according to Item 11, wherein the stimulus is light.
[0156] Item 13. The association of a first fusion polypeptide containing an IDR tract with another fusion peptide containing an effector domain is promoted by a light-inducible binding partner pair; A light-inducible binding partner pair consists of a first binding partner and a second binding partner, where the first binding partner and the second binding partner associate in the presence of light; and one of the binding partners is part of the first fusion polypeptide, and the other of the binding partners is associated with an effector domain; Item 13. The system according to item 12.
[0157] Item 14. The light-inducible binding partner pair is: - SspB-iLID, ΔPhyA-FHY-1 and ΔPhyA-FHL; - PhyB / PIF3 and PhyB / PIF6; - UVR8 / COP1;CRY2 / CIB1; - FKF1 / GI; - VVD variant of Neurospora crassa; - AsLOV2-ePDZ; - cpLOV2-SspB; - BphP1 / PpsR2; - BphP1 / Q-PAS1; - MagRed Item 14. The system of item 13, wherein the system is selected from the group consisting of:
[0158] Item 15. The system according to Item 13 or 14, wherein the light-inducible binding partner pair is SspB and iLID (sequence numbers 003 and 004).
[0159] Item 16. The system: a. a first fusion polypeptide comprising an IDR sequence tract containing the intrinsically disordered region, [optionally a first fluorescent marker polypeptide], and a first member of a light-inducible binding partner pair; b. a second fusion polypeptide comprising a second member of said light-inducible binding partner pair, [optionally a second fluorescent marker polypeptide] and a binding domain capable of specifically binding to a non-endogenous peptide epitope; c. a third fusion polypeptide comprising a plurality of said non-endogenous peptide epitopes and an effector domain capable of covalently modifying said target. 16. The system according to any one of items 1 to 15, comprising:
[0160] Item 17. The system according to Item 14, wherein the binding domain capable of specifically binding to a non-endogenous peptide epitope is an scFv (single-chain variable) antibody fragment.
[0161] Item 18. The system according to Item 16 or 17, wherein the non-endogenous peptide epitope is SEQ ID NO: 005.
[0162] Item 19. The system according to any one of Items 16 to 18, wherein the first fusion polypeptide does not contain a fluorescent marker polypeptide.
[0163] Item 20. The system according to any one of Items 16 to 19, wherein the second fusion polypeptide does not contain a fluorescent marker polypeptide.
[0164] Item 21. The system described in any one of Items 16 to 18, wherein the first fusion polypeptide is characterized by a sequence selected from SEQ ID NOs: 008, 009, 010, and 011.
[0165] Item 22. The system described in any one of Items 16 to 18, wherein the second fusion polypeptide is characterized by SEQ ID NO: 012 (sequence of scFv-GFP-iLID and, if applicable, scFv-iLID).
[0166] Item 23. The system described in any one of Items 16 to 22, wherein the third fusion polypeptide is characterized by SEQ ID NO: 13 (bioID-SunTag sequence).
[0167] Item 24. One or more nucleic acid sequences encoding a system as defined in any one of items 1 to 22.
[0168] Item 25. A nucleic acid sequence encoding a third fusion polypeptide as defined in Item 23.
[0169] Item 26. An expression vector encoding the system or polypeptide component of the system according to any one of Items 1 to 23, The expression vector is particularly selected from the group of plasmids, RNA expression vectors and viruses.
[0170] Item 27. A method for modifying a target (particularly a target protein) associated with endogenous intracellular condensates formed by an intracellular protein containing an intrinsically disordered region, the method comprising the steps of: a) providing a cell containing an intracellular protein that includes an intrinsically disordered region; b) expressing in said cells a system as defined in any one of items 1 to 23. The method comprising:
[0171] Item 28. The method comprises the following steps: c) recovering a preparation containing target biomolecules (particularly proteins) from said cells; d) isolating the target biomolecule (particularly a protein) modified by said effector polypeptide capable of modifying the target biomolecule. 28. The method of claim 27, further comprising:
[0172] Item 29. The method of Item 27 or 28, wherein the effector polypeptide is capable of biotinylating a target biomolecule, and isolation of the biotinylated target biomolecule is achieved by binding the biotinylated protein from a preparation containing the biomolecule to a matrix.
[0173] Item 30. The method according to Item 28 or 29, wherein the isolated target biomolecule is identified, particularly by mass spectrometry.
[0174] The present invention is further explained by the following examples and figures, from which further embodiments and advantages can be derived, which are intended to illustrate the invention without limiting its scope. [Brief explanation of the drawings]
[0175] [Figure 1]Intrinsically disordered regions (IDRs) target endogenous transcriptional condensates. (A and B) Using the PONDR VSL2 algorithm, the intrinsically disordered regions of BRD4 and NELFA are shown as thick black bars. The Y-axis represents the VSL2 score, and the X-axis represents the amino acid (aa) position. For BRD4, the IDR portion of BRD4 (aa 462 to aa 1362) is used. For NELFA, all amino acids of NELFA are used. (C-F) The BRD4 IDR colocalizes with pre-existing endogenous Mediator condensates. (C) Endogenous Mediator complex subunit 19 is labeled with HaloTag and shows transcriptional condensates (cyan). (D) The BRD4 IDR, fused to mCherry for visualization, shows clusters of IDRs in the nucleus (red). (E) Overlay of the BRD4 Mediator and IDRs confirms colocalization of Mediator condensates with IDR clusters. (F) Intensity profile through the orange arrowhead in panel E. Compared to background levels, the BRD4 IDR at the condensation site accumulates more than Med19. (H-K) NELFA colocalizes with pre-existing endogenous Pol II condensates. (G) Pol II forms endogenous condensates (cyan). (H) NELFA fused to mCherry forms NELFA clusters in the nucleus (red). (I) Overlay of Pol II and NELFA confirms the colocalization of Pol II condensates and NELFA clusters. (J) Intensity profile through the orange arrowhead in panel I. Panels B and C were generated from pondr.com. [Figure 2]Colocalized IDRs are highly mobile. (A) A representative cell from a FRAP (Fluorescence Recovery after Photobleaching) experiment. At 0 seconds, one condensate is bleached (yellow box), while another is not (blue box). (B) The average intensity profile over time after bleaching shows that the BRD4 IDRs in the condensate are highly mobile. The recovery rate of the BRD4 IDRs is 90%, with a half-recovery time of approximately 3.4 seconds. The inset shows typical photobleaching of the BRD4 IDRs in the condensate over time. [Figure 3] (A) iLID (an improved light-inducible dimer) and its partner sspB reversibly bind under blue light. The dissociation constants of iLID and sspB are 130 nM under blue light and 4700 nM in the dark. iLID and sspB colocalize within seconds under blue light and dissociate within minutes in the dark. (B) Multivalent interaction between scFv (single-chain variable fragment) and SunTag, which has 24 binding sites for scFv. This is typically used to improve the signal-to-noise ratio in fluorescence imaging. The dissociation constant between scFv and SunTag is 40 pM. (C) A multivalent optogenetic system using three components. IDR-mCh is fused to sspB for interaction with iLID under blue light (IDR-mCh-sspB). Due to the IDR moiety, this component remains in endogenous transcription condensates, as shown in Figure 1D and H. The scFv and iLID are fused to sfGFP (scFv-sfGFP-iLID). 24×SunTag is employed to efficiently target specific proteins into condensates and better visualize the binding between iLID and sspB under blue light. 24×SunTag and scFv-sfGFP-iLID form a complex protein containing up to 24 iLIDs. Under blue light, this complex begins to interact with IDR-mCh-sspB in the nucleus and condensates. We hypothesize that the multiple iLIDs allow the complex to find condensates faster and have a higher binding affinity to them. [Figure 4]The Light-induced targeting of endogenous condensate (LiTEC) system can accumulate GFP within condensates using blue light. (A-F) Cells expressing the constructs (IDR-mCh-sspB, scFv-sfGFP-iLID, and 24xSunTag) demonstrate light-induced targeting and accumulation of GFP within endogenous condensates. (A) The IDR of BRD4 is present within endogenous condensates. Three representative condensates are shown in yellow boxes. (B) GFP is uniformly distributed before blue light is turned on. (C-E) GFP accumulation is detected at the endogenous condensate site 5, 10, and 30 seconds after blue light irradiation. The longer the blue light is turned on, the brighter the GFP accumulation within the yellow box becomes. (F) A time-dependent profile of GFP intensity shows that GFP begins to accumulate in response to blue light. Intensity is normalized to the background intensity in the nucleus. The blue light was turned on at time t = 0. The average intensity of the three condensate positions (yellow boxes) gradually increases with blue light, while the intensity of the reference position decreases with photobleaching. (G) When IDR is not expressed in cells, scFv-sfGFP-iLID is uniformly distributed from the nucleus. (H) Proposed structure for cargo transport into condensates. The cargo must be a protein selected by the inventors, which can provide information about the properties of condensates by manipulating their function or composition. [Figure 5]LiTEC enhances IDR accumulation in endogenous condensates under blue light (BL). (A) Representative image of BRD4 IDR before BL. (B) Image of BRD4 IDR after 50 seconds of BL. IDR clusters become larger and brighter than clusters before BL. (C) Intensity profile of BRD4 IDR through the orange arrow shown in panel B. mCh intensity in condensates increases upon BL. I0s is the peak mCh intensity of the IDR cluster before BL. I50s is the peak mCh intensity of the IDR after 50 seconds of BL. IB is the background level and is set to 1. FCBL50s is defined as the ratio of I50s-IB to I0s-IB. The mCh signal is normalized by the average mCh intensity in the nucleus. (D) Image of scFv-sfGFP-iLID after 50 seconds of BL. One representative cluster is highlighted in a yellow box. (E) GFP intensity within the yellow box measured over time under BL. The GFP signal is normalized by the average GFP intensity of the nucleus at each time point. The black dots represent the GFP signal measured over time, and the red line represents the one-phase association curve fitted to the black dots. (F) The difference in relative GFP intensity and the difference in relative mCh intensity (I30s-I0s) show a positive correlation (n=218). This indicates that the accumulation of GFP during BL recruits the outer IDRs into the condensate. (G) Schematic diagram of the effect of the increase in IDRs during BL on the compositional change. [Figure 6]Endogenous BRD4 is unaffected by 70 seconds of blue light (BL)-induced enhancement of BRD4 IDR. (A) Procedure for a Blue-On experiment. The BRD4 IDR is first imaged for 5 seconds, and the mCh intensity before BL is measured ("Before BL"). GFP and Halo-BRD4 are then simultaneously imaged for 70 seconds using a 488 nm + 642 nm laser and dual camera. Finally, the BRD4 IDR is imaged for 5 seconds, and the mCh intensity after BL is measured ("After BL"). (B) At the marked condensate (orange arrow), the fold-change mCh,BL70s is 1.70, and the fold-change BRD4,BL70s is 0.77. The relative BRD4 intensity over time, the relative GFP intensity over time, and the relative mCh intensity profile around the condensate are shown. (C) Procedure for a Blue-Off experiment. The IDR of BRD4 was measured for 5 seconds ("Before"), followed by 70 seconds of measurement of Halo-BRD4 using a 642 nm laser. Blue light was not used. The IDR of BRD4 was then imaged again for 5 seconds ("After"). (D) At the marked condensate (orange arrow), the Fold-Change mCh, NoBL70s was 0.93, while the Fold-Change BRD4, NoBL70s was 0.78. The relative intensity of BRD4 over time and the relative mCh intensity profile around the condensate are shown. (E) Statistical analysis of the relative mCh intensity for each condition (Blue On: n = 50, Blue Off: n = 39). For Blue On, the median difference between "After BL" and "Before BL" was 0.251, and the mean was 0.288. For Blue Off, there was no statistically significant difference. (F) Statistical analysis of the Fold-Change 70s of BRD4 for Blue On and Blue Off experiments. Generally, the relative intensity of BRD4 decreases over time. Comparing the degree of decrease in each condition allows us to predict the effect of IDR enhancement. Here, BRD4 does not show a statistically significant difference between "Blue On" and "Blue Off" (median difference = 0.00234, mean difference = 0.041; Blue On median, mean = 0.8109, 0.8375 (±0.0176, SEM); Blue Off median, mean = 0.8085, 0.7965 (±0.0128, SEM)).This suggests that blue light-induced enhancement of IDR may not affect endogenous BRD4 in transcriptional condensates.For all conditions in this experiment, three-dimensional z-stack images were taken and projected onto a two-dimensional plane. [Figure 7]Enhancement of BRD4 IDR by 70 seconds of blue light (BL) slightly increases endogenous MED. In this experiment, the same procedure as in the previous figure was performed, except that endogenous MED was measured instead of endogenous BRD4. (A) Endogenous MED was measured under blue light for 70 seconds. BRD4 IDR was imaged for 5 seconds before and after BL ("Before" and "After"). (B) At the marked condensate (orange arrow), Fold-Change mCh,BL70s is 1.62, and Fold-Change MED,BL70s is 0.68. The relative intensity of BRD4 over time, the relative intensity of GFP over time, and the relative intensity profile of mCh around the condensate are shown. (C) Endogenous MED was measured for 70 seconds without blue light. BRD4 IDR was imaged for 5 seconds before and after this measurement ("Before" and "After"). (D) For the marked condensate (orange arrow), the Fold-Change mCh, NoBL70s is 1.04, while the Fold-Change MED, NoBL70s is 0.77. The relative intensity of BRD4 over time and the relative mCh intensity profile around the condensate are shown. (E) Statistical analysis of the relative mCh intensity for each condition (Blue On: n = 76, Blue Off: n = 55). For Blue On, the median difference between the "After" and "Before" conditions is 0.269, and the mean is 0.315. For Blue Off, there is no statistically significant difference. (F) Statistical analysis of the Fold-Change 70s MED for the Blue On and Blue Off experiments. In general, the relative intensity of MED decreases over time. Comparing the degree of decrease for each condition allows us to predict the effect of IDR enhancement. Here, we show a small (approximately 5%) but statistically significant difference in MED between the "Blue On" and "Blue Off" conditions. (Median difference = 0.0439, mean difference = 0.0498; Blue ON median, mean = 0.8215, 0.8189 (±0.0096, SEM); Blue OFF median, mean = 0.7717, 0.7750 (±0.0141, SEM). This suggests that enhancement of IDR by blue light may recruit small amounts of endogenous MED into transcriptional condensates. For all conditions in this experiment, three-dimensional z-stack images were taken and projected onto a two-dimensional plane. [Figure 8]A schematic diagram of LiTEC-based BioID is shown. (A) The principle of an exemplary, non-limiting embodiment of the system of the present invention (also referred to herein as LiTEC) has three components: scFv-sfGPF-iLID for multivalency and optogenetics, IDR-mCh-sspB for condensate targeting and optogenetics, and BioID-24×SunTag for biotinylation and multivalency. BioID-24×SunTag binds to 24 iLIDs, thus forming a large protein complex (the "BioID complex" or "LiTEC cargo"). The BioID complex is uniformly distributed in the dark but begins to accumulate in condensates upon blue light illumination. (B) A diagram of the expected LiTEC process. All three components are expressed intracellularly. BioID-24×SunTag and scFv-sfGFP-iLID form the BioID complex, and IDR-mCh-sspB accumulates in transcriptional condensates. When blue light is turned on, an optogenetic interaction between sspB and iLID is initiated. This interaction pulls the BioID complex into the condensate. When the blue light is turned off, the optogenetic interaction disappears. The bound BioID complex is now free to diffuse out of the condensate. [Figure 9]Doxycycline-inducible BioID-SunTag for BioID. (A) (Left) Cell line used for BioID experiments. miniTurboID is fused to 24×SunTag (BioID-SunTag). (Right) To reduce background levels of biotinylation in the nucleus, BioID-SunTag expression is controlled by a doxycycline-inducible system. (B) GFP accumulation after 30 seconds of blue light irradiation is dependent on doxycycline. (Top) (Left) Example of fitting GFP accumulation over time using a monophasic binding curve. Blue light is on for 30 seconds. (Center) Fitted line of GFP accumulation from a No-SunTag cell line. (Right) Fitted line of GFP accumulation from a Dox-inducible SunTag cell line without Dox (Dox 0). (Bottom) Linear fits of GFP accumulation from the Dox-inducible SunTag cell line treated with Dox at 1 μg / m (left), 2 μg / m (center), and 4 μg / m (right) for 48 hours, respectively. This indirectly demonstrates the successful induction of BioID-SunTag by the addition of doxycycline. (C) Mean differences in GFP relative intensity between before and after 30 seconds of blue light exposure in each condition. From left to right, the means are 0.206 (±0.144, SD, n=39), 0.315 (±0.158, SD, n=39), and 2.382 (±1.135, SD, n=43). [Figure 10]Biotinylated proteins specifically accumulate in transcription condensates under blue light. (A-D) Diagrams of the expected IDR clusters, GFP accumulation, and protein biotinylation under various conditions. (A) In the "All" condition, blue light causes GFP to accumulate within transcription condensates (marked by the BRD4 IDR), resulting in BioID biotinylation of the condensates. Biotinylation is shown in purple. (B) In the absence of blue light, BioID is not attracted to the condensates, resulting in no biotinylation at the condensate site. In the "All" and "Cont1" conditions, free BioID can biotinylate nuclear proteins; therefore, this background level is shown in a darker purple than the background levels in the "cont2" and "cont3" conditions. (C) Without external biotin, miniTurboID does not function properly. Even though blue light causes GFP to accumulate in condensates, the accumulated miniTurboID does not biotinylate proteins in the condensates. (D) Without doxycycline treatment, the BioID-SunTag complex is not expressed, resulting in no overall biotinylation. (E–H) Immunofluorescence (IF) of biotinylated proteins using Streptavidin-Alexa647 under different conditions. This confirms the expected results in panels A–D. Arrows indicate the location of transcriptional condensates. (I and J) Statistical analysis of the relative GFP intensity and Streptavidin-Alexa647 intensity under each condition. Only "All" indicates biotinylation accumulated in condensates. The center line within the box indicates the median, and the plus sign (+) indicates the mean. ("All": n = 36, "Cont1": n = 23, "Cont2": n = 22, "Cont3": n = 21). [Figure 11]A schematic diagram of the sample preparation process for mass spectrometry analysis using streptavidin-biotin pull-down is shown. First, nuclear proteins are extracted after appropriate treatments, such as blue light irradiation, biotin incubation, and doxycycline incubation. The extract contains biotinylated and non-biotinylated proteins. Because streptavidin and biotin have a very strong affinity (dissociation constant approximately 1 × 10-14 M), biotinylated proteins are pulled down using streptavidin beads. To account for nonspecific binding of proteins from the beads, a control bead experiment (streptavidin beads blocked with free biotin) should also be performed. [Figure 12]The mass spectrometry results were verified by immunofluorescence analysis of CCNT1, CDK7, and CDK9. Primary antibodies for CCNT1, CDK7, and CDK9 were incubated, followed by visualization with a secondary antibody containing Alexa647. (A) Transcriptional condensates were identified by the BRD4 IDR using mCh (green). (B) CCNT1 was visualized by indirect immunofluorescence. CCNT1 formed foci within the nucleus. (C) Images of BRD4 IDR and CCNT1 were overlaid. Orange arrows indicate colocalized foci, while light blue arrows indicate non-colocalized IDR-only foci. This indicates that most transcriptional condensates colocalize with CCNT1 clusters. (D) Zoomed-in red box in panel C. Scale bar = 1 μm. (E) Relative intensity profile through the arrowhead in panel D. This confirms the colocalization of transcriptional condensates with cyclin T1 clusters. (F-H) CDK7 is labeled by indirect immunofluorescence, and transcriptional condensates are marked by IDRs. CDK7 forms numerous foci, but these foci do not colocalize with transcriptional condensates. (I) Enlargement of the red box in panel H. Scale bar: 1 μm. (J) Relative intensity profile through the arrowhead in panel I, confirming that transcriptional condensates do not colocalize with CDK7 clusters. (K-M) CDK9 is visualized by indirect immunofluorescence. CDK9 forms foci in the nucleus, some of which colocalize with transcriptional condensates. (N) Enlargement of the red box in panel M. Scale bar: 1 μm. (O) Relative intensity profile through the arrowhead in panel N, confirming that only relatively small transcriptional condensates colocalize with CDK9 clusters. [Example]
[0176] Development of a system for light-induced targeting of endogenous condensates Example 1: Targeting endogenous transcriptional condensates To study the properties of biomolecular condensates, new methods are needed to maintain the internal and external environments of the condensates intact within cells during experiments. Isolating intact biomolecular condensates from cells is feasible. However, isolating condensates disrupts their surroundings, making isolation of small biomolecular condensates (<1 μm), especially those in liquid form, extremely challenging. Previous studies have isolated P granules and nucleoli that are relatively large (>1 μm). Therefore, we set out to develop a new in vivo system that could target and manipulate endogenous transcriptional condensates to study transcriptional condensates, which are in liquid form and typically smaller than 1 μm.
[0177] Zip-code of condensates; Intrinsically Disordered Regions (IDR) The formation of biomolecular condensates results from thermodynamic interactions between different biomolecules. Many studies have shown that weak and multivalent interactions, such as ionic interactions from intrinsically disordered regions, promote phase separation (Figure 1). We hypothesized that IDRs with the same syntax as transcriptional condensates can be preferentially located in transcriptional condensates, and therefore overexpression of these IDRs could be used to target condensates. Unlike structured regions such as DNA-binding domains, IDRs are thought to have no specific function other than weak interactions with other proteins. Therefore, using IDRs seemed a rational approach because their lack of binding ability minimizes functional or compositional changes in condensates that may occur when IDRs are overexpressed in living cells.
[0178] We selected two natural IDRs from the transcription-related proteins BRD4 and NELFA to test whether they can indeed colocalize with transcription condensates. BRD4 is a transcription factor that is enriched in super-enhancer regions and upregulates gene transcription by binding to acetylated histones. NELF (negative elongation factor) is a four-subunit protein complex that downregulates transcription by pausing RNA Pol II. We selected the NELFA subunit for this test. We intentionally selected functionally opposite proteins to demonstrate that the IDR targeting concept does not depend on protein function but on the structural similarity of the IDRs that form condensates.
[0179] We used PONDR (pondr.com), a website-based amino acid sequence analysis algorithm, to identify intrinsically disordered regions in the BRD4 and NELFA proteins. Figures 1A and 1B show the PONDR VSL2 scores for NRD4 and NELFA. The X-axis is the amino acid (aa) position, and the Y-axis is the VSL2 score, which indicates how highly disordered the amino acid sequence is (regions with high VSL2 scores are intrinsically disordered regions). The IDR is indicated by a bold black line in the center of the Y-axis. For BRD4, we used only the IDR portion of the BRD4 sequence, from position 462 to the C-terminus of the BRD4 sequence (position 1362). For NELFA, the IDR of this protein is located in the center of NEFLA, so the entire NEFLA protein was used for testing. To express these proteins in mESCs, two plasmids were generated by standard recombination methods: one containing the BRD4 IDR and mCherry (mCh) sequence, and the other containing the entire NELFA sequence and the mCh sequence. mCh is used to image the BRD4 IDR and NELFA, and the DNA sequence is then integrated into the mESC genome by lentiviral transduction.
[0180] Endogenous Mediator or Pol II condensates were imaged in live mESCs using epifluorescence microscopy (Figures 1C and 1G, respectively). Using the CRISPR / Cas9 genome editing system, HaloTag sequences were added to the N-terminus of Rpb1 (a Pol II subunit) and Med19 (a Mediator subunit), followed by Halo-JF646 for imaging. Next, the BRD4 IDR and NELFA were imaged as shown in Figures 1D and 1H. Both the BRD4 IDR and NELFA clustered within the nucleus, and overlay with endogenous Mediator or Pol II images confirmed that these clusters colocalized with Mediator or Pol II condensates (Figures 1E and 1I, respectively). The intensity profiles of the BRD4 IDR and NELFA through the orange arrows in panels E and I are shown in Figures 1F and 1J, respectively. This shows that both the BRD4 IDR and NELFA have a higher condensate-to-background intensity ratio than either Mediator or Pol II. Overall, this result confirms that the BRD4 IDR and NELFA can be used to target endogenous transcription condensates in live mESCs.
[0181] The colocalization of BRD4 IDRs with transcriptional condensates was expected due to their functional relationship to super-enhancers and transcription. However, the colocalization of NELFA was unexpected, as transcriptional condensates are thought to upregulate transcription by concentrating the transcriptional machinery. However, although NELF downregulates transcription by pausing Pol II, it may have a similar biomolecular grammar to other transcription factors. This reflects the fact that the function of transcriptional condensates has not yet been fully elucidated. IDRs that can target specific biomolecular condensates can also be referred to as "zipcodes" of biomolecular condensates. Because different IDRs target different types of biomolecular condensates, altering the IDR "zipcode" can target specific condensates.
[0182] We performed FRAP experiments on BRD4 IDR clusters to examine their mobility. Figure 2A shows time-lapse images of a representative cell. Two BRD4 IDR clusters are shown in yellow (top) and blue (bottom) boxes. The cluster in the yellow box is photobleached at time t = 0, while the cluster in the blue box is the control. At t = 0, the cluster is photobleached to an intensity level similar to the nuclear signal. However, after 1 second of photobleaching, it quickly begins to recover. Panel B shows the average intensity profile of the BRD4 IDR clusters as a function of time (n = 5). This demonstrates the rapid recovery of the BRD4 IDR clusters. The half-recovery time is 3.4 seconds, and the recovery rate is 90%. The inset in Panel B shows the intensity profile of the control cluster in the blue box. This decreases monotonically upon photobleaching.
[0183] In summary, we assayed the IDR portions of two proteins, BRD4 and NELF, for their ability to target transcriptional condensates of Mediator and Pol II in mESCs. We confirmed that both the BRD4 IDR and NELF formed clusters in mESCs and that these clusters colocalized with transcriptional condensates. Intensity analysis showed that even the condensate-to-background ratio was higher than that of Mediator or Pol II. The BRD4 IDR also exhibited high mobility. Thus, by using the BRD4 IDR, we successfully targeted endogenous transcriptional condensates. This confirms the first requirement of our novel system.
[0184] Optogenetics and Multivalency in Exemplary Embodiments of the System After demonstrating that targeting transcriptional condensates with appropriate zipcodes (IDRs) is possible, the next step was to find a way to manipulate transcriptional condensates transiently but effectively. Temporal manipulation is important here. If manipulation of condensates is performed continuously, it may be impossible to distinguish whether the resulting condensate characteristics are solely due to the natural properties of the condensate or are influenced by the artificial manipulation of the condensate.
[0185] Optogenetics is a biological technique that uses light to control the activity of cells, tissues, or even organs. The concept of optogenetics arose from neuroscientific studies of rhodopsin, a light-gated ion channel. Many groups have used optogenetic tools in various fields for precise and simultaneous response control. In the field of biomolecular condensates, research has been conducted using optogenetics to create artificial biomolecular condensates (see Bracha et al., (2018) Cell, 175(6):1467-1480.e13; Shin et al., (2017) Cell, 168(1-2):159-171.e14; Shin et al., (2018) Cell, 175(6):1481-1491.e13; Shunsuke et al., (2021) Nature, 599(7885):503-506). Here, instead of creating artificial condensates, we employed the optogenetic components used by these authors to target proteins of interest and recruit them into pre-existing endogenous transcriptional condensates with blue light.
[0186] The improved light-inducible dimer (iLID) and sspB, one of the optogenetic systems, are shown in Figure 3A. When iLID is irradiated with blue light (450-500 nm), it undergoes a conformational change, dramatically increasing its binding affinity to sspB.
[0187] The dissociation constant K between iLID and sspB D The binding potential of the antibody is 4700 nM in the dark but 130 nM under blue light, meaning that binding under blue light is 36 times stronger than binding in the dark. Furthermore, the binding process is reversible: binding begins within seconds of blue light irradiation and dissociates within minutes after the blue light is turned off, allowing instantaneous signal control by adjusting the intensity of the blue light.
[0188] Furthermore, the SunTag system can be used to more efficiently capture proteins of interest, and its high signal-to-noise ratio facilitates imaging. Figure B shows SunTag and scFv (single-chain variable fragment). 24xSunTag provides a multivalent interaction because the scFv has 24 peptide epitopes to which it can bind. The dissociation constant K between scFv and SunTag is D The SunTag is commonly used to increase the fluorescence signal-to-noise ratio (SNR).
[0189] The design of an exemplary, non-limiting embodiment of the system according to the present invention is shown in Figure 3C. It contains three components: the BRD4 IDR fused to mCh and sspB (IDR-mCh-sspB), the iLID fused to scFv and sfGFP (scFv-sfGFP-iLID), and 24x SunTag. The scFv and SunTag bind constantly, forming a complex containing up to 24 sfGFP and iLID. Subsequently, upon blue light irradiation, this complex begins to bind to up to 24 IDR-mCh-sspB. Shin et al. used similar components to generate artificial droplets in HEK293m or U2OS cells and called this system CasDrop (because dCas9 is fused to SunTag). In this study, we use these three components to recruit a protein of interest to pre-existing endogenous condensates. This is due to the IDR remaining in the endogenous condensates within mESCs. The iLID-SunTag complex will enter the endogenous condensate to interact with sspB fused to the IDR. We refer to this three-component system by the acronym "LiTEC," which stands for "light-induced targeting of endogenous condensate."
[0190] To test whether LiTEC can actually recruit proteins originally uniformly distributed throughout the nucleus into pre-existing endogenous condensates using blue light, we monitored the localization of GFP-tagged proteins. All three components were expressed in mESCs. IDR-mCh-sspB was imaged using a 561 nm laser (Figure 4A). As previously observed, the BRD4 IDR accumulates within endogenous transcriptional condensates. Three representative clusters are highlighted in yellow. Next, scFv-sfGFP-iLID was imaged using a 488 nm laser. Because the 488 nm laser itself is blue light, imaging this component immediately induces optogenetic binding. However, because it takes several seconds for iLID and sspB to interact, at the moment of blue light activation (time t = 0), the GFP signal is (still) uniformly distributed throughout the nucleus (Figure 4B). In the cell line containing only the scFv-sfGFP-iLID construct, cells exhibit only a uniform distribution of GFP under blue light. This is because iLID lacks its binding partner, sspB (Figure 4G). However, in LiTEC, GFP begins to accumulate within the IDR clusters after a few seconds of blue light irradiation (Figure 4C-E). Figure 4F shows the average GFP intensity as a function of time in three representative IDR clusters. While it increases rapidly over time, the GFP intensity in the reference cluster decreases monotonically upon photobleaching. This result indeed confirms the successful insertion of the iLID-SunTag complex into pre-existing transcriptional condensates by blue light.
[0191] However, this experiment shows that proteins lacking enzymatic activity (GFP and SunTag) are inserted into the condensates. We wondered whether it would be possible to insert proteins with enzymatic function to manipulate the condensates or to effectively control their transcription. Our hypothesis was that a protein (cargo) fused to 24xSunTag might be transported into the condensates upon blue light irradiation, which fuses SunTag to the IDR (Figure 4H). In principle, any cargo could be transported.
[0192] However, before further exploring this model, we analyzed the biophysical effects of LiTEC on transcriptional condensates. Because this system artificially inserts several exogenous components that are not naturally present in the condensate, these components may affect the properties of the transcriptional condensate. In the next subsection, we present the results of experiments aimed at verifying how the LiTEC system affects the composition of transcriptional condensates.
[0193] Analysis of quantitative composition changes during LiTEC The underlying hypothesis of this study is that due to the multivalency of the iLID-SunTag complex, some BRD4 IDRs ("outside IDRs") originally located on the outside of the condensate may enter the condensate upon blue light irradiation. It is likely that the iLID-SunTag complex first binds to the "outside IDRs" before finding the "inside IDRs" present within the condensate. However, once one iLID in the iLID-SunTag complex binds to the "inside IDR," other iLIDs in the iLID-SunTag complex will easily find an IDR to bind to due to the high IDR concentration in the condensate. Therefore, the iLID-SunTag complex will remain inside the condensate. During this process, the "outside IDRs" may enter the condensate along with the iLID-SunTag complex, resulting in a higher IDR intensity in the condensate compared to its intensity before blue light irradiation.
[0194] To test this hypothesis, IDR-mCh-sspB was imaged for several seconds without blue light using a 561 nm laser. Then, scFv-sfGFP-iLID was imaged for 1 minute with blue light (488 nm laser). Finally, IDR-mCh-sspB was imaged again for several seconds without blue light. Ideally, the time course of the mCh signal under blue light would be observed, but this is not possible due to crosstalk between mCh and GFP. Figure 5A and B show the distribution of IDR in the nucleus before blue light irradiation ("Before BL") and after 50 seconds of blue light irradiation ("After BL (50 s)"), respectively. Figure 5D and E show the distribution of GFP in the nucleus at the moment of 50 seconds of blue light irradiation and the time course profile of the relative GFP intensity under blue light irradiation, respectively. Here, "relative intensity" refers to the maximum intensity in the condensate divided by the average intensity of the nuclear background.
[0195] Relative intensity = maximum intensity in the condensate / average intensity of the nuclei
[0196] A relative intensity equal to 1 means that the intensity in the condensate is equal to the intensity of the background. A similar term, "relative change," can be used if the background level is set to zero.
[0197] Relative change = (intensity in condensate - background intensity) / background intensity
[0198] Relative change is the true signal from background, so it is similar to the concept of signal to noise ratio, but normalized (background is 1). When determining the difference in relative intensity at different time points, the relative intensity difference can be calculated:
[0199] Relative intensity difference = relative change at time A - relative change at time B = Relative intensity at time A - Relative intensity at time B
[0200] Another term is "Fold-Change (FC) of a component" BL / NoBL,Time", which is introduced as the ratio of the relative change of a component after time with and without blue light (BL) to the initial relative change of the component. Two examples illustrate this concept:
number
[0201] The intensities of the IDR and iLID of BRD4 can be compared before and after blue light irradiation. Figure 5C shows the intensity profile of the IDR of BRD4 through the orange arrow shown in panel B. The black line represents the initial relative intensity of the IDR of BRD4. The red line represents the relative intensity of the IDR of BRD4 after 50 seconds of blue light irradiation. For the first condensate, the initial relative intensity is I 0s The relative intensity of the blue light after 50 seconds is expressed as I 50s The fold change of this condensate is expressed as mCh Fold-Change BL50s is (I 50s -I B ) / (I 0s -I B ) In this typical condensate, the Fold-Change BL50s = 1.91. This confirmed that blue light further accumulates the BRD4 IDR in transcriptional condensates. More specifically, after 50 seconds of blue light, the relative change in mCh was 91% higher than the initial relative change. Therefore, calculating the "fold change" is a good way to compare the initial and final relative changes based on the initial relative change. Similar to the mCh signal, the relative intensity of the GFP signal as a function of time is shown in Figure 5E. This again confirms that GFP accumulates in condensates during blue light irradiation. The time-dependent accumulation of GFP in condensates can be interpreted as the association kinetics of iLID and sspB. Therefore, the time-dependent relative intensity of GFP can be fitted by a one-phase binding equation.
number
[0202] Finally, we compared the "difference in GFP relative intensity" and the "difference in mCh relative intensity" (relative intensity BL30s -Relative Strength BL0s We verified the correlation between GFP accumulation and IDR enhancement by plotting a scatter plot of the relative intensity of mCh (p < 0.05). In this experiment, we measured the relative intensity of mCh before and after blue light exposure, and the relative intensity of GFP before and after 30 seconds of blue light exposure. Figure 5F shows that the difference in relative intensity of GFP and the difference in relative intensity of mCh are positively correlated. In other words, the more GFP accumulates in the condensate under blue light, the more the "outer IDR" accumulates in the condensate. How does IDR accumulation affect the properties of transcriptional condensates? Unlike SunTag, the IDR of BRD4 is multivalent and may therefore interact with other components, potentially affecting the composition of the condensate. Figure 5G briefly illustrates the unknown effect of IDR enhancement by blue light. While the blue light is on, the IDR concentration in the condensate begins to increase, which may increase, decrease, or remain unchanged the concentrations of endogenous components such as BRD4, MED, and Pol II. To investigate how endogenous proteins respond to increased IDRs, we labeled endogenous BRD4 with HaloTag and determined how BRD4 condensates respond to increased IDRs. Two experiments were designed:
[0203] A) Blue light was turned on for 1 minute, and changes in the GFP signal were observed. Endogenous BRD4 was also observed while the blue light was on ("blue-on" experiment). Figure 6A shows the procedure for the first experiment. Before the blue light was turned on ("pre-BL"), the BRD4 IDR was first imaged for 5 seconds by capturing the mCh signal with a 561 nm laser. Then, the GFP and Halo-BRD4 signals were simultaneously captured for 70 seconds using 488 nm and 642 nm lasers and a dual camera. Prior to this imaging, 100 nM Halo-JF646 dye was added for 15 minutes and washed twice with 2i medium. For all images, 11 slices of the z-stack were captured with a 300 nm gap to ensure that the condensates of interest did not fall out of focus during imaging. After 70 seconds of blue light ("post-BL"), the BRD4 IDR was again imaged for 5 seconds to confirm the extent to which it increased with the 70 seconds of blue light. In Figure 6B, the relative intensity of (endogenous) BRD4 and the relative intensity of GFP are plotted as a function of time. Generally, the BRD4 signal decreases due to bleaching, while the GFP signal peaks within 1 minute. BRD4,BL70s is 0.77 for the indicated condensate (orange arrow). This means that the relative change in BRD4 after 70 seconds of blue light is 23% less than the initial relative change. The relative change in mCh around the indicated condensate before and after blue light also shows that the relative change in IDR is 70% more than the initial relative change. (Fold-Change of this condensate) mCh,BL70s (The fold-change coefficient is 1.70). However, this experiment alone does not allow us to conclude whether the decrease in BRD4 fold-change is due to enhanced IDR or photobleaching. A control experiment is shown in Figure 6C. Halo-BRD4 is imaged for 70 seconds without blue light exposure ("Blue-Off" experiment). Before and after BRD4 imaging, the IDR of BRD4 was also imaged for 5 seconds ("Before" and "After"). Therefore, no GFP signal is measured. Figures 4-6D show the relative intensity profiles of BRD4 and mCh over time through the indicated condensates. The BRD4 signal decreases over time, while the mCh signal does not change significantly. In this case, the Fold-Change coefficient isBRD4,NoBL70s is 0.78, and the ratio mCh,BL70s is 0.93.
[0204] To obtain meaningful conclusions, we statistically compared the fold change of the Halo-BRD4 signal between the two groups. We analyzed 50 "Blue On" condensates and 39 "Blue Off" condensates. When blue light was turned off ("Blue Off"), no difference in the relative mCh intensity was observed before and after 70 seconds of BRD4 imaging. However, when blue light was turned on ("Blue On"), the relative mCh intensity increased; the median and mean values of the difference between "After BL" and "Before BL" were 0.251 and 0.288, respectively (Figure 6E). The median and mean values of the fold change of mCh (after BL relative to before BL) in the "Blue On" condition were 1.548 and 1.537, respectively, while those in the "Blue Off" condition were 1.000 and 1.021 (data not shown). Finally, the fold change of BRD4 in the "Blue On" and "Blue Off" conditions is analyzed in Figures 4-6F. The median and mean values were 0.8109 and 0.8375 (±0.0176, SEM) for "Blue On" and 0.8085 and 0.7965 (±0.0128, SEM) for "Blue Off," indicating that the difference between the two is not significant (numerically, median difference = 0.00234, mean difference = 0.041). In other words, the enhancement of IDR (in this case, a 54.8% increase in BRD4 IDR) does not affect endogenous BRD4 strength.
[0205] Endogenous mediators (MEDs) are labeled with HaloTag and imaged using the same protocol used in the BRD4 experiments (Figure 7). Seventy-six "blue on" condensates and 55 "blue off" condensates are analyzed. Panel E shows that mCh relative intensity changes after blue light but not without blue light. For "blue on," the median and mean difference are 0.269 and 0.315, respectively. The median and mean fold change in mCh (after BL vs. before BL) in the "blue on" condition was 1.438 and 1.495, whereas those in the "blue off" condition were 1.098 and 1.126 (data not shown). Finally, the fold change in MED for "blue on" and "blue off" conditions is analyzed in Figure 7F. The median and mean values were 0.8215 and 0.8189 (±0.0096, SEM) for "Blue On" and 0.7717 and 0.7750 (±0.0141, SEM) for "Blue Off." The median difference was 0.0439, and the mean difference was 0.0498, slightly higher than the median and mean differences in BRD4 fold change in the previous experiment. Statistically, the fold changes in MED between the "Blue On" and "Blue Off" groups are distinguishable (only a 4.4% increase due to blue light), although the median and mean differences are not significant. This translates to 4.4% more endogenous MED in the condensates after 70 seconds of blue light compared to the control ("Blue Off") experiment. This suggests that the blue light-induced increase in IDR (in this case, the IDR of BRD4 increased by 43.8%) may alter the thermodynamics of the condensate slightly more favorably for the mediator, allowing a small amount of endogenous mediator that was originally outside the condensate to enter the condensate.
[0206] In both cases, after 70 seconds of blue light exposure, the mCh signal of the condensates increased by more than 40%, resulting in no difference in the number of endogenous BRD4s, but a slight increase in the number of endogenous MEDs within the condensates. Our current hypothesis for these results is as follows: BRD4 binds to acetylated histones using two bromodomains. Increasing the number of IDRs of BRD4 in the condensates does not change the number of acetylated histones in the condensates, so the number of BRD4 binding sites remains the same. Therefore, despite the increase in the number of IDRs, it does not actually lead to an increase in endogenous BRD4 in the condensates. However, increasing the IDRs may increase the multivalency of the condensates, thereby slightly increasing their ability to interact with other biomolecules. Therefore, condensates may recruit more biomolecules, especially those that directly interact with the IDRs of BRD4, such as Mediator. Of course, the IDRs of BRD4 also interact with other BRD4s, but weak interactions between IDRs may not be the main factor in BRD4 recruitment. Nevertheless, the conclusion that enhancing IDRs does not affect BRD4 but increases the number of Mediators in the condensate is difficult to conclude with certainty. The LiTEC system does not significantly affect the properties of transcriptional condensates, despite the insertion of many external biomolecules. Thus, the LiTEC system reliably preserves the native properties of the condensates studied.
[0207] Example 2: Modification of endogenous transcription condensates In Example 1 above, we demonstrated the potential of the LiTEC system for studying the characteristics of transcriptional condensates. The system demonstrated there was able to transport iLID-SunTag complexes and auxiliary IDRs into condensates in response to blue light. Therefore, is it possible to transport proteins of interest using LiTEC? We hypothesize that cargoes with appropriately fused IDRs could be inserted into condensates, likely multiplexed by the SunTag complex, as depicted in Figure 4H.
[0208] To test this hypothesis, we considered proximity-based labeling methods such as DamID and BioID. DNA adenine methyltransferase identification (DamID) is used to detect the binding sites of eukaryotic DNA and chromatin-binding proteins by fusing Dam (DNA adenine methyltransferase) protein. Dam protein recognizes a specific DNA sequence (GATC) within its vicinity (approximately 10 nm) and methylates its adenine residue. Because adenine methylation does not occur naturally in eukaryotes, the methylated adenine region likely reflects the binding site of a DNA-binding protein. Biotin identification (BioID) is used to detect proteins that interact with a protein of interest by fusing BirA protein. BirA protein biotinylates nearby proteins within approximately 10 nm. Biotinylated proteins can be pulled down using streptavidin beads and sequenced by mass spectrometry. The list of identified proteins will reflect the list of proteins that closely interact with the protein of interest.
[0209] BioID was chosen as an exemplary candidate because of the ability of the biotin tag to reveal what other biomolecules are contained within the transcriptional condensate, and because biotinylated proteins are also easily identified by isolation with streptavidin followed by mass spectrometry (MS).
[0210] Biotinylation Identification (BioID) and LiTEC To test biotinylation by BioID using LiTEC, we first fused BioID to 24xSunTag to generate the cell line shown in Figure 9A. Here, miniTurboID was chosen as the BioID due to its small size, fast biotinylation rate (10 min), and relatively low biotinylation level even without exogenous biotin. Despite the low biotinylation level from miniTurboID without exogenous biotin, we aimed to minimize nonspecific biotinylation during cell culture. Therefore, we employed a doxycycline (Dox)-inducible system to control the expression level of the BioID-SunTag complex. The doxycycline-inducible system uses the tetO (tetracycline operator) sequence and rtTA (reverse tetracycline-controlled transactivator) (Gossen et al., Science, 268(5218):1766-1769, 1995; Atze, Curr. Gene Ther., 16(3):156-167, 2016). The tetO sequence serves as a promoter for the gene of interest. Only upon tetracycline or doxycycline treatment does rtTA begin to bind to the tetO site and promote transcription of the gene of interest. Throughout this experimental description, unless a specific incubation time for Dox is mentioned, all BioID experiments are performed after 48 hours of Dox incubation.
[0211] To test whether the Dox-inducible system works well and how doxycycline concentration affects the LiTEC system, cells were incubated with different Dox concentrations. A cell line containing only IDR-mCh-sspB and sfFv-sfGFP-iLID but not SunTag ("No SunTag") was used as a control. The Dox-inducible BioID-SunTag cell line was incubated with 0, 1, 2, or 4 μg / ml of Dox for 48 hours, and GFP accumulation in the condensates was imaged under blue light. Figure 9B shows the fitting of the relative GFP intensity over time for each condition. The first panel shows an example of the raw data and fitted line of the relative GFP intensity over time (Dox 1 condition). Fitting was performed using a monophasic binding curve, as described in Section 4.6. The control experiment shows low GFP accumulation. Under blue light, sspB and iLID bind, allowing a small amount of GFP in the condensates to interact with sspB. Similar GFP accumulation results were observed in cells to which Dox was not added, indicating that expression of the BioID-SunTag complex was suppressed.
[0212] Treatment with Dox 1, 2, and 4 (μg / ml) significantly enhanced GFP accumulation, indicating successful expression of the BioID-SunTag complex. The mean difference in GFP relative intensity in each condition between before and after 30 seconds of blue light exposure is shown in Figure 9C. For "No SunTag" and "Dox 0," the mean values were 0.206 (±0.144, SD, n = 39) and 0.315 (±0.158, SD, n = 39). For "Dox 1," "Dox 2," and "Dox 4," the mean values were 2.382 (±1.135, SD, n = 43), 2.362 (±1.371, SD, n = 46), and 2.654 (±1.217, SD, n = 51), respectively. These results confirm that the expression level of BioID-SunTag complexes is doxycycline-inducible, thereby reducing the background level of biotinylation from miniTurboIDs, and indirectly confirm that blue light induces co-intercalation of BioIDs into condensates.
[0213] Streptavidin conjugated with Alexa647 (Streptavidin-Alexa647) allows direct visualization of biotinylated proteins. In this experiment, all cells were fixed with 4% PFA for 10 minutes at room temperature, followed by treatment with 0.5% Triton X-100 for 5 minutes at room temperature. The cells were then blocked overnight with 2% BSA. Finally, Streptavidin-Alexa647 (1:1500) was incubated in 2% BSA for 1 hour. Between each step, the samples were washed three times with 1x PBS.
[0214] There are three variables in this experiment: blue light, external biotin, and doxycycline. The four most important conditions to test are as follows: In the "All" condition, cells are treated with all three variables. Cells are first incubated with Dox 2 μg / ml for 48 hours, and external biotin (50 μM) is added just before turning on the blue light. The blue light is on for 30 minutes to allow sufficient time for biotinylation. These conditions are referred to as BL (blue light, 30 minutes), Bio (biotin 50 μM, 30 minutes), and Dox (doxycycline 2 μg / ml, 48 hours). The first control ("Cont1") is cells treated without BL but with Bio and Dox. The second control ("Cont2") is cells treated without Bio but with BL and Dox. Finally, the third control ("Cont3") is cells treated without Dox, but with BL and Bio. The expected results from these conditions are shown in Figures 10A-D, where biotinylation is indicated by a purple color. The higher the expected biotinylation concentration, the darker the purple color.
[0215] In the "All" condition, we predicted a high accumulation of biotinylation in transcriptional condensates (dark purple dots in the nucleus), but there was also a significant amount of background biotinylation. This is because BioIDs that are not drawn into condensates may randomly biotinylate nuclear proteins. In the "Cont1" condition, which does not use blue light, we predicted the same background biotinylation as in the "All" condition, but no accumulation of biotinylation in the condensates. For "Cont2" and "Cont3," we predicted that the amount of biotinylation in the nucleus would be very low because the cells do not have enough biotin to use ("Cont2") or because the cells do not express BioIDs ("Cont3"). Of course, we also predicted no accumulation of biotinylation in the condensates. If this prediction is correct, we may be able to obtain a list of "transcriptional condensate proteins" (proteins that remain in transcriptional condensates) by directly comparing the results of the "All" condition with those of the "Cont1" condition. Additional validation is achieved by comparing the result of the condition "All" with the result of the conditions "Cont2" or "Cont3".
[0216] Figures 10E–H show the intracellular signals of BRD4 IDR, GFP, and Streptavidin-Alexa647 under four conditions. Arrows indicate the location of transcription condensates based on the IDR-mCh-sspB images. The GFP signal indirectly indicates the location of BioID-24×SunTag. Streptavidin-Alexa647 indicates biotinylated proteins within the cells. In the "All" case, GFP accumulates within transcription condensates, and Streptavidin-Alexa647 also accumulates within the condensates. This suggests that biotinylation occurs within the condensates via BioID-24×SunTag. However, in the other three controls, no accumulation of Streptavidin-Alexa647 is observed within the transcription condensates. In "Cont1," GFP does not accumulate within the condensates, indicating that BioID-24×SunTag does not accumulate either. In "Cont2," GFP accumulates in the condensates, as shown in the "All" experiment, but BioID does not efficiently biotinylate neighboring proteins due to the lack of biotin. In "Cont3," GFP weakly accumulates in the condensates, but this accumulation is due to direct binding between sspB and iLID, not to enhancement of GFP by SunTag. Therefore, cells in this condition do not fully express the BioID-SunTag complex, which results in the absence of biotinylated proteins in the condensates. Note also the difference in the background biotinylation levels in the nuclei. In "All" and "Cont1," biotinylation occurs primarily inside the nucleus, allowing the morphology of nuclei and nucleoli to be visualized. In "Cont2" and "Cont3," biotinylation does not primarily occur in the nucleus, resulting in a relatively lower nuclear signal than in the cytoplasm. Statistical analysis of the relative GFP and Streptavidin-Alexa647 intensities is shown in Figures 10I and 10J (again, relative intensity means the peak intensity at the condensate site divided by the average background intensity). In the case of GFP, "All" and "Cont2" indicate accumulation within the condensates. In "Cont3," sspB and iLID interact directly, resulting in almost no GFP accumulation, and in "Cont1," neither accumulation is observed.The mean (+ symbol in the box plot) and median (midline in the box plot) values for "All" were 4.730 (±1.914, SD, n=36) and 4.169, respectively, for "Cont1" were 1.118 (±0.069, SD, n=23) and 1.116, respectively, for "Cont1" were 4.548 (±2.455, SD, n=22) and 3.536, respectively, and for "Cont3" were 1.959 (±0.911, SD, n=21) and 1.724, respectively. For streptavidin-Alexa647, only the "All" condition showed accumulation in aggregates. The mean values (+ marks in the box plots) and medians (midline in the box plots) were 2.291 (±0.543, SD, n=36) and 2.245 for "All," 1.116 (±0.049, SD, n=23) and 1.102 for "Cont1," 1.181 (±0.107, SD, n=22) and 1.155 for "Cont2," and 1.107 (±0.082, SD, n=21) and 1.117 for "Cont3."
[0217] This proof-of-concept experiment clearly demonstrates that BioIDs are accumulated in transcription condensates by blue light and biotinylate neighboring proteins within the condensates. Thus, we conclude that the LiTEC system can attract BioIDs into condensates using blue light, and that the attracted BioIDs successfully biotinylate proteins within the condensates.
[0218] mass spectrometry The number of cells is amplified to prepare sufficient nuclear extracts for mass spectrometry experiments. The basic procedure for mass spectrometry sample preparation is shown in Figure 11. First, cells are grown under appropriate treatments, such as blue light irradiation, exogenous biotin incubation, and doxycycline incubation. For the first preliminary data, the conditions used in Figure 10 are maintained. Then, the cells are sacrificed and only the nuclear extract is isolated. The nuclear extract contains both biotinylated and non-biotinylated proteins. Streptavidin has a very specific and strong affinity for biotin (the dissociation constant is approximately 1 × 10). -14M) Streptavidin beads can be used to pull down biotinylated proteins. However, it should be noted that nonspecific binding proteins to the beads exist, and these proteins must also be taken into account. "Control" beads are incubated with free biotin to block all available streptavidin. After blocking, the control beads are incubated with nuclear extract to check for nonspecific binding proteins. After reading the mass spectrometry results, a list of true biotinylated proteins can be obtained by comparing the list of control beads with the original beads. After incubating the nuclear extract with two types of streptavidin beads, the streptavidin-biotin bond is cleaved by chemical treatment and desalted. Finally, the sample is loaded for LC-MS / MS analysis.
[0219] Similar to the imaging experiment in Figure 10, four mass spectrometry samples were obtained under four different conditions. The goal of this experiment was to identify proteins in small condensates, which means that the number of proteins in small condensates may be very low. Therefore, this analysis focuses on highly biotinylated proteins rather than the amount of biotinylated proteins.
[0220] Pilot mass spectrometry experiments were performed under various conditions. In the "All" group, blue light (BL) was applied for 30 minutes, 50 μM biotin (Bio) was added for 30 minutes (with BL), and 2 μg / ml doxycycline was added for 48 hours. In control experiments, one of these treatments was omitted.
[0221] Positive cofactor 4 (PC4), cyclin T1 (CCTN1), cyclin T2 (CCTN2), cyclin-dependent kinase 9 (CDK9), bromodomain-containing protein 4 (BRD4), and mediator subunit 15 (MED15) were found to be highly enriched proteins and are considered proteins in transcription condensates. RNA polymerase II was not detected in this experiment, nor was CDK7.
[0222] PC4, BRD4, MED15, CDK9, CCNT1, and CCNT2 were identified as present and successfully labeled in pilot MS analysis experiments. Their enrichment ratios were greater than 2, suggesting their possible presence in transcriptional condensates. These proteins are also functionally associated with transcriptional activity. In particular, this result is relevant, given that transcriptional condensates are already known to contain BRD4 and Mediator. CDK9, cyclin T1, and cyclin T2 are known subunits of P-TEFb (Positive Transcription Elongation Factor b), which phosphorylates the CTD of the Pol II Rpb1 subunit, DSIF, and NELF, thereby helping to release paused RNA Pol II. This may suggest that transcriptional condensates upregulate transcription by concentrating P-TEFb subunits in the condensates. Interestingly, RNA polymerase II subunits were not found in all of the samples. CDK7, which is involved in transcription initiation, was also not found in these samples.
[0223] Validation of mass spectrometry results by immunofluorescence In the first experiment, we selected anti-CCNT1, anti-CDK7, and anti-CDK9 antibodies. Although CDK7 was not included in the protein list based on mass spectrometry analysis, we were interested to see how CDK7 differed in its nuclear distribution compared to CCNT1 or CDK9, given its important role in transcriptional activity. It could also be used as a negative control. Figure 12 shows the immunofluorescence results for CCNT1, CDK7, and CDK9 in fixed cells. IDR-mCh-sspB (green) was used to detect transcriptional condensates in the nucleus (Figure 12A, F, K). Indirect immunofluorescence was used to visualize CCNT1, CDK7, and CDK9. Secondary antibodies with Alexa647 were imaged and stained magenta (Figure 12B, G, L). Next, the IDR image and the anti-protein image were overlaid to confirm their distribution and colocalization within the nucleus (Figure 12C, H, M). The yellow dashed line indicates the nuclear membrane of the cell. Additionally, in panels D-E, I-J, and N-O, the red boxes of each immunofluorescence (IF) image are enlarged to show the relative intensity profile along the light blue arrowheads. Both orange and light blue arrows are present. Both arrows indicate the location of transcriptional condensates based on the IDR image, but light blue is used when the IDR clusters do not colocalize with the IF foci. The orange arrow indicates a cluster that colocalizes with the IF foci.
[0224] Similar to IDR-induced transcription condensates (Figure 12A), CCNT1 forms uniform background-level foci within the nucleus (Figure 12B). Most transcription condensates colocalize with CCNT1 foci, but some condensates do not contain CCNT1. In this case, most of the large, intense transcription condensates appear to contain CCNT1 (Figure 12C). The relative intensity profile in panel E again confirms the colocalization of transcription condensates and CCNT1 clusters. This result again demonstrates that the BioID with LiTEC system successfully detects CCNT1.
[0225] However, unlike CCNT1, CDK7 does not colocalize with transcriptional condensates, even though it forms numerous foci within the nucleus (Figure 12G). From the magnified images and relative intensity profiles in panels I and J, CDK7 even appears to be excluded from large, intense transcriptional condensates. This may imply that transcriptional condensates are not associated with transcription initiation. This also accurately reflects the results of the negative control, suggesting that the previous mass spectrometry results are valid again.
[0226] CDK9 exhibits numerous foci within the nucleus (Figure 12L). Interestingly, some condensates colocalize with CDK9 foci, while others do not (Figure 12M). The area highlighted in red is shown as an example. The magnified images and relative intensity profiles in panels N and O reveal that relatively small condensates colocalize with small CDK9 foci, but CDK9 does not exhibit any foci at the locations of large condensates. Generally, small transcriptional condensates with low intensity colocalize with small CDK9 foci, whereas large transcriptional condensates with high intensity do not. It can be concluded that the size of transcriptional condensates may correlate with their functionality, although the causal relationship remains unclear. These results demonstrate that BioID using the LiTEC system can successfully detect CDK9, even when present in large amounts in small transcriptional condensates.
[0227] In conclusion, immunofluorescence of CCNT1, CDK7, and CDK9 successfully validated the mass spectrometry results. The IDRs of BRD4 may serve as basal scaffolds for different types of transcriptional condensates, which may have distinct functions in transcription.
[0228] method Standard protocols were followed (see: Celis, Cell Biology: A Laboratory Handbook, Elsevier; Green and Sambrook, Molecular Cloning, CSH Press; Helgason and Miller, Basic Cell Culture Protocols, Springer).
[0229] JPEG2026500369000006.jpg113153
[0230] JPEG2026500369000007.jpg221159JPEG2026500369000008.jpg94159JPEG2026500369000009.jpg144159JPEG2026500369000010.jpg207159JPEG2026500369000011.jpg182159JPEG2026500369000012.jpg123159
Claims
1. 1. A system for covalently modifying a target, comprising: the target is associated with endogenous intracellular condensates formed by intracellular proteins containing intrinsically disordered regions (IDRs); or 1. A system for modifying the chemical or physical properties of endogenous intracellular condensates, comprising: The system comprises: a. an IDR sequence tract containing the intrinsically disordered region; and b. an effector domain capable of modifying said target Including, wherein the IDR sequence tract and the effector domain are located on separate polypeptide molecules that can be induced to associate in response to light.
2. The system of claim 1 , wherein the IDR-containing protein is a protein located in the nucleus of a eukaryotic cell.
3. IDR-containing proteins include: BRD4, NELFA, NELFB, CDK9, P-TEFb, and the Mediator complex; RNA polymerase (RPB1); is selected from the group consisting of In particular, the IDR-containing protein is BRD4 or NELFA.
3. The system according to claim 1 or 2.
4. The system of any one of claims 1 to 3, wherein the IDR sequence tract is selected from SEQ ID NO: 001 and SEQ ID NO: 002 (ΔN-BDR4 IDR; NELFA IDR), or a sequence variant thereof characterized by at least 85% sequence identity with either SEQ ID NO: 001 or SEQ ID NO:
002.
5. The system according to any one of claims 1 to 4, wherein the target is a protein, in particular an intracellular protein, more in particular an intracellular nuclear protein, even more in particular a nuclear protein associated with RNA polymerase II activity.
6. The effector polypeptide capable of covalently modifying said target may be capable of modifying the target: a. Biotinylation; b. Ubiquitination c. methylating; d. demethylating; e. Acetylation; f. deacetylating; g. phosphorylating; h. Dephosphorylate The system according to any one of claims 1 to 5, wherein the polypeptide is selected from polypeptides capable of:
7. the association of a first fusion polypeptide comprising an IDR tract with another fusion peptide comprising an effector domain is promoted by a light-inducible binding partner pair; The light-inducible binding partner pair consists of a first binding partner and a second binding partner, wherein the first binding partner and the second binding partner associate in the presence of light; and one of the binding partners is part of a first fusion polypeptide, and the other of the binding partners is associated with an effector domain; A system according to any one of claims 1 to 6.
8. The system of claim 7, wherein the light-inducible binding partner pair is SspB and iLID (SEQ ID NOs: 003 and 004).
9. The system is: a. a first fusion polypeptide comprising an IDR sequence tract containing the intrinsically disordered region, [optionally a first fluorescent marker polypeptide], and a first member of a light-inducible binding partner pair; b. a second fusion polypeptide comprising a second member of the light-inducible binding partner pair, [optionally a second fluorescent marker polypeptide], and a binding domain capable of specifically binding to a non-endogenous peptide epitope; c. a third fusion polypeptide comprising a plurality of said non-endogenous peptide epitopes and an effector domain capable of covalently modifying said target. The system according to any one of claims 1 to 8, comprising:
10. The system of claim 9, wherein the binding domain capable of specifically binding to a non-endogenous peptide epitope is an scFv (single chain variable) antibody fragment.
11. One or more nucleic acid sequences encoding a system as defined in any one of claims 1 to 10.
12. A method for modifying a target (particularly a target protein) associated with endogenous intracellular condensates formed by an intracellular protein containing an intrinsically disordered region, comprising the steps of: a) providing a cell containing an intracellular protein that includes an intrinsically disordered region; b) expressing in said cell a system as defined in any one of claims 1 to 10; The method comprising:
13. The method comprises: c) recovering a preparation containing the target biomolecule (particularly a protein) from said cells; d) isolating the target biomolecules (particularly proteins) modified by said effector polypeptides capable of modifying the target biomolecule. The method of claim 12 further comprising:
14. The method of claim 12 or 13, wherein the effector polypeptide is capable of biotinylating a target biomolecule, and isolation of the biotinylated target biomolecule is achieved by binding the biotinylated protein from a preparation containing the biomolecule to a matrix.