Bifunctional photo-crosslinking probes for covalent capture of protein-nucleic acid complexes in cells

By using the new compound A-L1-(C)n-((L2)n-B)m to bind nucleic acid and photoreactive functional groups, and activate crosslinking with ultraviolet light, the problems of low capture efficiency and unstable crosslinking in the prior art were solved, and efficient, selective and stable protein-nucleic acid crosslinking was achieved.

CN119948169APending Publication Date: 2025-05-06UNIV OF SOUTHERN CALIFORNIA
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202380066289.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2022-07-15
Filing Date
2023-07-14
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

The prior art has problems of inefficiency, instability of cross-linking and impaired antibody activity when capturing and isolating protein-nucleic acid complexes in cells, especially in studies of low abundance transcription factor.

Method used

Using a new compound A-L1-(C)n-((L2)n-B)m, through which it is combined with nucleic acid binding functional groups and photoreactive functional groups (such as bisaziridine), the photoreactive functional groups are activated by ultraviolet light to achieve efficient crosslinking of protein-nucleic acids.

Benefits of technology

Efficient, selective and stable protein-nucleic acid complex capture is achieved, avoiding the problems of low cross-linking efficiency and impaired antibody activity in the prior art, and improving the repeatability and accuracy of experimental data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119948169A_ABST
    Figure CN119948169A_ABST
Patent Text Reader

Abstract

A new class of molecular probes is provided for truly, efficiently, stably and selectively capturing protein-nucleic acid complexes in cells. The molecular probe has a nucleic acid binding functional group and a photoreactive bis-aziridine-based functional group separated by a linker having a selected length or having a multi-arm core to create a photo-crosslinking between closely adjacent nucleic acids and proteins. This is useful for various chromatin studies, including studies of interactions between transcription factors and DNA.
Need to check novelty before this filing date? Find Prior Art

Description

CROSS-REFERENCE TO RELATED APPLICATIONS

[0001] Pursuant to 35 U.S.C. §119(e), this application claims the benefit of priority to U.S. Provisional Patent Application No. 63 / 389,580, filed on July 15, 2022, the contents of which are incorporated herein by reference in their entirety. References to sequence listings

[0002] This application contains a sequence listing submitted as an electronic xml file, named "065715-000130WOPT", created on July 11, 2023, and 13,066 bytes in size. The information contained in this electronic file is hereby incorporated by reference in its entirety. Technical Field

[0003] The present invention relates to functional small molecules for proximity ligation to identify and / or label protein-nucleic acid complexes. Background Art

[0004] Protein-nucleic acid interactions are the basis of a wide range of cellular processes from genomic DNA replication, repair and transcription to RNA processing, translation and regulation. Nucleic acids (e.g., cytoplasmic DNA and viral RNA) also regulate cell signaling pathways involved in immune responses, aging, and a variety of human diseases. The main challenge in studying protein-nucleic acid interactions in situ is to capture and separate protein-nucleic acid complexes within cells, because most of these non-covalent complexes are dynamic and dissociate during the separation process. As part of the effort to address this problem, a variety of techniques have been developed to capture protein-nucleic acid complexes in cells, including direct UVC (254nm) crosslinking between RNA and RNA-binding proteins or using UVA (365nm) to crosslink RNA metabolically labeled with 4-thio-uridine (4SU) or 6-thio-guanosine (6SG). Although these methods have greatly facilitated the study of protein / RNA interactions, low crosslinking efficiency, the use of short-wavelength UVC (which damages proteins, DNA, and RNA), or the need to use 4SU and 6SG for metabolic labeling remain major limitations.

[0005] Chromatin immunoprecipitation followed by sequencing (ChIP-Seq) has been widely used in the study of interactions between transcription factors (TFs) and DNA. However, the major limitations of current methods have been increasingly recognized, especially for low-abundance TFs; and in traditional protocols, the cell fixation step by using formaldehyde can be the main cause of data irreproducibility, as it negatively affects the activity of antibodies and other aspects of ChIP-based experiments (such as ChIP-seq, Hi-C, or HiChIP). A large number of replicate ChIP-seq datasets in the ENCODE database have low correlation (r~0.5-0.6). Another major issue is that a large fraction (45%-80%) of the detected DNA sequences lack the expected binding motifs, which raises the question of whether the DNA fragments associate with TF targets through indirect mechanisms or simply due to nonspecific capture (see Figure 1A ). These limitations severely undermine the usefulness of ChIP-seq for mechanistic studies, such as analyzing the functional impact of genetic variants in TF binding sites. Part of the problem is due to the instability and limited quality of antibodies, especially for transcription factors. It is also becoming increasingly clear that another step in the protocol, fixation of cells with formaldehyde, can be a major cause of irreproducibility in ChIP-based experimental data.

[0006] The widespread use of formaldehyde as a crosslinking agent in the above techniques is based on the long-held but poorly understood belief that formaldehyde can crosslink proteins to DNA. Formaldehyde can crosslink two primary amine groups via a Schiff base intermediate, forming a methylene bridge between two spatially adjacent lysine residues. The reaction is very facile, which explains the high efficiency of formaldehyde-based crosslinking of protein complexes. However, crosslinking between proteins and DNA is quite different. The exocyclic amino groups in nucleic acids are inefficient nucleophiles due to delocalized binding to the aromatic ring system of the nucleoside bases ( Figure 1E). Although cross-linking products between certain amino acids and short oligonucleotides were observed by mass spectrometry under extreme conditions, the yields were very low and the cross-linking products were unstable. These observations raise the question of whether formaldehyde can indeed directly cross-link proteins and DNA into stable complexes. Earlier studies have shown that formaldehyde is unable to cross-link purified transcription / DNA complexes in vitro, although it is highly efficient in cross-linking higher-order chromatin complexes in vivo. On the other hand, certain transcription factors, such as NF-kB, STAT3, and fly insulator factor Elba, form ring structures or higher-order complexes wrapped around DNA, which can be stably cross-linked to DNA when secondary protein cross-linkers such as disuccinimidyl suberate (DSS) are used. These observations suggest that the surface success of formaldehyde-based cross-linking of protein-DNA complexes in cells is unlikely to be due to a direct cross-linking reaction between the two, but is achieved through the cross-linking of protein complexes that capture DNA. This mechanism of capturing DNA by cross-linking protein complexes may be the main source of signal noise in current CHIP-seq data ( Figure 1A ). In addition, formaldehyde cross-linking, an empirical and poorly characterized step, can lead to many other problems described in the literature.

[0007] First, direct modification of transcription factors by formaldehyde at the DNA binding surface (which is usually rich in lysine residues) may result in the inability to capture protein-DNA complexes (see Figure 5G ), especially for TFs that bind DNA highly dynamically, or may even cause artifacts. Second, formaldehyde cross-linking may affect the activity of the antibodies used to capture the TF, either directly by modifying residues in the epitope or indirectly by locking the protein structure making the epitope inaccessible. Therefore, different cell fixation conditions (formaldehyde concentration and reaction time) may exacerbate the variability of antibody reactivity. These problems may be more significant in the study of low-abundance TFs (relative to histones), where high concentrations of formaldehyde fixation may be required to capture sufficient protein / DNA complexes. Overall, the highly reactive but non-specific modification of proteins by formaldehyde and the inefficient cross-linking of DNA are the main limitations of current ChIP-based techniques.

[0008] Various reports from different laboratories have recognized that formaldehyde is unable to crosslink certain proteins to nucleic acids, even when they are in close proximity within the nucleus. The main reason behind this lies in the crosslinking chemistry required by formaldehyde: an active nucleophilic attack from one side of the crosslinked target on the carbon atom in the Schiff base imine that has been formed by crosslinking from the other side of the target. Although amine groups are nucleophiles in general, those exocyclic amine groups from nucleic acids (e.g., DNA) are almost inactive because the electrons on their nitrogen are somewhat restricted by delocalized conjugation from the nearby aromatic base ring system. This makes nucleic acids inherently inactive nucleophiles to be crosslinked by formaldehyde-mediated crosslinking in these ChIP-based methods ( Figure 1E ) directly cross-links proteins. In contrast, the locus-specific information extracted from these techniques is, to some extent, a rough estimate / simulation of the results of nearby protein-protein (e.g., histone-transcription factor, since histones bind extensively to DNA) cross-links. Therefore, directly modifying transcription factors at the DNA binding surface with formaldehyde will result in the inability to capture protein-DNA complexes, especially for transcription factors that bind DNA highly dynamically. Therefore, these results obtained with formaldehyde are biased from reality and may even lead to unwanted artifacts, since it is the histones rather than the DNA itself that are cross-linked to nearby proteins. Secondly, due to the large number of proteins in the cell nucleus, the good cross-linking ability of formaldehyde to these proteins often leads to high and undesirable noise, even if antibody extraction is performed, because the desired target of the antibody is usually cross-linked to multiple other nearby proteins. And due to this abundance of proteins relative to DNA in the cell nucleus, these protein-protein cross-linking noises are often so high that certain regions on chromosomes are off-limits for formaldehyde-mediated ChIP ( Figure 1E ). Third, because the regional physical and chemical characteristics of various proteins within the nucleus are very heterogeneous, the differential reactivity of formaldehyde to different proteins, protein complexes, and various subcellular regions and structures has been suggested to be the source of many problems encountered in formaldehyde-based protocols. This often results in certain proteins or protein complexes being more or less cross-linked by formaldehyde (e.g., due to differences in the amount and structural distribution of surface lysine residues), making certain genomic regions over- or under-sampled in ChIP-based assays ( Figure 1E Finally, formaldehyde cross-linking may affect the activity of the antibody used to capture TF, either directly by modifying residues in the epitope or indirectly by locking the protein structure and making the epitope inaccessible ( Figure 1E ). Therefore, although ChIP-based methods can have certain reproducible results, true and precise protein-nucleic acid cross-linking remains a challenge for formaldehyde-based cross-linkers.

[0009] In addition to formaldehyde, UV irradiation (e.g., short-wavelength UVC) has also been used to crosslink proteins to RNA and DNA. UV crosslinking can produce stable covalent products that can be digested by proteases and nucleases to generate peptide / oligonucleotide conjugates for subsequent mass spectrometry analysis (XL-MS). Although this represents a promising method for mapping protein-nucleic acid interactions in vitro and in cells, the main disadvantages of these means are the short-wavelength UVC (~250nm) and low crosslinking efficiency required to induce crosslinks between native proteins and nucleic acids. Short wavelengths (E=hc / wavelength, where hc is a constant) bring high UV doses (continuous or pulsed UV lasers), and increasing crosslinking yields under these conditions can lead to extensive UV damage to both proteins and DNA or RNA.

[0010] Given the significant drawbacks of formaldehyde and UV cross-linking discussed above, the continued use of such cross-linking agents remains challenging.

[0011] It is therefore an object of the present invention to provide new compounds and systems for capturing protein-DNA complexes in cells with high efficiency, selectivity (i.e., targeting only DNA binding proteins) and stability (so as to enable robust isolation of protein-DNA complexes for subsequent analysis by DNA sequencing or protein profiling).

[0012] All publications herein are incorporated by reference to the same extent as if each individual publication or patent application was specifically and individually indicated to be incorporated by reference. The following description includes information that may be helpful in understanding the present invention. It is not an admission that any of the information provided herein is prior art or relevant to the presently claimed invention, nor is it an admission that any publication specifically or implicitly referenced is prior art. Summary of the invention

[0013] In various embodiments, the present invention provides compounds of formula (I): A-L1-(C) n -((L2) n -B) mFormula (I), wherein: A represents a nucleic acid binding functional group derived from psoralen, methyltrioxsalen, benzophenone, 4',6-diamidino-2-phenylindole (DAPI), Hoechst dye, polyamide or G quartet binding molecule, ethopropanol or their derivatives; L1 is absent or represents the first linker; when n=1, C represents a core part having at least two functional groups, each functional group is used to attach to L1 and to at least one arm represented by L2-B respectively; or when n=0, C is absent; when n=1, L2 independently represents the second linker of each arm represented by L2-B; or when n=0, L2 is absent; for each of the arms, B independently represents: a photoreactive functional group, the photoreactive functional group includes diazirine or a derivative thereof or an aryl azide or a derivative thereof, optionally the aryl azide or a derivative thereof is selected from phenyl azide, o-hydroxyphenyl azide, m-hydroxyphenyl azide, tetrafluorophenyl azide, o-nitrophenyl azide, m-nitrophenyl azide or azidomethylcoumarin; or a detectable functional group; wherein, in at least one of the arms, B represents a photoreactive functional group; n=0 or 1; m represents (L2) n -B represents the number of arms, wherein when n=1, m is an integer of 1 or more, or when n=0, m=1.

[0014] In some embodiments, at least one of L1 and L2 is not absent, and the at least one of L1 and L2 is cleavable.

[0015] In some embodiments, L1, L2, or both independently include one or more of a sulfoxide-containing mass spectrometry (MS) cleavable bond, an acid-cleavable CS bond, a disulfide group, and an azo group.

[0016] In some embodiments, n=0, m=1, and the compound is represented by Formula (II): A-L1-B Formula (II), wherein L1 is absent or is the first linker.

[0017] In some embodiments, A is an amine-containing or amine-reactive derivative of psoralen, an amine-containing or amine-reactive derivative of methyltrimethoxalen, an amine-containing or amine-reactive derivative of benzophenone, an amine-containing or amine-reactive derivative of 4',6-diamidino-2-phenylindole (DAPI), an amine-containing or amine-reactive derivative of Hoechst dye, an amine-containing or amine-reactive derivative of polyamide, or an amine-containing or amine-reactive derivative of a G-quadruplex binding molecule, or an amine-containing or amine-reactive derivative of ethoxybutyraldehyde, optionally, A is derived from succinimidyl-[4- (psoralen-8-yloxy)]-butyrate (SPB) or 4'-aminomethyltrimethsalen (4AMT); B includes diazirine or diazirine alkyne, optionally aminodiazirine alkyne (AAD); and L1 is absent or is a first linker, wherein the first linker comprises one or more of the following: (i) a cleavable bond, (ii) an oligomer or polymer having -OCH2CH2- repeating units, and (iii) an unsaturated portion, the unsaturated portion being optionally selected from a carbon-carbon double bond, a carbon-carbon triple bond or an aromatic group.

[0018] In some embodiments, L1-B is derived from succinimidyl 6-(4,4'-azidopentanamido)hexanoate (NHS-LC-SDA), succinimidyl 2-((4,4'-azidopentanamido)ethyl)-1,3'dithiopropionate (NHS-SS-diaziridine) or 2-(3-(but-3-yn-1-yl)-3H-diaziridine-3-yl)ethane-1-amine (AAD); and / or wherein A is derived from 4'-aminomethyltrimethsalin (4AMT) or succinimidyl-[4-(psoralen-8-yloxy)]-butyrate (SPB); and wherein optionally, the photocrosslinking molecule is represented by formula (IIa) or formula (IIc):

[0019] In some embodiments, A is derived from succinimidyl-[4-(psoralen-8-yloxy)]-butyrate (SPB) or 4'-aminomethyltrimethalin (4AMT); B comprises diazirine or diazirine alkyne, optionally aminodiazirine alkyne (AAD); and L1 is a first linker comprising one or more of: (i) a cleavable bond, (ii) an oligomer or polymer having -OCH2CH2- repeating units, and / or (iii) an unsaturated portion, the unsaturated portion being optionally selected from a carbon-carbon double bond or an aromatic group; and wherein optionally, the photocrosslinking molecule is represented by formula (IIb), formula (IId), formula (IIe), or formula (IIf):

[0020] In some embodiments, A is selected from the group consisting of: in: R 1 are independently H, halo, OH, optionally substituted alkoxy or optionally substituted alkyl; R 2 are independently H, halo, OH, optionally substituted alkoxy or optionally substituted alkyl; a is 0, 1, 2, 3, 4 or 5; and b is 0, 1, 2, 3, or 4; in: R 3 are independently H, halo, OH, optionally substituted alkoxy or optionally substituted alkyl; R 4 is H, halo, OH, optionally substituted alkoxy or optionally substituted alkyl; R 5 are independently H, halo, OH, optionally substituted alkoxy or optionally substituted alkyl; c is 0, 1, 2, 3 or 4; and d is 0, 1, 2, 3 or 4; in: R 6 is H, halo, OH, optionally substituted alkoxy or optionally substituted alkyl; R 7 is H, halo, OH, optionally substituted alkoxy or optionally substituted alkyl; R 8 is H, halo, OH, optionally substituted alkoxy or optionally substituted alkyl; and R 9 is H, halo, OH, optionally substituted alkoxy or optionally substituted alkyl; in: R 10 is H, halo, OH, optionally substituted alkoxy or optionally substituted alkyl; R 11 is H, halo, OH, optionally substituted alkoxy or optionally substituted alkyl; and R 12 is H, halo, OH, optionally substituted alkoxy or optionally substituted alkyl; L1 is absent or L is selected from the group consisting of: in: q is 0, 1, 2, 3, or 4; in: p is 0, 1, 2, 3 or 4; in: R 13 is independently H, halo, OH, optionally substituted alkoxy or optionally substituted alkyl; and e is 0, 1, 2, 3 or 4; in: r is 0, 1, 2, 3, or 4; in: s is 0, 1, 2, 3, or 4; in: t is 0, 1, 2, 3, or 4; u is 0, 1, 2, 3, or 4; and B is selected from the group consisting of:

[0021] In some embodiments, A is selected from the group consisting of: L1 is absent or L1 is selected from the group consisting of: and B is selected from the group consisting of:

[0022] In some embodiments, the compound is:

[0023] In some embodiments, L1 comprises a length of 2 to 20 carbons or 20-100 carbons.

[0024] In some embodiments, n=1, m is an integer of 2 or greater, and C represents a core portion having at least three functional groups, each functional group being used to attach to L1 and to attach to at least two arms each represented by (L2-B), such that the compound is represented by formula (III):

[0025] In some embodiments, in one of the at least two arms, B comprises diaziridine or diaziridine azide, and in the other of the at least two arms, B represents a detectable functional group, and the detectable functional group comprises a fluorophore, biotin, a chromophore, a chromogen, a quantum dot, a fluorescent microsphere or a nanoparticle. In some embodiments, L1, L2 or both independently comprise one or more of the following: (i) a cleavable bond, (ii) an oligomer or polymer having -OCH2CH2- repeating units, and (iii) an unsaturated portion. In some embodiments, C represents a dendritic core portion, and the dendritic core portion comprises at least three surface functional groups, each of which is used to attach to L1 and to the at least two arms represented by each free L2-B. In some embodiments, L1, L2 or both independently comprise a triazole bonded to A.

[0026] In various embodiments, the present invention provides a method for crosslinking nucleic acids and proteins in a system, the method comprising: providing a compound of the present invention; providing a system, wherein the system comprises nucleic acids and proteins; contacting the compound with the system; and irradiating the system and the compound with ultraviolet light under conditions effective to crosslink the nucleic acids and the proteins. In some embodiments, the system is a living cell. In some embodiments, the wavelength of the ultraviolet light is between 300nm and 370nm. In some embodiments, the method further comprises using the system for immunoprecipitation, DNA and RNA complex extraction using organic solvents, chromatographic separation, chromatin precipitation, 3D chromatin conformation capture, mass spectrometry, and electrophoresis. In some embodiments, the elements L1, L2, or both of the compound are independently cleavable, and the method further comprises adding a cleavage agent to the system to cleave the elements L1, L2, or both; or wherein the element A of the compound is derived from psoralen, and the method further comprises applying ultraviolet light of a wavelength of about 230nm to cleave the element A; thereby generating a fingerprint of crosslinked proteins adjacent to nucleic acids in the system.

[0027] In some embodiments, the present invention provides a method for preparing a compound of formula (III), comprising: providing an azide derivative of a nucleic acid-binding photoreactive reagent, the reagent comprising psoralen, methyltrimethoxalin, benzophenone, 4',6-diamidino-2-phenylindole (DAPI), Hoechst dye, polyamide or G quadruplex binding molecule, ethobutanone or a derivative thereof; providing an azide derivative of a photoreactive reagent comprising a diazirine moiety to obtain an azide-diazirine bifunctional photoreactive reagent, and the photoreactive reagent optionally further comprises an alkyne group, or providing an aryl azide, The aryl azide is optionally selected from phenyl azide, o-hydroxyphenyl azide, m-hydroxyphenyl azide, tetrafluorophenyl azide, o-nitrophenyl azide, m-nitrophenyl azide or azido-methylcoumarin; optionally, an azide derivative of a detectable reagent is provided, the detectable reagent comprising a fluorophore, biotin, a chromophore, a chromogen, a quantum dot, a fluorescent microsphere or a nanoparticle; a multi-arm reagent having at least three functional groups is provided, each functional group independently comprising an alkyne; and each azide derivative and the aryl azide (if provided) are combined with the multi-arm reagent in a reaction vessel to prepare the compound.

[0028] In some embodiments, the multi-arm reagent has at least three functional groups, each of which independently comprises a cyclooctyne group. In some embodiments, the nucleic acid-binding photoreactive reagent comprises a first primary amine functional group, and providing an azide derivative of the nucleic acid-binding photoreactive reagent comprises converting the first primary amine functional group into a first azide-containing moiety, optionally by reacting the nucleic acid-binding photoreactive reagent with imidazole-1-sulfonyl azide; and / or wherein the photoreactive reagent comprising a diazirine moiety further comprises a second primary amine functional group or is modified with a second primary amino functional group, and providing an azide derivative of the photoreactive reagent comprises converting the second primary amine functional group into a second azide-containing moiety, optionally by reacting the photoreactive reagent with imidazole-1-sulfonyl azide. BRIEF DESCRIPTION OF THE DRAWINGS

[0029] Exemplary embodiments are shown in the referenced drawings.The embodiments and drawings disclosed herein are intended to be considered illustrative rather than restrictive.

[0030] Figure 1A-Figure 1C Depicts that formaldehyde cross-linking of proteins in the cell nucleus traps large-scale protein-DNA complexes in which many DNA and proteins are associated together. ( Figure 1A ) is not the capture of the targeted transcription factor (TF, blue oval) by the selected antibody (Ab, inverted Y shape) and its binding site (TFBS, blue bar); ( Figure 1B) Antibody pull down drags in many other proteins (multiple colors and shapes) and nonspecific sites (NS, yellow bars); ( Figure 1C ) ChIP-seq traces contain a high percentage of false signals.

[0031] Figure 1D Describes the overall design of the BFPX strategy invented in this article.

[0032] Figure 1E It is molecularly elaborated why formaldehyde is an undesirable crosslinking agent for protein-nucleic acid crosslinking. The main reason is because amines from nucleic acids (e.g., DNA) are not active nucleophiles, thus making the critical attack on the imine intermediate low yield and leading to capture or crosslinking failure or inaccuracy. Other disadvantages include high noise from other proteins in the nucleus, different protein crosslinking effects, histone-mediated inaccuracies or artifacts, protein epitope masking or distortion.

[0033] Figure 2A-2C is a schematic diagram of an exemplary probe design. The probe is "bifunctional" because on the one hand it endows nucleic acids with true free primary amine groups through the photoreactive intercalator psoralen, and on the other hand we equip it with a spectrally photoreactive crosslinker diaziridine group through this amine linkage, which can crosslink nearby biomolecules to form carbenes upon photoactivation. Figure 2A Depicted is an example of a psoralen-based BFPX probe, where the DNA binding / crosslinking head is a psoralen derivative (4'-aminomethyltrimethylsalen, 4AMT) and the protein crosslinking head is a diaziridine group. The dashed box is the linker region, which can be synthetically engineered to have a variety of lengths, cleavage properties, or functional groups (indicated by R) for enrichment and / or fluorescent labeling of crosslinked protein-DNA complexes. Figure 2B A schematic diagram of an exemplary probe design is depicted. Figure 2C A schematic diagram of an exemplary probe design is depicted.

[0034] Figure 3A and Figure 3BDepicted is an exemplary method for extending bifunctional probes by a multi-arm PEG DBCO copper-free clickable core to obtain multifunctional probes. The PEG DBCO clickable arms can be increased to more than the 4 shown here. Multiple protein photocrosslinkable ends can also be clicked into the core for capturing multiple nearby proteins. The PEG arm length can be selected. The nucleic acid binding photocrosslinkable ends are based on psoralen, benzophenone, DAPI, polyamide or G-quadruplex binding molecules, each of which is derivatized with an amine group; and by imidazole-1-sulfonyl azide, the amine-based nucleic acid binding photocrosslinkable ends are converted to azide-based (azido) nucleic acid binding photocrosslinkable ends. Azide-based compounds (e.g., azide-based nucleic acid binding photocrosslinkable ends; azide fluorophores; azobiotin-azide; and azide-based protein photocrosslinkable ends) can all be conjugated to a multi-arm core (e.g., PEG-DBCO) by copper-free click chemistry.

[0035] Figure 3C Depicted is an exemplary synthetic route for forming an exemplary BFPX probe using a DAPI molecule as a DNA binding head using a cleavable linker that is cleavable during mass spectrometry analysis due to collision-induced dissociation (CID). Through these designed functions, the molecule not only binds to the DNA double strand, but also contains (1) a fluorescent localizer; (2) an MS CID (collision-induced dissociation) cleavable fingerprint, which can be introduced by an azide-labeled acid-cleavable disuccinimidyl bissulfoxide (azide-A-DSBSO; bis(2,5-dioxopyrrolidin-1-yl) 3,3′-((2-(3-azidopropyl)-2-methyl-1,3-dioxane-5,5-diyl) bis(methylenesulfinyl)) dipropionate; a MS-cleavable cross-linker for studying protein-protein interactions) to facilitate MS analysis using the dissociated fraction; and (3) an SS (disulfide)-cleavable extraction biotin handle, which simplifies the enrichment process by streptavidin beads. Azide-A-DSBSO has two N-hydroxysuccinimide (NHS) ester groups for targeting amines, spacer length, two symmetrical acid-cleavable CS bonds and a central bioorthogonal azide tag; the cleaved spacer (after CS bond cleavage) generates labeled peptides for unambiguous identification by collision-induced dissociation in tandem MS.

[0036] Figure 3D Depicted are exemplary unsaturated moieties as linkers and exemplary synthetic routes thereof in forming BFPX probes.

[0037] Figure 3ESchematic diagram depicting the structure-based design of BFPX probes using DAPI or Hoechst 33258 as DNA-binding heads. Upper panel: Left, crystal structure of DAPI bound to double-stranded DNA, right, the DNA-binding face of DAPI (orange shaded area) should be avoided in synthetic modifications; blue arrows indicate potential sites for the introduction of linkers with protein-capturing groups (e.g., diazirine heads), and green arrows indicate potential sites for the introduction of photoaffinity labeling groups that crosslink to DNA, as these positions of DAPI are close to DNA. Lower panel: Left, crystal structure of Hoechst 33258 bound to double-stranded DNA, right, the DNA-binding face of Hoechst 33258 (orange shaded area) should be avoided in synthetic modifications; blue arrows indicate potential sites for the introduction of linkers with protein-capturing groups (e.g., diazirine heads), and green arrows indicate potential sites for the introduction of photoaffinity labeling groups that crosslink to DNA, as these positions are close to DNA.

[0038] Figure 4 The chemical structures of four exemplary psoralen-based BFPX probes synthesized in preliminary studies in the Examples section are shown. Figure 5A-Figure 5B , Figure 6A-6B and Figure 8A-B Exemplary synthetic routes and mass spectrometric confirmation of the synthetic products for each of 4AMT-LC-SDA, 4AMT-SDAD (also denoted as 4AMT-LC-SDAD), and SPB-PEG3-AAD are shown in FIG.

[0039] Figure 5A and Figure 5B Probe synthesis and confirmation by mass spectrometry (MS) are depicted. The probe has a flexible linkage between and an exemplary final probe molecule can be easily synthesized by amine-N-hydroxysuccinimide ester (NHS) conjugation chemistry, for example, the reaction of 4'-aminomethyltrimethylsalen (4-AMT) and NHS-LC-diaziridine (NHS-LC-SDA, succinimidyl 6-(4,4'-azidopentanamido)hexanoate) under one-step mild conditions with relatively good yields estimated by mass spectrometry. Figure 5A In the figure, the green box surrounds the DNA binding head and the red box surrounds the protein binding head. Figure 5B The main peak with good yield in the reaction mixture shows multiple reaction settings +Na + The final compound in the form of (480+23) or +Na + The mass of the bis-final compound (960+23) in the form (indicated by the arrow).

[0040] Figure 5C-5EGel electrophoresis results are depicted, demonstrating that probes from various signaling pathways (DNA, protein, larger pore size) successfully and efficiently and selectively cross-linked exemplary double-stranded DNA and DNA binding proteins.

[0041] Fig. 5F Depicted is the electrophoretic analysis of 4AMT-LC-SDA in suspension with GM12878 cell line with different sonication times compared to formaldehyde and UV alone in suspension, demonstrating the overall expected efficiency of this probe for in vivo application.

[0042] Figure 5G Describe the non-denaturing (native) EMSA assay of the effects of formaldehyde (FA), BFPX probe (4AMT-LC-SDA) and UVA (365nm) on the DNA binding of NFAT1. All binding reactions contain 10 μM 27merdsDNA (labeled with 5'6-FAM) with NFAT binding sites, as well as various combinations of NFAT1 protein (14 μM), FA (1% v / v), 4AMT-LC-SDA (20 μM), with or without UV irradiation (365nm, LED, 30 watts, sample distance 2cm, exposure time 15min at 4°C). 10 μL binding reactions are run on 10% PAGE 0.5xTBE gels. The gel is visualized by FAM fluorescence to DNA (top), and by Coomassie brilliant blue staining to protein (bottom). The higher diffuse bands in lanes 3, 8, 9 and 10 are likely to be non-specific protein complexes, which usually appear when excess protein is used to ensure that all DNA is bound.

[0043] Figure 5H Denaturing SDS gel analysis of covalent capture of in vitro assembled NFAT1 / DNA complexes by BFPX probes 4AMT-LC-SDA (lanes 1, 2, and 3 are indicated as "LC-SDA") and 4AMT-LC-SDAD (lanes 4, 5, and 6 are indicated as "LC-SDAD") is depicted. All reactions contained 10 μM 27mer dsDNA with NFAT1 binding sites (labeled with 5'6-FAM), as well as 14 μM NFAT1 and 25 μM BFPX probes (4AMT-LC-SDA, lanes 1 to 3; or 4AMT-LC-SDAD, lanes 4 to 6). Probe binding and UV crosslinking were performed as shown. Figure 5G and Figure 7 The cross-linked complexes (lanes 3 and 6) were further treated with DTT (100 mM, heated at 50°C for 30 min). The reaction mixtures were run on SDS gels and visualized by FAM fluorescence for DNA (top) and Coomassie brilliant blue staining for protein (bottom). Figure 7shown.

[0044] Fig. 6A Depicted is the synthesis of the cleavable BFPX probe 4AMT-SDAD, which involves coupling between 4AMT and NHS-SS-diaziridine.

[0045] Figure 6B This is the ESI mass spectrum confirming the synthesis of 4AMT-SDAD.

[0046] Figure 7 Depicted are denaturing SDS gel analysis of covalent capture of in vitro assembled MEF2 / DNA and NFAT1 / DNA complexes by the BFPX probe SPB-AAD. Reactions in lanes 1-7 contained 10 μM 44mer dsDNA with a MEF2 binding site (labeled with 5'6-FAM) and 40 μM MEF2A; reactions in lanes 8-12 contained 10 μM 27mer dsDNA with a NFAT1 binding site (labeled with 5'6-FAM) and 14 μM NFAT1; protein / DNA complexes were incubated at room temperature for 30 min and treated with various combinations of BFPX probe (SPB-AAD) and UV irradiation, as shown. Figure 5G As shown. The SPB-AAD concentrations in lanes 3, 4, 5, 6 and 7 and lanes 8, 9, 10, 11, 12 were 125 μM, 125 μM, 12.5 μM, 1.25 μM, and 0.125 μM, respectively. SDS loading dye (2% SDS) was added to the reaction mixture, boiled for 10 min, and then run on an SDS gel (Bio-Rad Any Kd gradient gel). The gel was visualized by FAM fluorescence for DNA (top) and Coomassie brilliant blue staining for protein (bottom).

[0047] Fig. 8A The synthesis of the BFPX probe SPB-PEG3-AAD is depicted.

[0048] Figure 8B This is the ESI mass spectrum confirming the synthesis of SPB-PEG3-AAD.

[0049] Figure 8C Depicted is a denaturing SDS gel analysis of covalent capture of in vitro assembled MEF2 / DNA and NFAT1 / DNA complexes by the BFPX probe SPB-PEG3-AAD. The reactions of MEF2 / DNA (lanes 1-4) and NFAT1 / DNA (lanes 5-8) are described in Figure 7The concentrations of SPB-PEG3-AAD in lanes 1, 2, 3, and 4 and lanes 5, 6, 7, and 8 were 100 μM, 100 μM, 1 μM, respectively. The reaction mixture was run on an SDS gel and DNA was visualized by FAM fluorescence (top) and protein was visualized by Coomassie Brilliant Blue staining (bottom), as shown in Figure 7 Described in .

[0050] Figure 9A-9D Depicted is the characterization of BFPX-mediated cross-linking reactions to identify covalent attachment sites on proteins. Fig. 9A ) was carried out corresponding to Figure 7 Lane 11 of the large-scale cross-linking reaction. The reaction mixture was separated on FPLC using a mono-Q column (yellow trace: conductivity, red trace: salt gradient (A: 10 mM Hepes pH 7.4; B: 10 mM Hepes 7.4, 1 M NaCl), blue trace: UV254 nm absorbance). Fig. 9B ) SDS PAGE analysis of the reaction and purification: lane 1, 5% reaction without UV 365 nm irradiation, lanes 2 and 3, two different batches of the cross-linking reaction, after UV365 nm cross-linking, in addition to free NFAT and DNA, larger complexes (putatively covalent NFAT / DNA complexes) appeared in lanes 2 and 3, which can be observed by fluorescence imaging (FAM, top) and CCB-G250 staining (protein staining, bottom) (interestingly, free NFAT can also be observed by fluorescence imaging). The flow-through contains only free NFAT (lane 4), peak I mainly contains complexes (lane 5), and peak II mainly contains free DNA (lane 6). ( Fig. 9C ) The purified NFAT / DNA complex was digested with trypsin, and the DNA-peptide conjugate was purified and sequenced by Edman degradation: Upper panel, schematic diagram of the DNA binding domain of human NFAT1 used in this study (residues 399-676), sequence of loop 478-491 (SEQ ID NO: 1) as shown above; Neutron image: Schematic diagram of the covalently captured NFAT / DNA complex, showing the DNA sequence used in this study ( (SEQ ID NO: 2) and its complementary strand (SEQ ID NO:3), where two potential probe (yellow wedge and star) binding sites (TpA) are marked. The underlined region is the NFAT binding motif. Lower sub-figure: Edman sequencing results (I479, T480 and G481) of the loop 478-491 (SEQ ID NO: 1) was identified as the attachment site. Fig.9DThe biochemical results of )(ac) fully match the structural model, in which NFAT binds to its consensus GGAAAA motif underlined in SEQ ID NO:2, and the BFPX probe inserts into the adjacent TpA site.

[0051] Fig. 10A Depicts the presence of MEF2A and DNA complex in GM12878 cells treated with UV365 and SPB-PEG3-AAD Western blot analysis. For each lane, one million GM12878 cells were used. Cell nuclei were extracted and resuspended in 500 μL PBS and treated with various combinations of SPB-PEG3-AAD (10 μM) and / or UV365 (10 min, 30 W, LED, sample distance 3 cm, cooled at 4 ° C). The samples were then sonicated (1 min) using a covariate and then subjected to limited MNase digestion (10 min, 37 ° C). The cleaved and partially digested samples were then mixed with SDS loading dye (2% SDS, boiled for 10 min) and run on SDS gels, using anti-MEF2 antibody (B-4) (SCBT, cat#sc-17785) to perform Western blot analysis on MEF2A.

[0052] Fig. 10B Western blot analysis of Hela cells transfected with FOXP3 labeled with AVI-TEV-FLAG is depicted. For each lane, approximately one million cells were used. Lane 1 was fixed with 1% formaldehyde (FA) for 10 min according to standard protocols. Lane 2, untreated control, lane 3, UV only, lanes 4 and 6 were treated with 1 μM or 10 μM of SPB-AAD, respectively, but without UV irradiation, and lanes 5 and 7 were treated with 1 μM or 10 μM of SPB-AAD, respectively, and irradiated with UV365 (10 min, 30 W, LED, sample distance 3 cm, cooled at 4 ° C). Cells were lysed using RIPA buffer and further digested by MNase. The sample mixture was mixed with SDS loading dye (2% SDS, boiled for 10 min) and run on SDS gels for Western blot analysis using anti-FLAG antibodies.

[0053] Figures 11A-11F Depicts the permeability and subcellular distribution of the BFPX probe SPB-AAD in Hela cells. Top panel ( Figures 11A-11C ): Buffer control experiment without adding SPB-AAD, ( Fig.11A ) bright field; Fig. 11B ) AlexaFluor fluorescence imaging; ( Fig. 11C ) DAPI staining of cell nuclei; bottom panel: treated with 25SPB-AAD, ( Fig.11D ) bright field; Fig.11E ) Alexa Fluor fluorescence imaging; Fig.11F ) DAPI staining of cell nuclei.

[0054] Fig. 12A The general design of the bifunctional photocross-linking (BFPX) strategy is depicted.

[0055] Fig. 12B Depicted are the psoralen-based BFPX probes synthesized and characterized in this study. Details of compound synthesis and associated spectroscopic validation are provided in the Examples section herein.

[0056] Fig. 12C A non-denaturing EMSA assay depicting the effects of varying concentrations of formaldehyde (FA), BFPX probes (SPB-AAD, SPB-PEG4-AAD, and SPB-spermidine-AD), and UVA (365 nm) on DNA binding of NFAT1. Binding reactions were performed in a buffer of 20 mM Hepes pH 7.6, 150 mM NaCl, 1 mM DTT, 12% glycerol. Protein and DNA were typically mixed to bind for 30 min at room temperature, followed by the addition of the BFPX probe for an additional 15 min, and then UV irradiation. In the case of FA treatment, the reaction time was 10 min at room temperature, followed by quenching with 250 mM glycine. When included (indicated by a positive sign), the binding reaction contained 10 μM of a 27mer dsDNA (labeled with Cy5) with a NFAT binding site (antigen receptor response element 2, ARRE2, from the IL-2 promoter, (SEQ ID NO:2), complementary strand unlabeled), 24 μM DNA binding domain (DBD) of human NFAT1 protein (residues 399-676, NFAT1-DBD). UV irradiation (indicated by a positive sign) was performed using an LED chip (21.4 mm x 21.4 mm, 365-370 nm, 30 watts, sample distance 2 cm, exposure time 2-15 min at 4°C, here exposure time 15 min). Half of the binding reaction (7.5 μL) was run on a 10% PAGE 0.5xTBE gel and visualized by Cy5 fluorescence on DNA. A small amount of single-stranded DNA is due to incomplete annealing. The higher diffuse band is likely to be a non-specific protein complex. None of these experimental defects affect the main conclusions of this experiment, because the double-stranded DNA band and the specific NFAT-DNA complex band (the sharp band in the middle) are well defined and can be monitored without interference from other bands.

[0057] Fig.12D The other half of the binding reaction from the above experiment is depicted ( Fig. 12C) SDS loading dye (2% SDS) was added, boiled for 10 min, and run on SDS gel (4%-15% gradient gel). The gel was visualized by Cy5 fluorescence for DNA.

[0058] Fig.12E Depicted is the further characterization of BFPX crosslinking of NFAT1 / DNA complexes using different DNA substrates and staining techniques. Two unlabeled DNA substrates containing the core NFAT binding site and different flanking bases and sequence lengths were used (c1: (SEQ ID NO:4), 21mer; c2: (SEQ ID NO: 5), 16mer) to bind NFAT1-DBD, such as Fig. 12C The binding reactions were then subjected to BFPX crosslinking using SPB-AAD (100 μM) with (lanes 2, 4) and without (lanes 1 and 3) UV irradiation. The reactions were then analyzed by SDS PAGE under denaturing conditions, as Fig.12D Described. Because DNA is unlabeled, gel is stained using dye (SYBR Safe) or Coomassie brilliant blue dye (CB) based on anthocyanin. As shown in the upper sub-figure, NFAT is cross-linked with c1 DNA (lane 2) and c2 DNA (lane 4) in the presence of UV, but is not cross-linked (lane 1 and lane 3, respectively) in the absence of UV irradiation. Although SYBR safe stains NFAT-c1DNA complex and NFAT-c2DNA complex well (SYBR staining), two complexes are weakly stained in the gel stained with Coomassie brilliant blue (CB staining), and free NFAT1 protein staining is very good. By using SYBR or CB staining agents, on the same gel (lower sub-figure) with three replicates side by side loading lanes 2 (2a, 2b and 2c) and lanes 4 (4a, 4b and 4c) samples, the different sizes of NFAT-c1DNA complex (lane 2) and NFAT-c2DNA complex (lane 4) are distinguished.

[0059] Fig.12F The time course of BFPX cross-linking of NFAT1 / DNA complexes is depicted. The binding reaction between NFAT1 DBD and 27mer 5'6-FAM labeled ARRE2 DNA is shown. Fig. 12C The cells were set up as described, with (+) or without (-) BFPX probe SPB-AAD (100 μM) and irradiated with UV for different amounts of time: 3 seconds (3s), 10 seconds (10s), 30 seconds (30s), 100 seconds (100s), 5 min (5m), 15 min (15m), 30 min (30m). Then Fig.12DThe reaction of different time points is analyzed on SDS gel as described. Cross-linked NFAT-DNA complex is visible as early as 3 seconds, and reaches plateau at about 100 seconds and 5min time point. When UV exposure time is longer than 5min, significant amount of cross-linked NFAT-DNA complex can be formed with UV irradiation in the absence of SPB-AAD.

[0060] Figure 12G The time course of BFPX cross-linking of MEF2 / DNA complexes is depicted. We also performed a time course study of BFPX-based cross-linking of another transcription factor, myocyte enhancer factor 2 (MEF2). The human MEF2A DBD (residues 2–95) was cross-linked to a region containing the consensus MEF2 site ( The binding reaction between the 44mer of (SEQ ID NO: 6)) and the 5' 6-FAM labeled DNA is as follows Fig. 12C The cells were set up as described, with (+) or without (-) BFPX probe SPB-AAD (100 μM), and irradiated with UV for different amounts of time: 3 seconds (3s), 10 seconds (10s), 30 seconds (30s), 100 seconds (100s), 5 min (5m), 15 min (15m), 30 min (30m). Then, as Fig.12D The reaction of different time points is analyzed on SDS gel as described. Similarly, cross-linked MEF2-DNA complex is visible as early as 3 seconds, and reaches plateau at about 100 seconds and 5min time points. Different from NFAT / DNA complex, even with a long UV exposure time of up to 30min, cross-linked MEF2 / DNA complex is not observed in the absence of BFPX probe (SPB-AAD). This observation shows that the cross-linking between MEF2 and DNA strictly depends on the use of BFPX probe, and cross-linking exceeds 5min, and a significant amount of cross-linked NFAT-DNA complex can be formed with UV irradiation in the absence of SPB-AAD.

[0061] Fig.12H BFPX cross-linking of Nkx2.5 / DNA complex is depicted. Fig. 12C As described, human Nkx2.5 DBD (homeodomain) was set up with a 19mer 5' 6-FAM labeled DNA containing a consensus sequence Nkx2.5 binding site ( (SEQ ID NO: 7)) and treated with formaldehyde (FA) or BFPX probe and UV irradiation. Various treatment combinations are listed above the gel. The reactions were then analyzed by SDS PAGE under denaturing conditions, as shown in Figure 2. Fig.12DAs described. All three probes SPB-AAD (S, lane 5), SPB-PEG4-AAD (E, lane 7) and SPB-spermidine-AD (M, lane 9) showed significant crosslinking of Nkx2.5 / DNA complexes, while almost no effect was shown on DNA alone (lanes 4, 6 and 8). Again, FA did not cause any protein-DNA crosslinking (lane 2) compared to the control containing only free DNA (lane 1). UV alone (lane 3) did not cause any Nkx2.5 / DNA crosslinking (lane 3) in the absence of the BFPX probe.

[0062] Fig.12I Depicted is the BFPX cross-linking of the p53 / DNA complex. Fig. 12C The human tumor suppressor p53 DBD was set up as described with a 37mer Cy5-labeled DNA (5'- / 5Cy5 / - (SEQ ID NO: 8)) and treated with BFPX probe and UV irradiation. Various treatment combinations are listed above the gel. The reaction was then analyzed by SDS PAGE under denaturing conditions, as shown in Figure 2. Fig.12D As described. All three probes SPB-AAD (S, lane 2), SPB-PEG4-AAD (E, lane 4) and SPB-spermidine-AD (M, lane 6) showed significant cross-linking of p53 / DNA complexes compared to the no protein control (lanes 1, 3 and 5). Since p53 binds to the DNA substrate in the form of a tetramer, the cross-linked p53 / DNA complex may contain one or more p53 protein molecules. Although the monomeric p53 / DNA complex appears to be the main cross-linked protein-DNA complex, higher molecular weight species can be seen in the SDS gel, which may correspond to cross-linked higher order p53 / DNA complexes.

[0063] Fig.12J Reverse BFPX crosslinking using the probe 4AMT-SDAD containing a cleavable linker is depicted. Because UV alone can crosslink NFAT1 to DNA at low to significant levels ( Fig.12F ), but failed to crosslink MEF2 ( Figure 12G We selected MEF2 / DNA complex for this test. Fig. 12C Human MEF2A DBD (residues 2-95, 26 μM for all reactions) was set up as described with a DNA sequence containing the consensus MEF2 site ( The binding reaction between 16mer 5'6-FAM labeled DNA (SEQ ID NO: 6), 10 μM for all reactions, where the amount of 4AMT-SDAD in lanes 3-1 and lanes 6-4 was increased to 10 μM, 50 μM and 100 μM, respectively. The binding reaction was then irradiated with UV for 15 min. The binding reactions in lanes 4, 5 and 6 were further incubated with 200 mM TCEP at 37°C overnight. Then, as Fig.12D All reactions were analyzed on SDS gels as described. The results indicate that cross-linked MEF2 / DNA complexes can be released by reductive cleavage of the disulfide bonds in the linker of 4AMT-SDAD used in the cross-linking reaction.

[0064] Fig.13A Depicted is the characterization of BFPX-mediated cross-linking reactions: Identification of covalent attachment sites on the cross-linked proteins. NFAT1 / DNA complexes were attached to SPB-AAD (corresponding to Fig. 12C Lane 8) and SPB-PEG4-AAD (corresponding to Fig. 12C Lane 11) of the cross-linking reaction was expanded 10 times. A schematic diagram of the DNA binding domain of human NFAT1 used in this study (residues 399-676) is shown at the top; the sequences of loops 434-452 (KAPTGGH (SEQ ID NO: 9) ... MENK (SEQ ID NO: 10)) and loops 478-497 (RITGKTV (SEQ ID NO: 11) ... IVGNTK (SEQ ID NO: 12)) are marked. The reaction mixture was separated on FPLC using a mono-Q column (yellow trace: conductivity, red trace: salt gradient (buffer A: 10mM Hepes pH 7.4; buffer B: 10mM Hepes 7.4, 1M NaCl), blue trace: UV254 nm absorbance). The purified NFAT / DNA complex (PD complex) was digested with trypsin, and the DNA-peptide conjugate was purified and sequenced by Edman degradation. The DNA sequence used in the study is shown ( (SEQ ID NO: 2) and its complementary strand (SEQ ID NO:3)), where two potential probe (yellow wedge and star) binding sites (TpA) are marked. The underlined region is the NFAT binding motif. Edman sequencing of the linked peptides identified IVGN (SEQ ID NO:13) and APTGGH (SEQ ID NO:14).

[0065] Fig. 13BDepicted are Edman sequencing results (IVGN) of the NFAT1 / DNA complex from SPB-AAD crosslinks, identifying the loop at 478-497 adjacent to the DNA (colored in cyan) as the attachment site. This result is consistent with the structural model, in which NFAT binds to its common GGAAAA motif and the SPB-AAD probe (in the space-filling model) inserts into the adjacent TpA site.

[0066] Fig. 13C Describe the Edman sequencing results (APTGGH) (SEQ ID NO: 14) of the NFAT1 / DNA complex cross-linked from SPB-PEG4-AAD, and identify the loop of 434-452 (colored in magenta) as the attachment site. Compared with the loop of 478-497, the loop of 434-452 is farther from DNA. This result is consistent with the structural model, in which NFAT binds to its common GGAAAA motif, and the SPB-PEG4-AAD probe (in the space filling model) inserts into the adjacent TpA site and adopts a conformation that is greatly extended to probe the 434-452 loop. Therefore, the biochemical results are very consistent with the structural model, proving the potential utility of BFPX cross-linking in drawing the protein-DNA interface structure of the transcription factor DNA complex assembled in vitro and separated in situ.

[0067] Fig.14A Characterization and testing of intracellular BFPX cross-linking are depicted. Hela cells were incubated with SPB-AAD or SPB-PEG4-AAD (10 μM or 100 μM) in the dark for 30 min to allow DNA binding and intercalation through the psoralen portion of the BFPX probe. After washing away excess BFPX probe, 365 nm UV irradiation was performed for 5 min. Alexa Fluor 647 picolyl azide molecules (from Click-iT TM Plus Alexa Fluor TM Alexa Fluor 647 Picolyl Azide Toolkit) reacts with the alkyne groups on SPB-AAD or SPB-PEG4-AAD via a click reaction. After click chemistry labeling, unreacted fluorescent molecules are removed by washing. Cell nuclei are analyzed using fluorescence imaging. Little or no fluorescent signal remains in the cells when no BFPX probe is added (columns 1 and 2) or when the probe is added without UV treatment (columns 3 and 6). In contrast, Alexa Fluor 647 picolyl azide molecules can be immobilized in cells via a click reaction with SPB-AAD (columns 4 and 5) or SPB-PEG4-AAD (columns 7 and 8) only when the BFPX probe is present and activated by UV cross-linking.

[0068] Fig. 14B Quantitative analysis is depicted, showing that the fixed fluorescence signal is dose-dependent on the concentrations of SPB-AAD and SPB-PEG4-AAD under UV irradiation.

[0069] Fig. 14C Depicts transfection of 10 cm plates of HEK293T cells (~8.8 million) with 2 μg AVI-TEV-FLAG-FOXP3 for approximately 48 hr. Cells were harvested into 1.6 mL of room temperature 1x DPBS buffer (calcium and magnesium free). Cells were aliquoted into 300 μL per sample, and SPB-AAD 10 μM (A10) or 100 μM (A100) or SPB-PEG4-AAD 100 μM (p100) was added to each sample. The samples were incubated at 37°C in the dark with rotation for 30 min. The samples were transferred to the middle well of a 24-well plate (Corning #3527). The height of 300 μL of liquid in a 24-well plate well was approximately 1.5 mm. The samples in the 24-well plate were then UV irradiated (LED chip: 21.4mm x 21.4mm, 365-37nm, 30 watts, sample distance (bottom of well plate to lamp plane) was 4cm, and exposure time was 2-5min at 4°C). The BFPX-treated cells were recovered from the 24-well plate and washed with 1x DPBS, centrifuged at 500g for 5min to pellet the cells and remove the DPBS buffer. 20μL 1x RIPA buffer was added to 500,000 BFPX-treated cells and shaken on ice for 15min to lyse the cells. 3μL 10xMNase, 0.25μL RnaseA, 12.5 units of Mnase were then added to the lysis mixture and ddw to 30μL, incubated at 37°C for 15min. 6μL SDS loading dye (2% SDS) was added to the lysis mixture and boiled at 95°C for 10min. Run on 4%-15% SDS-PAGE gel at 200 V for 40 min and analyze by Western blotting (transfer overnight at 18 V (30 mA), block in 2% BSA in PBS for 45 min, primary antibody: anti-FLAG 1:2000; room temperature, 2 hr; secondary antibody: anti-mouse light chain kappa 1:100,000; Femto Signal ECL). Fig. 14CAs shown, higher molecular weight substances containing FOXP3 were observed in cells treated with SPB-AAD (A) and SPB-PEG4-AAD (P) (lanes 4-9) compared with the control (lanes 1-3), and the formation of cross-linked FOXP3 complexes was dose-dependent on probe concentration (compare lane 4 (SPB-AAD, 10 μM, A10) and lane 6 (SPB-AAD, 100 μM, A100)) and UV exposure time (compare lane 4 (2 min) and lane 5 (5 min) when the SPB-AAD concentration was 10 μM (A10); also compare lane 8 (2 min) and lane 9 (5 min) when the SPB-PEG4AAD concentration was 100 μM (P100)).

[0070] Fig.14D Describes the Fig. 14C ) were treated with the same BFPX-treated HEK293T cells transfected with AVI-TEV-FLAG-FOXP3 as described in the previous study, and genomic DNA was extracted using TRIzol. Briefly, for each aliquot of 250,000 cells, 200 μL TRIzol TMThe samples were lysed and homogenized in 4% paraformaldehyde (PNA) reagent. Incubate for 5 minutes to completely dissociate the nucleoprotein complex. Then 40 μL of chloroform was added to the sample and incubated for 3 minutes. The sample was then centrifuged at 12,000 × g for 15 minutes at 4°C to separate into a lower red phenol / chloroform-intermediate layer and a colorless upper aqueous phase. The upper aqueous phase was removed by pipetting. 60 μL of 100% ethanol was added to the remaining organic phase and the interface layer, mixed, incubated for 3 minutes, and then centrifuged at 2,000 × g for 5 minutes at 4°C to pellet the genomic DNA. The DNA pellet was resuspended in 200 μL of 0.1 M sodium citrate / 10% ethanol (pH 8.5), incubated for 30 minutes, and then centrifuged at 2,000 × g for 5 minutes at 4°C to pellet the genomic DNA again. The 0.1 M sodium citrate / 10% ethanol wash was repeated once more. The DNA was resuspended in 400 μL of 75% ethanol, incubated for 20 min, and centrifuged at 2,000×g for 5 min at 4°C to pellet the genomic DNA. The DNA pellet was air-dried for 10 min. The DNA pellet was dissolved in 300 μL of freshly prepared 8 mM NaOH. Once the DNA was dissolved, the pH was adjusted to ~ pH 7.5 using 3M sodium acetate (pH 5.2, add about 10-11 μL). 10xBenzonase buffer (final concentration 1x) and 0.2 units of Benzonase were then added to the genomic DNA sample to completely digest the genomic DNA (checked by agarose gel). SDS loading dye (2% SDS) was then added to the sample and boiled at 95°C for 10 min. Run on 4%-15% SDS-PAGE gel at 200 V for 40 min and analyze by Western blotting (transfer overnight at 20 V (40 mA), block in 2% BSA in PBS for 45 min, primary antibody: anti-FLAG 1:2000; room temperature, 2 hr; secondary antibody: anti-mouse light chain kappa 1:100,000; SuperSignal TM Western Blot Substrate Bundle,Femto). like Fig.14D As shown, only cells treated with UV and probe (SPB-AAD, 100 μM, A100) showed FOXP3 protein (lanes 4 and 5) compared to the control (lanes 1-3). This observation suggests that in cells treated with BFPX, FOXP3 protein is covalently linked to genomic DNA and can be extracted by TRIzol protocol under strong denaturing conditions, while in control cells, the non-covalent nucleoprotein complex is completely dissociated. In addition, Benzonase digestion not only releases FOXP3 with a molecular weight similar to its free monomer, but also releases higher molecular weight species, suggesting that FOXP3 may bind to certain genomic regions in the form of higher-order structures that are resistant to Benzonase digestion.

[0071] Fig.14E Figure 3. Capture of endogenous transcription factor MEF2 bound to genomic DNA in GM12878 cells. GM12878 cells (7.5 million per sample) were treated with different combinations of UV irradiation and BFPX probes, and genomic DNA was extracted using the TRIzol protocol, digested with Benzonase, and analyzed by SDS-PAGE / Western blotting, as shown in Figure 3. Fig.14D As shown. Here, the Trizol extraction procedure was scaled up to 7.5 million cells per mL of TRIzol reagent. The final DNA pellet was dissolved in 300 μL 8 mM NaOH, neutralized with 10-11 μL 3 M sodium acetate (pH 5.2), and digested with 25U Benzonase for 1 hour. The results showed that MEF2C covalently bound to genomic DNA extracted with TRIzol under strong denaturing conditions was only observed in cells treated with UV and probe (lanes 3, 4, and 5) compared to the control group (lane 1). The UV-only control (lane 2) showed a weak signal, indicating that UV alone may cross-link some proteins to DNA at a low level. However, the cross-linking efficiency in cells treated with BFPX was significantly higher than that of the UV-only control (compare lanes 3-lane 5 and lane 2). DETAILED DESCRIPTION

[0072] All references cited herein are incorporated by reference in their entirety as if fully set forth. Unless otherwise defined, technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which the invention belongs. Singleton et al., Dictionary of Microbiology and Molecular Biology 3 rd ed., Revised, J. Wiley & Sons (New York, NY 2006); March, Advanced Organic Chemistry Reactions, Mechanisms and Structure 7 th ed., J. Wiley & Sons (New York, NY 2013); and Sambrook and Russel, Molecular Cloning: A Laboratory Manual 4 thed., Cold Spring Harbor Laboratory Press (Cold Spring Harbor, NY 2012) provides those skilled in the art with a general guide to many of the terms used in this application. For reference on how to make antibodies, see D. Lane, Antibodies: A Laboratory Manual 2 nd ed. (Cold Spring Harbor Press, Cold Spring Harbor NY, 2013); Kohler and Milstein, (1976) Eur. J. Immunol. 6:511; Queen et al., U.S. Pat. No. 5,585,089; and Riechmann et al., Nature 332:323 (1988); U.S. Pat. No. 4,946,778; Bird, Science 242:423-42 (1988); Huston et al., Proc. Natl. Acad. Sci. USA 85:5879-5883 (1988); Ward et al., Nature 334:544-54 (1989); Tomlinson I. and Holliger P. (2000) Methods Enzymol, 326, 461-479; Holliger P. (2005) Nat. Biotechnol. Sep; 23(9):1126-36).

[0073] Those skilled in the art will recognize that many methods and materials similar or equivalent to those described herein can be used in the practice of the present invention. Other features and advantages of the present invention will become apparent through the following detailed description combined with the accompanying drawings (which illustrate the various features of embodiments of the present invention by way of example). In fact, the present invention is by no means limited to the methods and materials described. For purposes of the present invention, the following terms are defined as follows. For convenience, some of the terms used herein in the specification, embodiments, and appended claims are collected here.

[0074] Unless otherwise stated, or implied from the context, the following terms and phrases include the meanings provided below. Unless otherwise expressly stated, or apparent from the context, the following terms and phrases do not exclude the meanings that the term or phrase has acquired in the field to which it belongs. Unless otherwise defined, all technical terms and scientific terms used herein have the same meanings as those generally understood by those of ordinary skill in the art to which the invention belongs. It should be understood that the present invention is not limited to the specific methods, protocols, reagents, etc. described herein, and can therefore vary. The definitions and terms used herein are provided to help describe specific embodiments and are not intended to limit the claimed invention, as the scope of the invention is limited only by the claims.

[0075] As used herein, the term "comprising or comprises" is used to refer to compositions, methods, systems, articles, devices, and their respective components that are useful for embodiments, but is still open to the inclusion of unspecified elements (whether useful or not). Those skilled in the art will understand that, in general, the terms used herein are generally intended to be "open" terms (e.g., the term "including" should be interpreted as "including but not limited to", the term "having" should be interpreted as "at least having", the term "includes" should be interpreted as "including but not limited to", etc.). Although the open term "comprising" is used herein to describe and claim the present invention as a synonym for terms such as including, containing, or having, the present invention or its embodiments may also be described alternatively using alternative terms (e.g., "consisting of" or "consisting essentially of".

[0076] Unless otherwise stated, the terms "a / an" and "the" and similar references used in the context of describing the specific embodiments of the present application (especially in the context of the claims) may be interpreted as covering both the singular and the plural. The enumeration of the value range herein is intended only to be used as a shorthand method for individually referring to each individual value falling within the range. Unless otherwise indicated herein, each individual value is incorporated into the specification as if it were individually listed herein. Unless otherwise indicated herein or clearly contradictory to the context, all methods described herein can be performed in any suitable order. The use of any and all examples or exemplary language (e.g., "such as") provided with respect to certain embodiments herein is intended only to better illustrate the present application and is not intended to limit the scope of the present application otherwise claimed. The abbreviation "eg" is derived from the Latin exempli gratia and is used herein to indicate non-limiting examples. Therefore, the abbreviation "eg" is synonymous with the term "for example". Any language in the specification should not be interpreted as indicating any unclaimed elements necessary for the practice of the present application.

[0077] "Optional" or "optionally" means that the subsequently described event may or may not occur, so that the description includes instances where the event occurs and instances where it does not.

[0078] In some embodiments, the numerical value of the amount of the representation reagent, characteristic (such as concentration), reaction conditions, etc. for describing and claiming some embodiments of the present invention will be understood to be modified by the term "about" in some cases.Therefore, in some embodiments, the numerical parameter set forth in the written description and the appended claims is an approximate value, which can be changed according to the desired characteristic sought to be obtained by a particular embodiment.In some embodiments, the numerical parameter should be explained according to the numerical value of the reported significant digit and by applying common rounding techniques.Although the numerical range and the parameter of the wide range of some embodiments of the present invention are approximate values, the numerical value set forth in the specific examples is reported as accurately as possible.The numerical value presented in some embodiments of the present invention may include certain errors that are inevitably caused by the standard deviation found in their respective test measurements.

[0079] The grouping of the alternative elements or embodiments of the invention disclosed herein should not be construed as having limitations. Each group member may be mentioned and claimed individually, or may be mentioned and claimed in any combination with other members of the group or other elements found herein. For convenience and / or patentability reasons, one or more members of the group may be included in the group or deleted from the group. When any such inclusion or deletion occurs, this specification is deemed to include the modified group at this time, so as to meet the written description of all Markush groups used in the appended claims.

[0080] As used herein, the term "electron donating group" is well known in the art and generally refers to a functional group or atom that pushes electron density from itself to other parts of a molecule (e.g., through resonance and / or inductive effects). Non-limiting examples of electron donating groups include OR c NR c R d , an alkyl group, wherein R c and R d Each is independently H, optionally substituted alkyl, optionally substituted heteroalkyl, optionally substituted aryl, optionally substituted heteroaryl, optionally substituted cyclyl or optionally substituted heterocyclyl.

[0081] As used herein, the term "electron withdrawing group" is well known in the art and generally refers to a functional group or atom that pulls electron density toward itself and away from other parts of a molecule (e.g., through resonance and / or inductive effects). Non-limiting examples of electron withdrawing groups include NO2, F, Cl, Br, I, CF3, CN, CO2R a 、C(=O)NR a R b 、C(=O)R a 、SO2R a 、SO2OR a 、SO2NR a R b PO3R a R b , or NO, where R a and R b Each is independently H, optionally substituted alkyl, optionally substituted heteroalkyl, optionally substituted aryl, optionally substituted heteroaryl, optionally substituted cyclyl or optionally substituted heterocyclyl.

[0082] As used herein, the term "alkyl" means a straight or branched saturated aliphatic radical having a carbon atom chain. x Alkyl and C x -C y Alkyl, where X and Y represent the number of carbon atoms in the chain. For example, C1-C6 alkyl includes alkyls having chains of 1 to 6 carbons (e.g., methyl, ethyl, propyl, isopropyl, butyl, sec-butyl, isobutyl, tert-butyl, pentyl, neopentyl, hexyl, etc.). Alkyl represented together with another radical (e.g., as in arylalkyl) means a straight or branched saturated alkyl divalent radical having the specified number of atoms, or a bond when no atoms are specified. For example (C6-C 10 )Aryl (C0-C3) alkyl includes phenyl, benzyl, phenethyl, 1-phenethyl, 3-phenylpropyl, etc. The main chain of the alkyl group may be optionally inserted with one or more heteroatoms (such as N, O or S).

[0083] In a preferred embodiment, a straight chain or branched alkyl group has 30 or fewer carbon atoms in its backbone (e.g., C1-C2 for a straight chain). 30 , for the branched chain C3-C 30 ), and more preferably 20 or less. Likewise, preferred cycloalkyl groups have 3-10 carbon atoms in their ring structure, and more preferably 5, 6 or 7 carbons in the ring structure. The term "alkyl" (or "lower alkyl") used throughout the application documents, examples and claims is intended to include "unsubstituted alkyl" and "substituted alkyl", wherein the latter refers to an alkyl moiety having one or more substituents (replacing hydrogen on one or more carbons of the hydrocarbon backbone).

[0084] Unless the number of carbon atoms is otherwise specified, "lower alkyl" as used herein means an alkyl group as defined above but having 1 to 10 carbons, more preferably 1 to 6 carbon atoms in its backbone structure. Likewise, "lower alkenyl" and "lower alkynyl" have similar chain lengths. Throughout the application, preferred alkyl groups are lower alkyls. In preferred embodiments, substituents identified herein as alkyls are lower alkyls.

[0085] Non-limiting examples of substituents for the substituted alkyl groups may include halogen, hydroxy, nitro, thiol, amino, azido, imino, amido, phosphoryl (including phosphonates and phosphites), sulfonyl (including sulfates, sulfonamido, sulfamoyl and sulfonates), and silyl groups, as well as ethers, alkylthiols, carbonyl (including ketones, aldehydes, carboxylates and esters), -CF3, -CN, and the like.

[0086] As used herein, the term "alkenyl" refers to a linear, branched or cyclic unsaturated hydrocarbon radical having at least one carbon-carbon double bond. x Alkenyl and C x -C y Alkenyl, wherein X and Y represent the number of carbon atoms in the chain. For example, C2-C6 alkenyl includes alkenyl (e.g., vinyl, allyl, propenyl, isopropenyl, 1-butenyl, 2-butenyl, 3-butenyl, 2-methylallyl, 1-hexenyl, 2-hexenyl, 3-hexenyl, etc.) having a chain of 2-6 carbons and at least one double bond. The alkenyl represented together with another atomic group (e.g., as in aryl alkenyl) refers to a straight or branched alkenyl divalent atomic group with a specified number of atoms. The main chain of the alkenyl may optionally be inserted with one or more heteroatoms (e.g., N, O, or S).

[0087] As used herein, the term "alkynyl" refers to an unsaturated hydrocarbon radical having at least one carbon-carbon triple bond. x Alkynyl and C x -Cy Alkynyl, wherein X and Y represent the number of carbon atoms in the chain. For example, C2-C6 alkynyl includes alkynyls having chains of 2-6 carbons and at least one triple bond (e.g., ethynyl, 1-propynyl, 2-propynyl, 1-butynyl, isopentynyl, 1,3-hexadiynyl, n-hexynyl, 3-pentynyl, 1-hexene-3-ynyl, etc.). Alkynyl represented together with another atomic group (e.g., as in arylalkynyl) means a straight or branched alkynyl divalent atomic group with a specified number of atoms. The main chain of the alkynyl may optionally be inserted with one or more heteroatoms (e.g., N, O or S).

[0088] The terms "alkylene", "alkenylene" and "alkynylene" refer to divalent alkyl, alkenylene and alkynylene radicals. The prefix C is often used. x and C x -C y , where X and Y represent the number of carbon atoms in the chain. For example, C1-C6 alkylene includes methylene (-CH2-), ethylene (-CH2CH2-), trimethylene (-CH2CH2CH2-), tetramethylene (-CH2CH2CH2CH2-), 2-methyltetramethylene (-CH2CH(CH3)CH2CH2-), pentamethylene (-CH2CH2CH2CH2CH2-), and the like.

[0089] As used herein, the term "alkylidene" refers to a group having the general formula =CR a R b A straight-chain or branched unsaturated aliphatic divalent atomic group. a and R b Non-limiting examples of are each independently hydrogen, alkyl, substituted alkyl, alkenyl or substituted alkenyl. C is often used x Alkane subunit and C x -C y Alkylene, wherein X and Y represent the number of carbon atoms in the chain. For example, C2-C6 alkynylene includes methylene (=CH2), ethylene (=CHCH3), isopropylene (=C(CH3)2), propyleneene (=CHCH2CH3), allylene (=CH-CH=CH2), etc.

[0090] As used herein, the term "heteroalkyl" refers to a linear or branched or cyclic carbon-containing atom group or a combination thereof comprising at least one heteroatom. Suitable heteroatoms include, but are not limited to, O, N, Si, P, Se, B, and S, wherein phosphorus and sulfur atoms are optionally oxidized, and nitrogen heteroatoms are optionally quaternized. Heteroalkyl can be substituted as defined above for alkyl groups.

[0091] As used herein, the term "halogen" or "halo" refers to an atom selected from fluorine (F), chlorine (Cl), bromine (Br) and iodine (I). The term "halogen radioisotope" or "halogenated radioisotope" refers to a radionuclide of an atom selected from fluorine (F), chlorine (Cl), bromine (Br) and iodine (I).

[0092] In some embodiments, "iodine" when used in the context of a halo functional group or a halogen functional group or as a halo substituent or a halogen substituent refers to an iodine atom (I).

[0093] In some embodiments, "bromine" when used in the context of a halo or halogen functional group or as a halo or halogen substituent refers to a bromine atom (Br).

[0094] In some embodiments, "chlorine" when used in the context of a halo or halogen functional group or as a halo or halogen substituent refers to a chlorine atom (Cl).

[0095] In some embodiments, "fluorine" when used in the context of a halo functional group or a halogen functional group or as a halo substituent or a halogen substituent refers to a fluorine atom (F).

[0096] As a separate group or part of a larger group, "halogen substituted moiety" or "halogenated moiety" means an aliphatic, alicyclic or aromatic moiety described herein substituted with one or more "halogen" atoms, as such terms are defined in this application. For example, halogen-substituted alkyl groups include haloalkyl, dihaloalkyl, trihaloalkyl, perhaloalkyl, etc. (e.g., halogen-substituted (C1-C3) alkyl groups include chloromethyl, dichloromethyl, difluoromethyl, trifluoromethyl (-CF3), 2,2,2-trifluoroethyl, perfluoroethyl, 2,2,2-trifluoro-1,1-dichloroethyl, etc.).

[0097] The term "aryl" refers to a monocyclic, bicyclic or tricyclic fused aromatic ring system. C x Aryl and C x -C y Aryl, where X and Y represent the number of carbon atoms in the ring system. For example, C6-C 12Aryl includes aryl groups having 6 to 12 carbon atoms in the ring system. Exemplary aryl groups include, but are not limited to, pyridyl, pyrimidinyl, furanyl, thienyl, imidazolyl, thiazolyl, pyrazolyl, pyridazinyl, pyrazinyl, triazinyl, tetrazolyl, indolyl, benzyl, phenyl, naphthyl, anthracenyl, azulenyl, fluorenyl, indanyl, indenyl, naphthyl, phenyl, tetrahydronaphthyl, benzimidazolyl, benzofuranyl, benzothiofuranyl, benzothiophenyl, benzoxazolyl, benzoxazolinyl, benzothiazolyl, benzotriazolyl, benzotetrazolyl, benzisoxazolyl, benzisothiazolyl, benzimidazolinyl, carbazolyl, 4aH-carbazolyl, carboline yl, chromanyl, chromenyl, cinnolinyl, decahydroquinolinyl, 2H,6H-1,5,2-dithiazinyl, dihydrofuran[2,3b]tetrahydrofuran, furanyl, furazanyl, imidazolidinyl, imidazolinyl, imidazolyl, 1H-indazolyl, indolenyl, indolinyl, indolizinyl, indolyl, 3H-indolyl, isatinoyl, isobenzofuranyl, isochromanyl, isoindazolyl, iso indolyl, isoindolyl, isoquinolyl, isothiazolyl, isoxazolyl, methylenedioxyphenyl, morpholinyl, naphthyridinyl, octahydroisoquinolinyl, oxadiazolyl, 1,2,3-oxadiazolyl, 1,2,4-oxadiazolyl, 1,2,5-oxadiazolyl, 1,3,4-oxadiazolyl, oxazolidinyl, oxazolyl, oxindolyl, pyrimidinyl, phenanthridinyl, phenanthrolinyl, phenazinyl, phenothiazinyl, phenoxathinyl, phenoxazinyl, phthalazinyl, piperazinyl, piperidinyl, piperidonyl, 4-piperidonyl, piperonyl, pteridinyl, purinyl, pyranyl, pyrazinyl, pyrazolidinyl, pyrazolinyl, pyrazolyl, pyridinyl oxazine, pyridoxazole, pyridoimidazole, pyridothiazole, pyridinyl / pyridyl, pyrimidinyl, pyrrolidinyl, pyrrolinyl, 2H-pyrrolyl, pyrrolyl, quinazolinyl, quinolinyl, 4H-quinolizinyl, quinoxalinyl, quinuclidinyl, tetrahydrofuranyl, tetrahydroisoquinolinyl, tetrahydroquinolinyl, tetrazolyl, 6H-1,2,5-thiadiazinyl, 1,2,3-thiadiazolyl, 1,2,4-thiadiazolyl, 1,2,5-thiadiazolyl, 1,3,4-thiadiazolyl, thianthrenyl, thiazolyl, thienyl, thienothiazolyl, thienooxazolyl, thienoimidazolyl, thienyl and xanthenyl, etc. In some embodiments, 1, 2, 3 or 4 hydrogen atoms of each ring can be substituted with a substituent.

[0098] The term "heteroaryl" refers to an aromatic 5-8 membered monocyclic, 8-12 membered fused bicyclic, or 11-14 membered fused tricyclic ring system having heteroatoms: 1-3 heteroatoms if monocyclic; 1-6 heteroatoms if bicyclic; or 1-9 heteroatoms if tricyclic; the heteroatoms being selected from O, N, or S (e.g., carbon atoms and 1-3 N, O, or S heteroatoms, 1-6 N, O, or S heteroatoms, or 1-9 N, O, or S heteroatoms, respectively, if monocyclic, bicyclic, or tricyclic). C is generally used. x Heteroaryl and C x -C yHeteroaryl, wherein X and Y represent the number of carbon atoms in the ring system. For example, C4-C9 heteroaryl includes heteroaryl having 4 to 9 carbon atoms in the ring system. Heteroaryl includes, but is not limited to, heteroaryl derived from benzo[b]furan, benzo[b]thiophene, benzimidazole, imidazo[4,5-c]pyridine, quinazoline, thieno[2,3-c]pyridine, thieno[3,2-b]pyridine, thieno[2,3-b]pyridine, indolizine, imidazo[1,2a]pyridine, quinoline, isoquinoline, phthalazine, quinoxaline, naphthyridine, quinolizine, indole, isoindole, indazole, indoline, benzoxazole, benzopyrazole, benzothiazole, imidazo[1,5-a]pyridine, pyrazolo[1,5-a]pyridine, imidazo[1,2-a]pyrim ... [1,2-c]pyrimidine, imidazo[1,5-a]pyrimidine, imidazo[1,5-c]pyrimidine, pyrrolo[2,3-b]pyridine, pyrrolo[2,3c]pyridine, pyrrolo[3,2-c]pyridine, pyrrolo[3,2-b]pyridine, pyrrolo[2,3-d]pyrimidine, pyrrolo[3,2-d]pyrimidine, pyrrolo[2,3-b]pyrazine, pyrazolo[1,5-a]pyridine, pyrrolo[1,2-b]pyridazine, pyrrolo[1,2-c]pyrimidine, pyrrolo[1,2-a]pyrimidine, pyrrolo[1,2-a]pyrazine, triazolo[1,5-a]pyridine pyridine, pteridine, purine, carbazole, acridine, phenazine, phenothiazene, phenoxazine, 1,2-dihydropyrrolo[3,2,1-hi]indole, indolizine, pyrido[1,2-a]indole, 2-(1H)-pyridone, benzimidazolyl, benzofuranyl, benzothiofuranyl, benzothiophenyl, benzoxazolyl, benzoxazolinyl, benzothiazolyl, benzotriazolyl, benzotetrazolyl, benzisoxazolyl, benzisothiazolyl, benzimidazolinyl, carbazolyl, 4aH-carbazolyl, carbolyl, chromanyl, chromenyl, cinnolinyl, decahydroquinolinyl, 2H,6H-1,5,2-dithiazinyl, dihydrofurano[2,3b]tetrahydrofuran, furanyl, furazolidinyl, imidazolinyl, imidazolyl, 1H-indazolyl, indolene, indolyl, indolizinyl, indolyl, 3H-indolyl, isatoyl, isobenzofuranyl, isochromanyl, isoindazolyl, isoindolyl, isoindolyl, isoquinolyl, isothiazolyl, isoxazolyl, methylenedioxyphenyl, morpholinyl, naphthyridinyl, octahydroisoquinolyl, oxadiazolyl, 1,2,3-oxadiazolyl, 1,2,4-oxadiazolyl, 1,2,5-oxadiazolyl, 1,3,4-oxadiazolyl, oxazolidinyl, oxazolyl, oxepanyl, oxetanyl, oxindolyl, pyrimidinyl, phenanthridinyl, phenanthrolinyl, phenazinyl, phenothiazinyl, phenoxathiyl, phenoxazinyl, phthalazinyl, piperazinyl, piperidinyl, piperidonyl, 4-piperidonyl, piperonyl, pteridinyl, purinyl, pyranyl, pyrazinyl, pyrazolidinyl, pyrazolinyl, pyrazolyl, pyridazinyl, pyridooxazole, pyridoimidazole, pyridothiazole, pyridinyl (pyridinyl / pyridinyl) yridyl), pyrimidinyl, pyrrolidinyl, pyrrolinyl, 2H-pyrrolyl, pyrrolyl, quinazolinyl, quinolyl, 4H-quinolizinyl, quinoxalinyl, quinuclidine, tetrahydrofuranyl, tetrahydroisoquinolyl, tetrahydroquinolyl, tetrazolyl, 6H-1,2,5-thiadiazinyl, 1,2,3-thiadiazolyl, 1,2,4-thiadiazolyl, 1,2,5-thiadiazolyl, 1,3,4-thiadiazolyl, thianthrenyl, thiazolyl, thienyl, thienothiazolyl, thienoxazolyl, thienoimidazolyl, thienyl and xanthenyl, and the like. Some exemplary heteroaryl groups include, but are not limited to, pyridyl, furyl, imidazolyl, benzimidazolyl, pyrimidinyl, thiophenyl, thienyl, pyridazinyl, pyrazinyl, quinolyl, indolyl, thiazolyl, naphthyridinyl, 2-amino-4-oxo-3,4-dihydropteridin-6-yl, tetrahydroisoquinolyl, etc. In some embodiments, 1, 2, 3, or 4 hydrogen atoms of each ring may be replaced by a substituent.

[0099] The term "cyclyl" or "cycloalkyl" refers to saturated and partially unsaturated cyclic hydrocarbon groups having 3 to 12 carbons (e.g., 3 to 8 carbons, and e.g., 3 to 6 carbons). C x Ring group and C x -C y Cyclic groups, wherein X and Y represent the number of carbon atoms in the ring system. For example, C3-C8 cyclic groups include cyclic groups having 3 to 8 carbon atoms in the ring system. Cyclic alkyl groups may additionally be optionally substituted, for example, by 1, 2, 3 or 4 substituents. 10 Cyclic groups include cyclopropyl, cyclobutyl, cyclopentyl, cyclohexyl, cyclohexenyl, 2,5-cyclohexadienyl, cycloheptyl, cyclooctyl, bicyclo[2.2.2]octyl, adamantane-1-yl, decahydronaphthyl, oxocyclohexyl, dioxocyclohexyl, thiocyclohexyl, 2-oxobicyclo[2.2.1]hept-1-yl, and the like.

[0100] Aryl and heteroaryl groups may be optionally substituted at one or more positions with one or more substituents, such as halogen, alkyl, aralkyl, alkenyl, alkynyl, cycloalkyl, hydroxy, amino, nitro, sulfhydryl, imino, amido, phosphate, phosphonate, phosphite, carbonyl, carboxyl, silyl, ether, alkylthio, sulfonyl, ketone, aldehyde, ester, heterocyclic, aromatic or heteroaromatic moiety, -CF3, -CN, and the like.

[0101] The term "heterocyclyl" refers to a non-aromatic 4-8 membered monocyclic, 8-12 membered bicyclic or 11-14 membered tricyclic ring system having heteroatoms: if monocyclic, 1-3 heteroatoms; if bicyclic, 1-6 heteroatoms; or if tricyclic, 1-9 heteroatoms; the heteroatoms being selected from O, N or S (e.g., carbon atoms and 1-3 heteroatoms of N, O or S, 1-6 heteroatoms of N, O or S or 1-9 heteroatoms of N, O or S, respectively, if monocyclic, bicyclic or tricyclic). C is generally used. x Heterocyclic and C x -C y Heterocyclyl, wherein X and Y represent the number of carbon atoms in the ring system. For example, C4-C9 heterocyclyl includes heterocyclyls having 4-9 carbon atoms in the ring system. In some embodiments, 1, 2 or 3 hydrogen atoms of each ring may be substituted by a substituent. Exemplary heterocyclyl groups include, but are not limited to, piperazinyl, pyrrolidinyl, dioxanyl, morpholinyl, tetrahydrofuranyl, piperidinyl, 4-morpholinyl, 4-piperazinyl, pyrrolidinyl, perhydropyrrolazinyl, 1,4-diazaperhydroepinyl, 1,3-dioxanyl, 1,4-dioxanyl, etc.

[0102] The terms "bicyclic" and "tricyclic" refer to multiple ring assemblies fused, bridged or linked by single bonds.

[0103] The term "cyclylalkylene" means a divalent aryl, heteroaryl, cyclyl or heterocyclyl group.

[0104] As used herein, the term "fused ring" refers to a ring that combines with another ring to form a compound having a bicyclic structure when the ring atoms common to the two rings are directly bonded to each other. Non-exclusive examples of common fused rings include: decalin, naphthalene, anthracene, phenanthrene, indole, furan, benzofuran, quinoline, etc. Compounds having fused ring systems can be saturated, partially saturated, cyclic, heterocyclic, aromatic, heteroaromatic, etc.

[0105] As used herein, the term "carbonyl" refers to the radical -C(O)-. It should be noted that the carbonyl radical can be further substituted with various substituents to form different carbonyl groups, including acids, acid halides, amides, esters, ketones, and the like.

[0106] The term "carboxy" means the radical -C(O)O-. It should be noted that the compounds described herein containing a carboxyl moiety may include protected derivatives thereof, i.e., wherein the oxygen is replaced by a protecting group. Suitable protecting groups for the carboxyl moiety include benzyl, tert-butyl, and the like. The term "carboxylic acid" means -COOH.

[0107] The term "cyano" refers to the radical -CN.

[0108] The term "heteroatom" refers to an atom that is not a carbon atom. Specific examples of heteroatoms include, but are not limited to, nitrogen, oxygen, sulfur, and halogens. A "heteroatom moiety" includes a moiety in which the atom to which the moiety is attached is not carbon. Examples of heteroatom moieties include -N=, -NR N -、-N + (O - )=, -O-, -S- or -S(O)2-, -OS(O)2- and -SS-, wherein R N is H or another substituent.

[0109] The term "hydroxy" refers to the radical -OH.

[0110] The term "imine derivative" means a derivative comprising the moiety -C(NR)-, wherein R comprises a hydrogen or carbon atom alpha to the nitrogen.

[0111] The term "nitro" refers to the radical -NO2.

[0112] "Oxaaliphatic," "Oxaalicyclic," or "Oxaaromatic" means an aliphatic, alicyclic, or aromatic as defined herein, except that one or more oxygen atoms (-O-) are positioned between carbon atoms of the aliphatic, alicyclic, or aromatic, respectively.

[0113] "Oxoaliphatic", "Oxoalicyclic" or "Oxoaromatic" means an aliphatic, alicyclic or aromatic as defined herein substituted with a carbonyl group. The carbonyl group may be an aldehyde, ketone, ester, amide, acid or acyl halide.

[0114] As used herein, the term "aromatic" means a moiety wherein the constituent atoms form an unsaturated ring system, all atoms in the ring system being sp 2 The total number of hybridized and pi electrons is equal to 4n + 2. An aromatic ring may be one in which the ring atoms are only carbon atoms (eg, aryl) or may contain carbon atoms and non-carbon atoms (eg, heteroaryl).

[0115] As used herein, the term "substituted" refers to the independent replacement of one or more (usually 1, 2, 3, 4 or 5) hydrogen atoms on a substituted moiety with a substituent, the substituent being independently selected from the group of substituents listed in the definition of "substituent" below or otherwise specified. In general, a non-hydrogen substituent can be any substituent that can be bonded to an atom of a given moiety that is specified to be substituted. Examples of substituents include, but are not limited to, acyl, acylamino, acyloxy, aldehyde, alicyclic, aliphatic, alkanesulfonylamino, alkanesulfonyl, alkaryl, alkenyl, alkoxy, alkoxycarbonyl, alkyl, alkylamino, alkylcarbamoyl, alkylene, alkylene, alkylthio, alkynyl, amide, amine, amino, amino, aminoalkyl, aralkyl, aralkylsulfonamido, aralkylsulfonyl, aromatic, aryl, arylamino, arylcarbamoyl, aryloxy, azido, carbamoyl, , carbonyl, carbonyl compounds (carbonyls) (including ketone, carboxyl, carboxylate / salt), CF3, cyano (CN), cycloalkyl, cycloalkylidene, ester, ether, haloalkyl, halogen, halogen, heteroaryl, heterocyclic radical, hydroxyl, hydroxyalkyl, imino, iminoketone, ketone, sulfhydryl, nitro, oxaalkyl, oxo, oxoalkyl, phosphoryl (including phosphonate / salt and phosphite / salt), silyl, sulfonamido, sulfonyl (including sulfate / salt, sulfamoyl and sulfonate / salt), thiol and urea moieties; each of which may also be optionally substituted or unsubstituted. In some cases, two substituents together with the carbon to which they are attached may form a ring.

[0116] Substituents may be protected as needed, and any protecting group commonly used in the art may be used. Non-limiting examples of protecting groups may be found in, for example, Greene et al., Protective Groups in Organic Synthesis, 3rd Ed. (New York: Wiley, 1999).

[0117] The term "alkoxyl" as used herein refers to an alkyl group as defined above having an oxygen radical connected thereto. Representative alkoxy groups include methoxy, ethoxy, propoxy, tert-butoxy, n-propoxy, isopropoxy, n-butoxy, isobutoxy, etc. "Ether" is two hydrocarbons covalently linked by oxygen. Therefore, the substituent of the alkyl group that makes this alkyl an ether is an alkoxy or is similar to an alkoxy, such as the alkoxy group can be represented by one of -O-alkyl, -O-alkenyl and -O-alkynyl. Aryloxy can be represented by -O-aryl or O-heteroaryl, wherein aryl and heteroaryl are defined as follows. The alkyl of alkoxy and aryloxy groups can be substituted as described above.

[0118] As used herein, the term "aralkyl" refers to an alkyl group substituted with an aryl group (eg, an aromatic or heteroaromatic group).

[0119] "Alkylthio" refers to an alkyl group as defined above with a sulfur atom attached thereto. In a preferred embodiment, the "alkylthio" moiety is represented by one of -S-alkyl, -S-alkenyl and -S-alkynyl. Representative alkylthio groups include methylthio, ethylthio, etc. The term "alkylthio" also encompasses cycloalkyl groups, alkene groups, cycloalkene groups and alkyne groups. "Arylthio" refers to an aryl or heteroaryl group.

[0120] The term "sulfinyl" refers to the radical -SO-. It should be noted that the sulfinyl radical can be further substituted with various substituents to form different sulfinyl groups, including sulfinic acid, sulfinamide, sulfinyl ester, sulfoxide, and the like.

[0121] The term "sulfonyl" refers to the radical -SO2-. It should be noted that the sulfonyl radical can be further substituted with various substituents to form different sulfonyl groups, including sulfonic acid (-SO3H), sulfonamide, sulfonate, sulfone, etc.

[0122] The term "thiocarbonyl" refers to the radical -C(S)-. It should be noted that the thiocarbonyl radical can be further substituted with various substituents to form different thiocarbonyl groups, including thioacids, thioamides, thioesters, thioketones, and the like.

[0123] As used herein, the term "amino" means -NH2. The term "alkylamino" means a nitrogen moiety having at least one straight or branched unsaturated aliphatic, cyclic or heterocyclic radical attached to the nitrogen. For example, representative amino groups include -NH2, -NHCH3, -N(CH3)2, -NH(C1-C 10 Alkyl), -N(C1-C 10 The term "alkylamino" includes "alkenylamino", "alkynylamino", "cyclylamino" and "heterocyclylamino". The term "arylamino" means a nitrogen moiety having at least one aryl radical attached to nitrogen. For example, -NHaryl and -N(aryl). The term "heteroarylamino" means a nitrogen moiety having at least one heteroaryl radical attached to nitrogen. For example, -NHheteroaryl and -N(heteroaryl). Optionally, the two substituents together with the nitrogen may also form a ring. Unless otherwise stated, the amino moiety-containing compounds described herein may include protected derivatives thereof. Suitable protecting groups for the amino moiety include acetyl, tert-butyloxycarbonyl, benzyloxycarbonyl, and the like.

[0124] The term "aminoalkyl" means alkyl, alkenyl and alkynyl as defined above, except that one or more substituted or unsubstituted nitrogen atoms (-N-) are located between the carbon atoms of the alkyl, alkenyl or alkynyl. For example, (C2-C6)aminoalkyl refers to a chain containing 2 to 6 carbon atoms and one or more nitrogen atoms located between the carbon atoms.

[0125] The term "alkoxyalkoxy" means -O-(alkyl)-O-(alkyl), for example, -OCH2CH2OCH3 and the like.

[0126] The term "alkoxycarbonyl" refers to -C(O)O-(alkyl), for example, -C(=O)OCH3, -C(=O)OCH2CH3, and the like.

[0127] The term "alkoxyalkyl" means -(alkyl)-O-(alkyl), for example, --CH2OCH3, -CH2OCH2CH3, and the like.

[0128] The term "aryloxy" means -O-(aryl), for example, -O-phenyl, -O-pyridyl, and the like.

[0129] The term "arylalkyl" means -(alkyl)-(aryl), for example, benzyl (ie, -CH2phenyl), -CH2-pyridyl, and the like.

[0130] The term "arylalkoxy" refers to -O-(alkyl)-(aryl), for example, -O-benzyl, -O-CH2-pyridyl, and the like.

[0131] The term "cycloalkoxy" refers to -O-(cycloalkyl), for example -O-cyclohexyl and the like.

[0132] The term "cycloalkylalkoxy" refers to -O-(alkyl)-(cycloalkyl), for example -OCH2cyclohexyl and the like.

[0133] The term "aminoalkoxy" refers to -O-(alkyl)-NH2, for example, -OCH2NH2, -OCH2CH2NH2, and the like.

[0134] The term "monoalkylamino" or "dialkylamino" means -NH(alkyl) or -N(alkyl)(alkyl), respectively, for example -NHCH3, -N(CH3)2, and the like.

[0135] The term "monoalkylaminoalkoxy" or "dialkylaminoalkoxy" means -O-(alkyl)-NH(alkyl) or -O-(alkyl)-N(alkyl)(alkyl), respectively, for example, -OCH2NHCH3, -OCH2CH2N(CH3)2, and the like.

[0136] The term "arylamino" refers to -NH(aryl), for example -NH-phenyl, -NH-pyridyl, and the like.

[0137] The term "arylalkylamino" refers to -NH-(alkyl)-(aryl), for example, -NH-benzyl, -NHCH2-pyridyl, and the like.

[0138] The term "alkylamino" means -NH(alkyl), for example -NHCH3, -NHCH2CH3, and the like.

[0139] The term "cycloalkylamino" refers to -NH-(cycloalkyl), for example -NH-cyclohexyl and the like.

[0140] The term "cycloalkylalkylamino" refers to -NH-(alkyl)-(cycloalkyl), for example, -NHCH2-cyclohexyl and the like.

[0141] With respect to all definitions provided herein, it should be noted that the definitions should be interpreted as open ended in the sense that other substituents beyond those specified may be included. Thus, C1 alkyl indicates the presence of a carbon atom but does not indicate what substituent is on that carbon atom. Thus, C1 alkyl includes methyl (i.e., -CH3) as well as -CR a R b R c , where R a , R b and R c Each of them may be independently hydrogen or any other substituent in which the atom alpha to the carbon is a heteroatom or a cyano group. Thus, CF3, CH2OH and CH2CN are all C1 alkyl groups.

[0142] Unless otherwise indicated, structures depicted herein are intended to include compounds that differ only in the presence of one or more isotopically enriched atoms. For example, in addition to replacing a hydrogen atom with deuterium or tritium or with an isotopically enriched 13 C- or 14 Compounds having the present structure except for the substitution of the carbon atom of C- for the carbon atom are within the scope of the present invention.

[0143] In various embodiments, the compounds of the invention disclosed herein can be synthesized using any synthetic method available to those skilled in the art. Non-limiting examples of synthetic methods for preparing various embodiments of the compounds of the invention are disclosed in the Examples section herein.

[0144] "Hi-C" refers to a technique in which chromatin is cross-linked with formaldehyde, then digested and re-ligated in such a way that only covalently linked DNA fragments form ligation products. It is based on chromosome conformation capture and can be used to comprehensively examine chromatin interactions in mammalian cell nuclei.

[0145] "Diaziridine" refers to a class of organic molecules, which are formed by combining a carbon atom with two nitrogen atoms, which are double-bonded to each other to form a cyclopropene-like ring, 3H-diaziridine. They are isomeric with the diazo carbon group, and like it, can be used as a precursor of carbene by losing a dinitrogen molecule. For example, irradiating diaziridine with ultraviolet light can cause carbene to be inserted into various CH, NH and OH bonds. Without wishing to be bound by a particular theory, the photoactivation of diaziridine generates a reactive carbene intermediate. Such intermediates can form covalent bonds at a distance corresponding to the length of the spacer arm of a particular reagent by addition reactions with any amino acid side chains or peptide backbones (e.g., proteins or other molecules containing nucleophilic or active hydrogen groups R').

[0146] In some embodiments, the diaziridine is:

[0147] In some embodiments, the diaziridine is: Where R d2 and R d3 Each is independently H, halo, CH3, CF3, optionally substituted alkyl, optionally substituted heteroalkyl, optionally substituted aryl, optionally substituted heteroaryl, optionally substituted cyclyl or optionally substituted heterocyclyl.

[0148] In some embodiments, the diaziridine is: Where R d1 is H, halo, CH3, CF3, optionally substituted alkyl, optionally substituted heteroalkyl, optionally substituted aryl, optionally substituted heteroaryl, optionally substituted cyclyl or optionally substituted heterocyclyl.

[0149] Psoralens are embedded in the DNA double helix structure, where they are ideally positioned to form one or more adducts with adjacent pyrimidine bases (preferably thymine) under UV photon excitation. Psoralens can also be activated by irradiation with long wavelength UV light. Although light in the UVA range is the clinical standard, UVB is more efficient in forming photoadducts. The photochemical reaction sites in psoralens are olefin-like carbon-carbon double bonds in the furan ring (five-membered ring) and the pyrone ring (six-membered ring). Without wishing to be bound by a particular theory, when appropriately embedded near a pyrimidine base, a four-center photocycloaddition reaction can lead to the formation of either of two cyclobutyl-type monoadducts. Typically, furan side monoadducts are formed in higher proportions. The furan monoadduct can absorb a second UVA photon, resulting in a second four-center photocycloaddition at the pyrone end of the molecule, thereby forming a diadduct or crosslink. The pyrone monoadduct does not absorb in the UVA range and therefore cannot form crosslinks by further UVA irradiation.

[0150] "Azobiotin-azide" refers to a linker containing a biotin moiety connected to an azide group via a spacer arm containing a diazo group that can be cleaved with a 50 mM sodium dithionite solution. Azo compounds are compounds with the functional group diazenyl RN=NR', where R and R' can be aryl or alkyl.

[0151] Herein, we developed a class of molecular probes that bind DNA and become activated upon irradiation with long-wavelength UVA (~360 nm) to efficiently crosslink nearby bound DNA and proteins ( Figure 1D). These molecular probes will serve as powerful tools for covalently capturing protein-DNA and protein-RNA complexes in in situ studies of a wide range of protein-nucleic acid interactions. This new class of molecules (also known as bifunctional photocrosslinking (BFPX) probes) allows for selective, efficient, and robust (stable) capture of protein-nucleic acid complexes within cells. Compared with existing crosslinking technologies, the probes developed in this study will have the following new and useful features: (1) High specificity: Unlike formaldehyde and other bifunctional chemical crosslinking probes that react with lysine (such as DSS), which will react with any protein in the cell and cause damage to the DNA binding surface of TF and antibody epitopes, our probes will preferentially bind only to nucleic acids through nucleic acid-specific recognition heads, thereby achieving regional selectivity for photochemical reactions at or near DNA binding sites. (2) High efficiency: We can introduce highly efficient photoaffinity labeling groups (such as diazirine) that can be activated by long-wavelength UVA (~360nm), thereby reducing the photodamage associated with UVC (250nm) irradiation. (3) Versatility and tunability: Through synthetic engineering, variable functional heads or linkers can be introduced to enable customized applications. For example, multiple arms of a photoaffinity tag group can be introduced to capture more than one TF, as in the NFAT / Fos-Jun / DNA ternary complex, allowing direct experimental determination of multi-TF complexes rather than relying on sequential ChIP-seq, which is not feasible for most TFs using current methods. Linker length can be varied to capture proximal DNA binding domains or more distal protein cofactors recruited by the TF. (4) Temporal and spatial selectivity: Crosslinking can be initiated by UV irradiation that is controlled in time and focus, allowing potential temporal and spatial control to capture protein / DNA complexes in selected temporal and subcellular regions. (5) Allows the study of the proteome bound to the genome: Due to the enhanced specificity of cross-linking only proteins bound to DNA, in addition to identifying DNA sequences bound by proteins of interest (such as traditional ChIP-seq), reversible linkers, isotope labels, and mass spectrometry can be used to label and identify all proteins bound to DNA in the entire genome after UV irradiation (conceptually the reverse process of ChIP-seq). This will provide unprecedented information on protein-DNA interactions on a genome-wide and proteome-wide scale.

[0152] We have developed a class of molecular probes that bind to DNA or RNA and become activated upon irradiation with long-wavelength UV (~350 nm or about 365 nm) to efficiently crosslink to nearby bound DNA and proteins ( Figure 2A , Figure 2BThese molecular probes will serve as powerful tools for the covalent capture of protein-DNA and protein-RNA complexes with high efficiency, selectivity (i.e., targeting only DNA-binding proteins), and stability (to allow subsequent XL-MS analysis) in situ studies of a wide range of protein-nucleic acid interactions.

[0153] To overcome the inherent disadvantages of formaldehyde cross-linking (which cannot directly cross-link proteins to DNA because amines in DNA are inactive), we designed bifunctional probes to endow DNA molecules and / or RNA molecules with true active primary amines. Through amine linkage (e.g., by click chemistry or amine NHS that converts amines to azides), a photoreactive head (functional group) is introduced to cross-link to nearby proteins, thus truly capturing and only capturing nucleic acids in close proximity to proteins. This provides a more precise location of chromatin proteins bound on DNA and preserves epitopes on proteins because it is not cross-linked to every protein it encounters like formaldehyde (in chromatin immunoprecipitation, ChIP, or Hi-C). This can also provide proteins with IDs for mass spectrometry identification (reverse ChIP).

[0154] Taking DNA as an example, studying protein binding to DNA throughout the genome in cells has been key to understanding cellular mechanisms. The conventional procedure in these studies is to covalently fix proteins to DNA via formaldehyde cross-linking, so that protein-bound DNA sites can be analyzed by ChIP-seq, or protein-mediated chromatin interactions can be analyzed by Hi-C, HiChIP, or similar techniques. However, increasing evidence shows that formaldehyde cross-linking has major limitations, including non-specific modification of proteins, thereby destroying the epitopes of antibodies used in ChIPseq, or inefficient capture of transcription factors such as Lac repressors and NF-κB.

[0155] To overcome these limitations, we developed a novel method for in situ covalent capture of protein-nucleic acid complexes using bifunctional photocrosslinking probes, providing a new class of molecular tools that possess two photocrosslinking functional groups, hence named bifunctional photocrosslinkers (BFPX).

[0156] In various embodiments, one such functional group-DNA photocrosslinker is formed from a non-specific DNA binding small molecule (e.g., psoralen, DAPI, benzophenone, etc. - for genome-wide capture of all protein / DNA complexes) or a specific DNA binding small molecule (e.g., polyamide - for targeted capture of all protein / DNA complexes) that has the inherent ability to crosslink to DNA at long UV wavelengths (~360 nm, such as psoralen; or between 330 nm and 370 nm) or contains a synthetically introduced photocrosslinking moiety (e.g., benzophenone or phenyldiaziridine). Other embodiments provide that the nucleic acid binding head / functional group is derived from methyl trimethoxalen; Hoechst dye, such as Hoechst 33342 (2'-[4-ethoxyphenyl]-5-[4-methyl-1-piperazinyl]-2,5'-bis-1H-benzimidazole trihydrochloride trihydrate), Hoechst 33258, Hoechst 34580, Hoechst S769121; polyamide; or a G-quadruplex binding molecule.

[0157] The DNA crosslinker will be connected to the protein capture photocrosslinker through a linker, which can be a single arm of diazirine, or multiple arms of multiple diazirine groups (for multiplex capture of more than one protein bound to DNA together).

[0158] The linker length can be of variable length to capture proximal direct DNA binding proteins (shorter linker length) or protein cofactors that are more distal to the DNA but recruited by the DNA binding proteins (longer linker length) and to improve capture efficiency. (Too short a linker will result in crosslinking back to the DNA, while too long a linker will capture non-specific proteins that are not directly or indirectly associated with the DNA). The linker can also be engineered to include a tag (e.g., an alkyne group) that can be used to enrich the captured protein / DNA complexes by click chemistry, or for fluorescent labeling to track the probe distribution within the cell by an azido fluorophore.

[0159] In various embodiments, the linker may include one or more of the following: (i) a cleavable bond, (ii) an oligomer or polymer having -OCH2CH2- repeating units, and (iii) an unsaturated moiety. Exemplary unsaturated moieties include, but are not limited to, carbon-carbon double bonds, triple bonds, and aromatic groups (e.g., Figure 3D ).

[0160] The linker can also be designed to be cleavable so that proteins cross-linked to the DNA complex can be released for mass spectrometry analysis. This is an important aspect of these embodiments, which will improve technologies designed to capture protein-bound DNA sequences (such as ChIP-seq and Hi-C) and open up new areas for mapping all proteins that bind to DNA (DNA-bound proteome). An example of a probe with a cleavable linker is by coupling NHS-SS-diaziridine (SDAD) with 4AMT to produce 4AMT-SS-SDAD. Similar to 4AMT-LC-SDA, 4AMT-SS-SDAD can be used to cross-link proteins to DNA. However, the added advantage of 4AMT-SS-SDAD is that after purification of the cross-linked protein-DNA complex, the protein or its protease-digested peptide fragments can be released by cleavage of the linker to facilitate subsequent protein analysis (such as by mass spectrometry). In addition to the disulfide linkers described, other cleavable connections ( Figure 3C ). Using these cleavable linkers, BFPX probes can be used not only to detect DNA sequences bound by a given protein (e.g., ChIP-seq analysis), but also to perform whole-genome analysis of all proteins that bind to DNA. In various embodiments, a sulfoxide-containing MS-cleavable cross-linker is used to introduce an MS-cleavable functional group into the linker of the BFPX probe; or the linker of the BFPX probe contains a sulfoxide-containing MS-cleavable CS bond. The two symmetrical CS bonds (introduced into the linker of the BFPX probe by these sulfoxide-containing MS-cleavable cross-linkers) can be used in tandem mass spectrometry (MS / MS or MS 2) during the MS cleavage using collision induced dissociation (CID), for example by implementing high energy collision dissociation (HCD) and / or electron transfer dissociation (ETD). This results in physical separation of nucleic acid-protein crosslinks, thereby generating unique peptide fragment pairs with defined mass relationships. Exemplary sulfoxide-containing MS cleavable crosslinkers for this purpose include DSSO (bis-(propionic acid NHS ester)-sulfoxide, bis(2,5-dioxopyrrolidin-1-yl) 3,3'-sulfinyl dipropionate), d0-DMDSSO (bis(2,5-dioxopyrrolidin-1-yl) 3,3'-sulfinyl bis(2-methylpropionate)), DHSO (3,3'-sulfinylbis(propane hydrazide); dihydrazide sulfoxide), BMSO (3,3'-sulfinylbis(propane hydrazide); dihydrazide sulfoxide). Acylbis(N-(2-(2,5-dioxo-2,5-dihydro-1H-pyrrol-1-yl)ethyl)propionamide)), alkyne-A-DSBSO(bis(2,5-dioxopyrrolidin-1-yl) 3,3′-((2-(but-3-yn-1-yl)-2-methyl-1,3-dioxane-5,5-diyl)bis(methylenesulfinyl)) dipropionate; alkyne-labeled acid-cleavable disuccinimidyl bissulfoxide) and azide-A-DSBSO.

[0161] Compared with formaldehyde cross-linking, our design has the following advantages: (1) High specificity: Unlike formaldehyde, which will react with any protein in the cell (which will lead to epitope damage and over-fixation of cellular protein complexes, causing difficulties in reversing cross-linked complexes), our probe will preferentially bind to DNA. After washing away excess unbound probe, the DNA-binding probe will only cross-link to DNA and nearby DNA-binding proteins; (2) Temporal and spatial selectivity: Cross-linking can be initiated by UV irradiation controlled in time and focus, allowing potential temporal and spatial control to capture protein / DNA complexes at selected times and subcellular regions; (3) Due to the enhanced specificity of cross-linking only proteins bound to DNA, in addition to identifying DNA sequences bound by proteins of interest (as in traditional ChIP-seq), reversible linkers, isotope labeling, and mass spectrometry can be used to label all proteins bound to DNA in the entire genome under UV irradiation (conceptually the reverse process of ChIP-seq, but for mapping the DNA-bound proteome). This will provide unprecedented information on protein-DNA interactions at the whole genome and whole proteome scale.

[0162] Various embodiments provide that the probes described herein are not limited to those containing psoralen (or isopsoralen, or derivatives such as xanthoxalin / methoxsalen, bergamotolactone, imperatorin, and nodakenetin) as the sole nucleic acid-binding photocrosslinkable group ( Figure 1A-1E , Figure 2A-2C), and the probe is not limited to binding only double-stranded DNA. A variety of other nucleic acid binding molecules can be used to form the nucleic acid binding head of the bifunctional probe of the present invention, as long as they are equipped with a linking group such as an amine or an azido group. Their nucleic acid targets can be molecularly specific.

[0163] In one embodiment, the nucleic acid binding head of the bifunctional probe is derived from 4',6-diamidino-2-phenylindole (DAPI), which binds dsDNA genome-wide and non-specifically and can be functionalized to click chemistry alkynes via an azido linker group ( Figure 3A-3E ). DAPI can easily and strongly bind to DNA even without photochemistry.

[0164] In other embodiments, the nucleic acid binding head of the bifunctional probe is a photoreactive nucleic acid binder, which can be phenylazide, phenyldiaziridine, or benzophenone - all of which can bind DNA under similar spectral illumination (300-360 nm) - and is modified by attaching an amine / azide group as the other functional head that reacts with proteins ( Figure 3A-3E ).

[0165] A further embodiment provides that the nucleic acid binding head of the bifunctional probe is a sequence-specific binding agent, such as a polyamide, which can be further modified to have an amine / azide linker group as another functional head that reacts with a protein ( Figure 3A-3E ). Suitable nucleic acid binding heads / functional groups are derived from pyrrole-imidazole polyamides, or hairpin polyamides containing N-methylpyrrole (Py), N-methylimidazole (Im), and N-methyl-3-hydroxypyrrole (Hp) residues to bind to specific predetermined sequences in nucleotides.

[0166] Among them, psoralen, phenylazide, phenyldiaziridine or benzophenone polyamide can also be used for RNA binding or single-stranded DNA binding.

[0167] In addition, special molecules that recognize specific nucleic acid targets (e.g., G-quadruplex recognition using polycyclic structures of anthracene or quinolinium derivatives) can also be modified into bifunctional probes via amino / azide linkages to more efficiently crosslink with nearby biomolecules ( Figure 3A-3C In addition, ethoxybutyraldehyde (1,1-dihydroxy-3-ethoxy-2-butanone), which binds guanine at the N1 and N2 positions, can be used as a nucleic acid binding head to develop BFPX probes targeting single-stranded DNA and RNA.

[0168] In addition, using this design strategy, the bifunctional probes can be further developed to have multiple other functions ( Figure 3B), for example for conjugation with cleavable azido-diazobiotin (as enrichment handle) and / or with fluorophore azide to allow detection / identification of spatial location. The disclosed bifunctional / multifunctional probes impart amine groups to DNA and the ability to photocrosslink nearby proteins via diazirine, and can be modified to molecules with azide groups at each end. When coupled to a multi-arm core, they can be assembled in a customized combination, such as by copper-free bioorthogonal click chemistry on a cyclooctyne core, for use in vivo or in vitro.

[0169] The photocrosslinking molecules (BFPX) provided in this article are suitable for a wide range of genomic research tools (such as ChIPmentation, Cut&Run, HiChIP), which traditionally rely on formaldehyde crosslinking but are also limited by formaldehyde crosslinking.

[0170] We have developed a class of molecular probes that are cell permeable, inert to cellular molecules, bind to DNA, and become activated upon irradiation with long-wavelength UVA (~360 nm) to covalently crosslink DNA and nearby DNA-bound proteins. These molecular probes can capture protein-nucleic acid complexes in vitro and in situ for a wide range of analyses at both the bulk and single-molecule levels.

[0171] For a given cell in a specific state, which proteins bind to the genome and where they bind on the genomic DNA are fundamental questions for understanding cellular function and disease mechanisms. Most protein-nucleic acid complexes are too unstable to be isolated from their native cellular environment due to their inherent, mainly ionic and hydrophilic interactions. A variety of techniques have been developed to capture protein-DNA complexes for biochemical and imaging analysis. Most of these methods rely on capturing protein-DNA complexes by formaldehyde crosslinking.

[0172] However, there is growing evidence that formaldehyde fixation may be a major problem undermining the effectiveness of current methods.

[0173] This is mainly due to the low efficiency of formaldehyde in cross-linking DNA to proteins, as well as its nonspecific damage and high reactivity to proteins (especially on DNA binding surfaces and antibody binding epitopes). Most of the DNA captured by formaldehyde is not covalently attached to proteins, but is trapped in fixed protein complexes, resulting in the capture of a large number of nonspecific DNA fragments. Direct UV cross-linking of proteins to DNA and RNA has been reported, but these methods are limited by low cross-linking efficiency and the use of short-wavelength UVC (~250nm) that damages proteins and nucleic acids. Here, we designed and synthesized a class of molecular probes that bind (or embed) to DNA (or RNA) and are enriched on DNA (or RNA), and can be activated by long-wavelength UVA (~360nm), thereby forming a covalent link between DNA or RNA and the protein to which it is bound. These molecular probes are biocompatible (inert to cell molecules in the absence of UV), can permeate cells, and after UVA activation, can capture protein-DNA (or protein-RNA) complexes within cells with high specificity and stability.

[0174] Our general strategy is to engineer molecular probes that have two different photocrosslinking groups, hence termed bifunctional photocrosslinkers (BFPX) ( Fig. 12A ). One of the functional groups is responsible for binding and cross-linking DNA or RNA under UVA irradiation, while the other functional group is responsible for covalently capturing proteins bound or recruited to DNA through a UVA-activated photochemical reaction. The two functional groups are connected by a linker that is engineered to introduce a molecular handle that is responsible for monitoring / labeling, separation, and analysis of cross-linked protein-DNA complexes.

[0175] For DNA binding and cross-linking heads, many natural or synthetic DNA binding molecules can be used to capture any protein / DNA complex genome-wide, including Hoechst dyes, psoralen or derivatives, 4',6-diamidino-2-phenylindole (DAPI), which bind to DNA almost non-specifically. Alternatively, specific DNA binding molecules (such as sequence-specific DNA binding polyamides or G-quadruplex binders) can be used to target the capture of protein-DNA complexes bound to specific genomic sites of interest. In the current study, we chose cell-permeable and non-toxic psoralen as the DNA anchoring head. Psoralen and its derivatives can bind and intercalate DNA almost non-specifically, with a binding site preference of 5'-TA>5'-AT>>5'-TG>5'-GT', and covalently cross-link with DNA with high efficiency (up to 80%) under long-wavelength (~360nm) UV irradiation.

[0176] Psoralens have a moderate affinity for double-stranded DNA and RNA (Kd ~ μM), resulting in their enrichment on DNA without disrupting DNA binding of most transcription factors (Kd ~ nM). For the protein capture head, any photochemically activatable group that forms a covalent linkage with the protein can be considered. In the current study, we chose the well-established photoaffinity labeling group diazirine, which has a long UVA wavelength activation spectrum (340-365nm) similar to that of psoralens. The two heads are connected by a synthetic linker with variable length, flexibility, and cleavability. Following the above design principles, we synthesized a series of BFPX probes with various linker lengths and functional characteristics ( Fig. 12B ) (For details, see the Examples section of this article).

[0177] We first used in vitro assembled transcription factor-DNA complexes in nondenaturing (electrophoretic mobility shift assay, EMSA, Fig. 12C ) and denaturation (SDS page, Fig.12D ) conditions, the activity of the BFPX probe was tested.

[0178] like Fig. 12C As shown, UV alone (lane 4) and in combination with increasing concentrations of the BFPX probes SPB-AAD (lanes 8-10), SPB-PEG4-AAD (lanes 11-13), and SPB-spermidine-AD (lanes 14-16) showed little effect on NFAT1 (nuclear factor of activated T cells) binding to DNA compared to the control (lane 3). In contrast, formaldehyde (lanes 5-7) showed dose-dependent inhibition of NFAT1 / DNA binding interactions and reduced most NFAT / DNA complexes (lane 7) at a typical concentration of 1% used in chromatin immunoprecipitation (ChIP) experiments. Similar experiments using other transcription factors showed that concentrations of up to 500 μM of the BFPX probe did not affect DNA binding of GATA3 in EMSA, while formaldehyde as low as 0.1% (v / v) inhibited DNA binding (data not shown).

[0179] When analyzed by denaturing SDS-PAGE Fig. 12C The same binding reaction in Fig.12D), the native NFAT1 / DNA complex dissociated as expected (lane 3), and formaldehyde cross-linking did not produce a stable protein-DNA complex (lanes 5-7), which is consistent with the view that formaldehyde cannot cross-link proteins to DNA in vitro. In contrast, all three BFPX probes SPB-AAD (lanes 8-10), SPB-PEG4-AAD (lanes 11-13), and SPB-spermidine-AD (lanes 14-16) showed dose-dependent cross-linking of NFAT1 / DNA complexes, which were stable under strong denaturing conditions (2% SDS, 95°C for 10 min), indicating that there is a stable covalent association between protein and DNA. The cross-linked NFAT / DNA complex can be well resolved by SDS gradient gels according to its expected molecular weight and can be stained by a variety of methods, including cyanine dyes (SYBR Safe) and Coomassie Brilliant Blue ( Fig.12E ).

[0180] SYBR Safe stains protein-DNA complexes more effectively than free proteins, whereas Coomassie Brilliant Blue stains free proteins more effectively than protein-DNA complexes ( Fig.12E ). Therefore, it is best to use fluorescently labeled DNA for quantitative analysis of cross-linking reactions. The apparent cross-linking efficiency (intensity of the complex band relative to the total DNA intensity) can be as high as 22.5% ( Fig.12D , lane 16). When normalized to total protein / DNA complexes (i.e., only 48% of the total DNA in lane 16 was present as protein-DNA complexes, see Fig. 12C ), under these experimental conditions, the cross-linking efficiency was estimated to be 47%. BFPX cross-linking efficiency varies depending on the protein's DNA affinity, the probe used, the binding conditions (salt and buffer), and the DNA sequence flanking the protein binding site. The cross-linking reaction is also dose-dependent on the UV irradiation time ( Fig.12F , Figure 12G ). Cross-linked protein-DNA complexes were observed as early as 3 seconds, and the reaction reached a plateau almost within 100 seconds. These observations raise the possibility of using BFPX to study the dynamic processes of protein-DNA interactions in cells. We have demonstrated BFPX-based cross-linking of a wide range of transcription factor / DNA complexes that belong to different DNA binding domain families and bind to DNA from either the major or minor groove. These include MEF2 ( Figure 12G )、Nkx2.5( Fig.12H ) and p53( Fig.12I) and GATA3 (not shown). Based on the molecular size, the cross-linked protein-DNA complex contains one strand of the duplex DNA, indicating that the psoralen head of the BFPX probe forms a monoadduct with the DNA. Psoralen-DNA cross-links can be reversed by short UVC (254 nm) or alkaline heating. However, we found that these literature-reported reversal procedures caused significant damage to both the protein and the DNA (data not shown). Another approach to release cross-linked proteins is to use probes with cleavable links, as we have shown in Fig.12J As shown in Figure 2 using 4AMT-SDAD. BFPX-based crosslinking is strictly dependent on DNA binding, because a large number of proteins in the binding solution that do not bind DNA (e.g., BSA) are not crosslinked, and when DNA is bound to homologous proteins, non-homologous DNA binding proteins are not crosslinked (data not shown). The specificity of the crosslinking reaction is further demonstrated at the structural level (see below).

[0181] We use Fig.13A The workflow shown in analyzes the covalent attachment sites on the NFAT / DNA complex cross-linked by SPB-AAD and SPB-PEG4-AAD. The cross-linked NFAT / DNA complex is purified on FPLC using a Mono-Q column, digested with trypsin, and purified by another round of Mono-Q on FPLC. The DNA attached with peptides is eluted like free DNA. The DNA-peptide conjugate is then sequenced by Edman degradation. The sequencing results are usually noisy, with multiple amino acids detected per cycle. However, a major peptide can be identified for each SPB-AAD and SPB-PEG4-AAD cross-linked complex. For the SPB-AAD cross-linked complex, the peptide (IVGN) is uniquely matched to the NFAT1 trypsin fragment between the 478th and 497th residues, while the peptide (APTGGH) from the SPB-PEG4-AAD cross-linked complex is uniquely matched to the trypsin fragment between the 434th and 452nd residues. These results were compared with the crystal structure of the NFAT / DNA complex (pdb accession 1a02, Chen et al., Nature 1998). If we assume that the BFPX probe inserts into the 5'TpA-3' site on the DNA and that the linker adopts an extended conformation, the diaziridine head group of SPB-AAD would be adjacent to the NFAT loop (residues 478-497) close to the DNA ( Fig. 13B ), while the diaziridine head group of SPB-PEG4-AAD will be adjacent to the NFAT loop (residues 434-452) which is relatively far from the DNA ( Fig. 13C). These structure-based predictions match well with the Edman sequencing results. These analyses strongly suggest that, despite synthetic modifications, the psoralen head in the BFPX probe retains its preference for binding to the 5'-TpA-3' site and that the carbene generated by UV activation of diazirine favors crosslinking to proteins in a proximity-dependent manner. These studies also demonstrate the utility of the BFPX probe in structural analysis of protein-DNA interactions in captured protein-DNA complexes both in vitro and in situ.

[0182] Next, we used azide-containing fluorescent molecules to test the cellular and nuclear permeability of the BFPX probes through click reactions with terminal alkynes in SPB-AAD and SPB-PEG4-AAD. Briefly, Hela cells were incubated with SPB-AAD or SPB-PEG4-AAD (10 μM or 100 μM) in the dark for 30 min to allow DNA binding and intercalation through the psoralen portion of the BFPX probe. After washing away excess BFPX probe and irradiating with UV 365 nm for 5 min, Alexa Fluor 647 pyridylmethyl azide was coupled to the alkyne group on SPB-AAD or SPB-PEG4-AAD through a click reaction. After click chemistry labeling, unreacted fluorescent molecules were removed by washing. The cell nuclei were then analyzed using fluorescence imaging ( Fig.14A When no BFPX probe was added (columns 1 and 2) or the probe was added without UV treatment (columns 3 and 6), almost no fluorescence signal remained in the cells. In contrast, only when the BFPX probe was present and activated by UV cross-linking, the Alexa Fluor 647 picolyl azide molecules could be immobilized in the cells via click reactions with SPB-AAD (columns 4 and 5) or SPB-PEG4-AAD (columns 7 and 8), and the immobilized fluorescence signal was dose-dependent on the BFPX probe ( Fig. 14B ).

[0183] Finally, we tested the capture of transcription factors bound to genomic DNA within cells. HEK293T cells transfected with FLAG-tagged FOXP3 were treated with UV irradiation and various combinations of probes (SPB-AAD: A; and P: SPB-PEG4-AAD). Cells were lysed in 1x RIPA buffer followed by RNase and MNase treatment. Whole cell lysates were then added with SDS loading dye, heated at 95°C and run on 4%-15% SDS PAGE gels and analyzed by western blotting using anti-FLAG antibodies.

[0184] like Fig. 14CAs shown, higher molecular weight FOXP3-containing species were observed in cells treated with SPB-AAD (A) and SPB-PEG4-AAD (P) (lanes 4-9) compared with controls (lanes 1-3), and the formation of cross-linked FOXP3 complexes was dose-dependent on probe concentration (compare lane 4 (SPB-AAD, 10 μM, A10) with lane 6 (SPB-AAD, 100 μM, A100)) and UV exposure time (compare lane 4 (2 min) with lane 5 (5 min) at 10 μM SPB-AAD (A10); also compare lane 8 (2 min) with lane 9 (5 min) at 100 μM SPB-PEG4AAD (P100)). We also used TRIzol to extract genomic DNA from BFPX-treated and control cells. TRIzol uses strong denaturing agents (e.g. 4M guanidine thiocyanate) to lyse cells and completely dissociate nucleoprotein complexes. The extracted genomic DNA is then extensively digested with Benzonase and analyzed by SDS PAGE gel and western blot. Fig.14D As shown, using the same HEK293T cells transfected with FLAG-tagged FOXP3, it is clear that only cells treated with UV and probe (SPB-AAD, 100 μM, A100) show FOXP3 protein (lanes 4 and 5) compared to the control (lanes 1-lane 3). This observation again suggests that FOXP3 protein is covalently linked to genomic DNA in BFPX-treated cells, which can be extracted by TRIzol protocol under strong denaturing conditions, while in control cells, the non-covalent nucleoprotein complex does completely dissociate. Interestingly, Benzonase digestion not only releases FOXP3 with a molecular weight similar to its free monomer, but also releases a larger molecular weight species, suggesting that FOXP3 may bind to certain genomic regions in a higher order structure that is resistant to Benzonase digestion. We further tested the capture of endogenous transcription factors in GM12878 using MEF2C as a target. GM12878 cells (7.5 million per sample) were treated with different combinations of UV irradiation and BFPX probes, and genomic DNA was extracted using the TRIzol protocol, digested with Benzonase, and analyzed by SDS-PAGE / Western blotting ( Fig.14E). Compared to the transfection system, the signal for endogenous transcription factors is substantially weaker due to low natural abundance. However, it is clear that MEF2C covalently bound to genomic DNA and extracted by TRIzol under strong denaturing conditions is observed only in cells treated with UV and probe (lanes 3, 4, and 5) compared to the control (lane 1). The UV-only control (lane 2) shows a weak signal, indicating that UV alone may crosslink some proteins to DNA at a low level. But in cells treated with BFPX, the crosslinking efficiency is significantly higher than the UV-only control (compare lanes 3-5 with lane 2).

[0185] The BFPX strategy is mechanically defined, and its modular design enables the development of general and customized tools for capturing protein-nucleic acid complexes in vitro and in cells for overall and single-molecule analysis. Unlike formaldehyde and other lysine-reactive bifunctional chemical cross-linking probes that react nonspecifically and uncontrollably with any protein in the cell, BFPX probes preferably bind only to nucleic acids through specific recognition heads, thereby achieving regional selectivity of photochemical reactions at or near DNA binding sites. Cross-linking can be initiated by UV irradiation at specific time points and concentrated in specific subcellular locations, allowing for temporal and spatial resolution of protein / nucleic acid interactions in situ. In addition to identifying the DNA sequence bound by the protein of interest, reversible linkers, isotope labeling, and mass spectrometry can also be used to identify all proteins bound to DNA throughout the genome at a specific time point. These new features provided by BFPX will provide unprecedented information on protein-DNA interactions at the whole genome and whole proteome scale.

[0186] Various embodiments of the present invention

[0187] Implementations include those listed below.

[0188] Embodiment 1. Photo-crosslinking molecules of formula (I): A-L1-(C) n -((L2) n -B) m(I), wherein A represents a nucleic acid binding functional group derived from psoralen, methyl trimethoxalin, benzophenone, 4',6-diamidino-2-phenylindole (DAPI), Hoechst dye, polyamide or G quadruplex binding molecule or a derivative thereof; L1 is absent or represents a first linker; when n=1, C represents a core portion having at least two functional groups, each functional group being used to attach to L1 and to at least one arm represented by L2-B, respectively; or when n=0, C is absent; when n=1, L2 independently represents a second linker for each arm represented by L2-B; or when When n=0, L2 does not exist; for each of the arms, B independently represents: a photoreactive functional group, the photoreactive functional group includes diaziridine or a derivative thereof or an aryl azide or a derivative thereof, optionally, the aryl azide or a derivative thereof is selected from phenyl azide, o-hydroxyphenyl azide, m-hydroxyphenyl azide, tetrafluorophenyl azide, o-nitrophenyl azide, m-nitrophenyl azide or azidomethylcoumarin; or a detectable functional group; wherein in at least one of the arms, B represents a photoreactive functional group; n=0 or 1; m represents (L2) n -B represents the number of arms, wherein when n=1, m is an integer of 1 or more, or when n=0, m=1.

[0189] Embodiment 2. The photo-crosslinkable molecule of embodiment 1, wherein at least one of L1 and L2 is not absent, and the at least one of L1 and L2 is cleavable.

[0190] Embodiment 3. The photocrosslinkable molecule of embodiment 2, wherein L1, L2 or both independently comprise one or more of a sulfoxide-containing mass spectrometry (MS) cleavable bond, an acid cleavable CS bond, a disulfide group, and an azo group.

[0191] Embodiment 4. A photo-crosslinked molecule as described in any one of embodiments 1-3, wherein n=0, m=1, and the photo-crosslinked molecule is represented by formula (II): A-L1-B(II).

[0192] Embodiment 5. A photocrosslinked molecule as described in embodiment 4, wherein: A is an amine-containing or amine-reactive derivative of psoralen, methyl trimethsalen, benzophenone, DAPI, Hoechst dye, polyamide or G quadruplex binding molecule, optionally, A is derived from succinimidyl-[4-(psoralen-8-yloxy)]-butyrate (SPB) or 4'-aminomethyl trimethsalen (4AMT); B includes diazirine or diazirine alkyne, optionally aminodiazirine alkyne (AAD); and L1 is absent or is a first linker, wherein the first linker comprises one or more of the following: (i) a cleavable bond, (ii) an oligomer or polymer having -OCH2CH2- repeating units, and (iii) an unsaturated portion, optionally selected from a carbon-carbon double bond, a carbon-carbon triple bond or an aromatic group.

[0193] Embodiment 6. The photocrosslinking molecule of embodiment 5, wherein L1-B is derived from succinimidyl 6-(4,4'-azidopentanamido)hexanoate (NHS-LC-SDA), succinimidyl 2-((4,4'-azidopentanamido)ethyl)-1,3'dithiopropionate (NHS-SS-diaziridine) or 2-(3-(but-3-yn-1-yl)-3H-diaziridine-3-yl)ethane-1-amine (AAD); and / or wherein A is derived from 4'-aminomethyltrimethsalin (4AMT) or succinimidyl-[4-(psoralen-8-yloxy)]-butyrate (SPB); and wherein optionally, the photocrosslinking molecule is represented by formula (IIa) or formula (IIc):

[0194] Embodiment 7. A photocrosslinkable molecule as described in embodiment 5, wherein: A is derived from SPB or 4AMT; B comprises diaziridine or diaziridine alkyne, optionally aminodiaziridine alkyne (AAD); and L1 is a first linker, the first linker comprising one or more of the following: (i) a cleavable bond, (ii) an oligomer or polymer having -OCH2CH2- repeating units, and / or (iii) an unsaturated portion, the unsaturated portion optionally selected from a carbon-carbon double bond or an aromatic group; and wherein optionally, the photocrosslinkable molecule is represented by formula (IIb), (IId), (IIe) or (IIf):

[0195] Embodiment 8. The photocrosslinkable molecule of embodiment 5, wherein L1 comprises a length of 2 to 20 carbons or 20-100 carbons.

[0196] Embodiment 9. A photo-crosslinkable molecule as described in any one of embodiments 1-3, wherein n=1, m is an integer of 2 or greater, and C represents a core portion having at least three functional groups, each functional group being used to attach to L1 and to attach to at least two arms each represented by (L2-B), such that the photo-crosslinkable molecule is represented by formula (III):

[0197] Embodiment 10. A photocross-linked molecule as described in embodiment 9, wherein B comprises diaziridine or diaziridine azide in one of the at least two arms, and B represents a detectable functional group in the other of the at least two arms, wherein the detectable functional group comprises a fluorophore, biotin, a chromophore, a chromogen, a quantum dot, a fluorescent microsphere or a nanoparticle.

[0198] Embodiment 11. A photocrosslinkable molecule as described in embodiment 9 or 10, wherein L1, L2 or both independently comprise one or more of the following: (i) a cleavable bond, (ii) an oligomer or polymer having a repeating unit of -OCH2CH2-, and (iii) an unsaturated portion.

[0199] Embodiment 12. A photo-crosslinked molecule as described in any one of embodiments 9-11, wherein C represents a branched core portion, and the branched core portion comprises at least three surface functional groups, each surface functional group is used to attach to L1 and to attach to the at least two arms represented by L2-B respectively.

[0200] Embodiment 13. The photocrosslinkable molecule of any one of embodiments 9-12, wherein L1, L2, or both independently comprise a triazole bonded to A.

[0201] Embodiment 14. A method for cross-linking a nucleic acid to an adjacent protein in a system, the method comprising: incubating the system with the photocross-linking molecule of any one of embodiments 1-13, and irradiating the system with ultraviolet light.

[0202] Embodiment 15. The method of embodiment 14, wherein the system is a living cell.

[0203] Embodiment 16. A method as described in embodiment 14 or 15, wherein the wavelength of the ultraviolet light is between 300nm and 360nm.

[0204] Embodiment 17. The method of any one of embodiments 14-16, further comprising using the system to perform one or more of immunoprecipitation, chromatin precipitation, 3D chromatin conformation capture, mass spectrometry, and electrophoresis.

[0205] Embodiment 18. A method as described in any one of embodiments 14-17, wherein the elements L1, L2 or both of the photo-crosslinked molecule are independently cleavable, and the method further includes adding a cutting agent to the system to cut the elements L1, L2 or both; or wherein the element A of the photo-crosslinked molecule is derived from psoralen, and the method further includes applying ultraviolet light with a wavelength of about 230 nm to cut the element A; thereby generating a fingerprint of cross-linked proteins adjacent to nucleic acids in the system.

[0206] Embodiment 19. A method for preparing a photocrosslinked molecule according to any one of embodiments 9 to 13, the method comprising: providing an azide derivative of a nucleic acid-binding photoreactive reagent, the reagent comprising psoralen, methyl trimethoxalin, benzophenone, 4',6-diamidino-2-phenylindole (DAPI), Hoechst dye, polyamide or G quadruplex binding molecule or a derivative thereof; providing an azide derivative of a photoreactive reagent comprising a diazirine moiety to obtain an azide-diazirine bifunctional photoreactive reagent, and the photoreactive reagent optionally further comprises an acetylenic group, or providing an aryl azide. , the aryl azide is optionally selected from phenyl azide, o-hydroxyphenyl azide, m-hydroxyphenyl azide, tetrafluorophenyl azide, o-nitrophenyl azide, m-nitrophenyl azide or azidomethylcoumarin; optionally providing an azide derivative of a detectable agent, the detectable agent comprising a fluorophore, biotin, a chromophore, a chromogen, a quantum dot, a fluorescent microsphere or a nanoparticle; providing a multi-arm reagent having at least three functional groups, each functional group independently comprising an alkyne; and combining each azide derivative and the aryl azide (if provided) with the multi-arm reagent in a reaction vessel to prepare a photocrosslinked molecule.

[0207] Embodiment 20. A method as described in embodiment 19, wherein the multi-arm reagent has at least three functional groups, each functional group independently comprising a cyclooctyne group.

[0208] Embodiment 21. A method as described in embodiment 19 or 20, wherein the nucleic acid-binding photoreactive reagent comprises a first primary amine functional group, and providing an azide derivative of the nucleic acid-binding photoreactive reagent includes converting the first primary amine functional group into a first azide-containing portion, optionally by reacting the nucleic acid-binding photoreactive reagent with imidazole-1-sulfonyl azide; and / or wherein the photoreactive reagent comprising a diazirine portion further comprises a second primary amine functional group or is modified with a second primary amino functional group, and providing an azide derivative of the photoreactive reagent includes converting the second primary amine functional group into a second azide-containing portion, optionally by reacting the photoreactive reagent with imidazole-1-sulfonyl azide.

[0209] Additional embodiments include those listed below.

[0210] In some embodiments, L1 comprises a length of 2 to 20 carbons or 20-100 carbons. In some embodiments, L1 comprises 2 to 20 carbons, 2 to 19 carbons, 2 to 18 carbons, 2 to 17 carbons, 2 to 16 carbons, 2 to 15 carbons, 2 to 14 carbons, 2 to 13 carbons, 2 to 12 carbons, 2 to 11 carbons, 2 to 10 carbons, 2 to 9 carbons, 2 to 8 carbons, 2 to 7 carbons, 2 to 6 carbons, 2 to 5 carbons, 2 to 4 carbons, 2 to 3 carbons.

[0211] In some embodiments, L1 comprises 20-100 carbons in length, 20-95 carbons in length, 20-90 carbons in length, 20-85 carbons in length, 20-80 carbons in length, 20-75 carbons in length, 20-70 carbons in length, 20-65 carbons in length, 20-60 carbons in length, 20-55 carbons in length, 20-50 carbons in length, 20-45 carbons in length, 20-40 carbons in length, 20-35 carbons in length, 20-30 carbons in length, or 20-25 carbons in length.

[0212] In some embodiments, L2 comprises a length of 2 to 20 carbons or a length of 20-100 carbons. In some embodiments, L2 comprises 2 to 20 carbons, 2 to 19 carbons, 2 to 18 carbons, 2 to 17 carbons, 2 to 16 carbons, 2 to 15 carbons, 2 to 14 carbons, 2 to 13 carbons, 2 to 12 carbons, 2 to 11 carbons, 2 to 10 carbons, 2 to 9 carbons, 2 to 8 carbons, 2 to 7 carbons, 2 to 6 carbons, 2 to 5 carbons, 2 to 4 carbons, 2 to 3 carbons.

[0213] In some embodiments, L2 comprises 20-100 carbons in length, 20-95 carbons in length, 20-90 carbons in length, 20-85 carbons in length, 20-80 carbons in length, 20-75 carbons in length, 20-70 carbons in length, 20-65 carbons in length, 20-60 carbons in length, 20-55 carbons in length, 20-50 carbons in length, 20-45 carbons in length, 20-40 carbons in length, 20-35 carbons in length, 20-30 carbons in length, or 20-25 carbons in length.

[0214] In some embodiments, the detectable functional group comprises a fluorophore, biotin, a chromophore, a chromogen, a quantum dot, a fluorescent microsphere, or a nanoparticle. In some embodiments, the detectable functional group comprises a fluorophore. In some embodiments, the detectable functional group comprises biotin. In some embodiments, the detectable functional group comprises a chromophore. In some embodiments, the detectable functional group comprises a chromogen. In some embodiments, the detectable functional group comprises a quantum dot. In some embodiments, the detectable functional group comprises a fluorescent microsphere. In some embodiments, the detectable functional group comprises a nanoparticle.

[0215] In some embodiments, the detectable functional group is a fluorophore, biotin, a chromophore, a chromogen, a quantum dot, a fluorescent microsphere, or a nanoparticle. In some embodiments, the detectable functional group is a fluorophore. In some embodiments, the detectable functional group is biotin. In some embodiments, the detectable functional group is a chromophore. In some embodiments, the detectable functional group is a chromogen. In some embodiments, the detectable functional group is a quantum dot. In some embodiments, the detectable functional group is a fluorescent microsphere. In some embodiments, the detectable functional group is a nanoparticle.

[0216] In some embodiments, the compounds of the present invention are photocrosslinked molecules. In some embodiments, the compounds of formula (I) are photocrosslinked molecules. In some embodiments, the compounds of formula (II) are photocrosslinked molecules. In some embodiments, the compounds of formula (III) are photocrosslinked molecules. In some embodiments, the compounds of formula (IIa) are photocrosslinked molecules. In some embodiments, the compounds of formula (IIb) are photocrosslinked molecules. In some embodiments, the compounds of formula (IIc) are photocrosslinked molecules. In some embodiments, the compounds of formula (IId) are photocrosslinked molecules. In some embodiments, the compounds of formula (IIe) are photocrosslinked molecules.

[0217] In some embodiments, the compound of the present invention is a bifunctional photocrosslinking (BFPX) probe. In some embodiments, the compound of formula (I) is a bifunctional photocrosslinking (BFPX) probe. In some embodiments, the compound of formula (II) is a bifunctional photocrosslinking (BFPX) probe. In some embodiments, the compound of formula (III) is a bifunctional photocrosslinking (BFPX) probe. In some embodiments, the compound of formula (IIa) is a bifunctional photocrosslinking (BFPX) probe. In some embodiments, the compound of formula (IIb) is a bifunctional photocrosslinking (BFPX) probe. In some embodiments, the compound of formula (IIc) is a bifunctional photocrosslinking (BFPX) probe. In some embodiments, the compound of formula (IId) is a bifunctional photocrosslinking (BFPX) probe. In some embodiments, the compound of formula (IIe) is a bifunctional photocrosslinking (BFPX) probe.

[0218] In some embodiments, the compound of formula (II) is a compound of formula (I). In some embodiments, the compound of formula (III) is a compound of formula (I). In some embodiments, the compound of formula (IIa) is a compound of formula (II). In some embodiments, the compound of formula (IIb) is a compound of formula (II). In some embodiments, the compound of formula (IIc) is a compound of formula (II). In some embodiments, the compound of formula (IId) is a compound of formula (II). In some embodiments, the compound of formula (IIe) is a compound of formula (II). In some embodiments, the compound of formula (IIa) is a compound of formula (I). In some embodiments, the compound of formula (IIb) is a compound of formula (I). In some embodiments, the compound of formula (IIc) is a compound of formula (I). In some embodiments, the compound of formula (IId) is a compound of formula (I). In some embodiments, the compound of formula (IIe) is a compound of formula (I).

[0219] In some embodiments, the compound of formula (I) is: A-L1-B. In some embodiments, the compound of formula (I) is: AB. In some embodiments, the compound of formula (II) is: A-L1-B. In some embodiments, the compound of formula (II) is: AB.

[0220] In some embodiments, the ultraviolet light comprises UVA light, UVB light, or UVC light, or a combination thereof. In some embodiments, the ultraviolet light comprises UVA and UVB light. In some embodiments, the ultraviolet light comprises UVB and UVC light. In some embodiments, the ultraviolet light comprises UVA and UVC light. In some embodiments, the ultraviolet light comprises only UVA light. In some embodiments, the ultraviolet light comprises only UVB light. In some embodiments, the ultraviolet light comprises only UVC light.

[0221] In some embodiments, the ultraviolet light is UVA light, UVB light, or UVC light, or a combination thereof. In some embodiments, the ultraviolet light is UVA and UVB light. In some embodiments, the ultraviolet light is UVB and UVC light. In some embodiments, the ultraviolet light is UVA and UVC light. In some embodiments, the ultraviolet light is only UVA light. In some embodiments, the ultraviolet light is only UVB light. In some embodiments, the ultraviolet light is only UVC light.

[0222] Without being bound by theory, in some embodiments, the wavelength of ultraviolet light is between 100 nm and 400 nm. Without being bound by theory, in some embodiments, the wavelength of UVA light is between 315 nm and 400 nm. Without being bound by theory, in some embodiments, the wavelength of UVB light is between 280 nm and 315 nm. Without being bound by theory, in some embodiments, the wavelength of UVC light is between 100 nm and 280 nm.

[0223] Without being bound by theory, in some embodiments, the wavelength of ultraviolet light is between 100 nm and 400 nm. Without being bound by theory, in some embodiments, the wavelength of UVA light is between 315 nm and 400 nm. Without being bound by theory, in some embodiments, the wavelength of UVB light is between 280 nm and 314 nm. Without being bound by theory, in some embodiments, the wavelength of UVC light is between 100 nm and 279 nm.

[0224] In some embodiments, the wavelength of the ultraviolet light is 300nm to 360nm. In some embodiments, the wavelength of the ultraviolet light is 300nm to 360nm. In some embodiments, the wavelength of the ultraviolet light is 300nm to 400nm, 300nm to 310nm, 300nm to 320nm, 300nm to 330nm, 300nm to 340nm, 300nm to 350nm, 300nm to 360nm, 300nm to 370nm, 300nm to 379nm, 300nm to 380nm, or 300nm to 390nm.

[0225] In some embodiments, the ultraviolet light has a wavelength of 400nm to 390nm, a wavelength of 400nm to 380nm, a wavelength of 400nm to 370nm, a wavelength of 400nm to 360nm, a wavelength of 400nm to 350nm, a wavelength of 400nm to 340nm, a wavelength of 400nm to 330nm, a wavelength of 400nm to 320nm, a wavelength of 400nm to 316nm, a wavelength of 400nm to 315nm, a wavelength of 400nm to 310nm, or a wavelength of 400nm to 300nm.

[0226] In some embodiments, the ultraviolet light has a wavelength of 315nm to 400nm, a wavelength of 315nm to 390nm, a wavelength of 315nm to 380nm, a wavelength of 315nm to 370nm, a wavelength of 315nm to 360nm, a wavelength of 315nm to 350nm, a wavelength of 315nm to 340nm, or a wavelength of 315nm to 330nm.

[0227] In some embodiments, the wavelength of ultraviolet light is 316nm to 400nm, 316nm to 390nm, 316nm to 380nm, 316nm to 379nm, 316nm to 370nm, 316nm to 360nm, 316nm to 350nm, 316nm to 340nm, or 316nm to 330nm. In some embodiments, the wavelength of ultraviolet light is 316nm to 379nm. In some embodiments, nucleic acid includes deoxyribonucleic acid (DNA) or ribonucleic acid (RNA) or a combination thereof. In some embodiments, nucleic acid includes deoxyribonucleic acid (DNA). In some embodiments, nucleic acid includes ribonucleic acid (RNA).

[0228] In some embodiments, the nucleic acid is deoxyribonucleic acid (DNA) or ribonucleic acid (RNA) or a combination thereof. In some embodiments, the nucleic acid is deoxyribonucleic acid (DNA). In some embodiments, the nucleic acid is ribonucleic acid (RNA).

[0229] Additional embodiments include those listed below.

[0230] Embodiment 1A. Compounds of formula (I): A-L1-(C) n -((L2) n -B) m Formula (I), in: A represents a nucleic acid binding functional group derived from psoralen, methyl trimethoxalan, benzophenone, 4',6-diamidino-2-phenylindole (DAPI), Hoechst dye, polyamide or G-quadruplex binding molecule, ethopropaldehyde or its derivatives; L1 does not exist or represents the first linker; When n=1, C represents a core moiety having at least two functional groups, each functional group being used to attach to L1 and to attach to at least one arm represented by L2-B, respectively; or when n=0, C is absent; When n=1, L2 independently represents the second linker of each arm represented by L2-B; or when n=0, L2 does not exist; For each of the arms, B independently represents: A photoreactive functional group, wherein the photoreactive functional group comprises diaziridine or a derivative thereof or an aryl azide or a derivative thereof, and optionally, the aryl azide or a derivative thereof is selected from phenyl azide, o-hydroxyphenyl azide, m-hydroxyphenyl azide, tetrafluorophenyl azide, o-nitrophenyl azide, m-nitrophenyl azide or azidomethylcoumarin; or Detectable functional groups; wherein, in at least one of the arms, B represents a photoreactive functional group; n = 0 or 1; m represents (L2) n -B represents the number of arms, wherein when n=1, m is an integer of 1 or more, or when n=0, m=1.

[0231] Embodiment 2A. The compound of embodiment 1A, wherein at least one of L1 and L2 is not absent, and at least one of L1 and L2 is cleavable.

[0232] Embodiment 3A. A compound as described in Embodiment 2A, wherein L1, L2 or both independently comprise one or more of a sulfoxide-containing mass spectrometry (MS) cleavable bond, an acid-cleavable CS bond, a disulfide group, and an azo group.

[0233] Embodiment 4A. A compound as described in any one of Embodiments 1A-3A, wherein n=0, m=1, and the compound is represented by formula (II): A-L1-B Formula (II), Wherein, L1 does not exist or is the first linker.

[0234] Embodiment 5A. A compound according to embodiment 4A, wherein: A is an amine-containing or amine-reactive derivative of psoralen, an amine-containing or amine-reactive derivative of methyltrimethsalen, an amine-containing or amine-reactive derivative of benzophenone, an amine-containing or amine-reactive derivative of 4',6-diamidino-2-phenylindole (DAPI), an amine-containing or amine-reactive derivative of Hoechst dye, an amine-containing or amine-reactive derivative of polyamide, or an amine-containing or amine-reactive derivative of a G-quadruplex binding molecule, or an amine-containing or amine-reactive derivative of ethoxybutyraldehyde, optionally A is derived from succinimidyl-[4-(psoralen-8-yloxy)]-butyrate (SPB) or 4'-aminomethyltrimethsalen (4AMT); B comprises diaziridine or diaziridine alkyne, optionally aminodiaziridine alkyne (AAD); and L1 is absent or is a first linker, wherein the first linker comprises one or more of the following: (i) a cleavable bond, (ii) an oligomer or polymer having -OCH2CH2- repeating units, and (iii) an unsaturated portion, the unsaturated portion being optionally selected from a carbon-carbon double bond, a carbon-carbon triple bond, or an aromatic group.

[0235] Embodiment 6A. A compound as described in Embodiment 5A, wherein L1-B is derived from succinimidyl 6-(4,4'-azidopentanamido)hexanoate (NHS-LC-SDA), succinimidyl 2-((4,4'-azidopentanamido)ethyl)-1,3'dithiopropionate (NHS-SS-diaziridine) or 2-(3-(but-3-yn-1-yl)-3H-diaziridine-3-yl)ethane-1-amine (AAD); and / or wherein A is derived from 4'-aminomethyltrimethsalen (4AMT) or succinimidyl-[4-(psoralen-8-yloxy)]-butyrate (SPB); and wherein optionally, the photocrosslinking molecule is represented by Formula (IIa) or Formula (IIc):

[0236] Embodiment 7A. A compound according to Embodiment 5A, wherein: A is derived from succinimidyl-[4-(psoralen-8-yloxy)]-butyrate (SPB) or 4'-aminomethyltrimethylsalen (4AMT); B comprises diaziridine or diaziridine alkyne, optionally aminodiaziridine alkyne (AAD); and L1 is the first linker, the first linker comprising one or more of the following: (i) a cleavable bond, (ii) an oligomer or polymer having a -OCH2CH2- repeating unit, and / or (iii) an unsaturated portion, the unsaturated portion being optionally selected from a carbon-carbon double bond or an aromatic group; and wherein optionally, the photo-crosslinking molecule is represented by formula (IIb), formula (IId), formula (IIe) or formula (IIf):

[0237] Embodiment 8A. A compound according to embodiment 4A, wherein: A is selected from the group consisting of: in: R 1 are independently H, halo, OH, optionally substituted alkoxy or optionally substituted alkyl; R 2 are independently H, halo, OH, optionally substituted alkoxy or optionally substituted alkyl; a is 0, 1, 2, 3, 4 or 5; and b is 0, 1, 2, 3, or 4; in: R 3 are independently H, halo, OH, optionally substituted alkoxy or optionally substituted alkyl; R 4 is H, halo, OH, optionally substituted alkoxy or optionally substituted alkyl; R 5 are independently H, halo, OH, optionally substituted alkoxy or optionally substituted alkyl; c is 0, 1, 2, 3 or 4; and d is 0, 1, 2, 3 or 4; in: R 6 is H, halo, OH, optionally substituted alkoxy or optionally substituted alkyl; R 7 is H, halo, OH, optionally substituted alkoxy or optionally substituted alkyl; R 8 is H, halo, OH, optionally substituted alkoxy or optionally substituted alkyl; and R 9 is H, halo, OH, optionally substituted alkoxy or optionally substituted alkyl; in: R 10 is H, halo, OH, optionally substituted alkoxy or optionally substituted alkyl; R 11 is H, halo, OH, optionally substituted alkoxy or optionally substituted alkyl; and R 12 is H, halo, OH, optionally substituted alkoxy or optionally substituted alkyl; and L1 is absent or L is selected from the group consisting of: in: q is 0, 1, 2, 3, or 4; in: p is 0, 1, 2, 3 or 4; in: R 13is independently H, halo, OH, optionally substituted alkoxy or optionally substituted alkyl; and e is 0, 1, 2, 3 or 4; in: r is 0, 1, 2, 3, or 4; in: s is 0, 1, 2, 3, or 4; in: t is 0, 1, 2, 3, or 4; u is 0, 1, 2, 3, or 4; and B is selected from the group consisting of:

[0238] Embodiment 9A. The compound of Embodiment 4A, wherein: A is selected from the group consisting of: L1 is absent or L1 is selected from the group consisting of: and B is selected from the group consisting of:

[0239] Embodiment 10A. The compound of Embodiment 1A or Embodiment 4A, wherein the compound is:

[0240] Embodiment 11A. A compound as described in Embodiment 5A, wherein L1 comprises a length of 2 to 20 carbons or 20-100 carbons.

[0241] Embodiment 12A. A compound as described in any of Embodiments 1A-3A, wherein n=1, m is an integer of 2 or greater, and C represents a core portion having at least three functional groups, each functional group being used to attach to L1 and to attach to at least two arms each represented by (L2-B), such that the compound is represented by Formula (III):

[0242] Embodiment 13A. A compound as described in embodiment 12A, wherein in one of the at least two arms B comprises diaziridine or diaziridine azide, and in the other of the at least two arms B represents a detectable functional group, wherein the detectable functional group comprises a fluorophore, biotin, a chromophore, a chromogen, a quantum dot, a fluorescent microsphere or a nanoparticle.

[0243] Embodiment 14A. A compound as described in embodiment 12A or embodiment 13A, wherein L1, L2 or both independently comprise one or more of the following: (i) a cleavable bond, (ii) an oligomer or polymer having repeating units of -OCH2CH2-, and (iii) an unsaturated portion.

[0244] Embodiment 15A. A compound as described in any of Embodiments 9A-14A, wherein C represents a dendritic core portion, wherein the core portion comprises at least three surface functional groups, each surface functional group is used to attach to L1 and to attach to the at least two arms each represented by L2-B.

[0245] Embodiment 16A. A compound as described in any of Embodiments 9A-15A, wherein L1, L2 or both independently comprise a triazole bonded to A.

[0246] Embodiment 17A. A method of crosslinking a nucleic acid to an adjacent protein in a system, the method comprising: incubating a compound of any one of Embodiments 1A-16A with the system, and irradiating the system with ultraviolet light.

[0247] Embodiment 18A. The method of embodiment 17A, wherein the system is a living cell.

[0248] Embodiment 19A. The method of Embodiment 17A or Embodiment 18A, wherein the wavelength of the ultraviolet light is between 300 nm and 370 nm.

[0249] Embodiment 20A. The method of any one of embodiments 17A-19A, further comprising using the system to perform one or more of immunoprecipitation, chromatin precipitation, 3D chromatin conformation capture, mass spectrometry, and electrophoresis.

[0250] Embodiment 21A. A method as described in any of embodiments 17A-20A, wherein element L1, L2 or both of the compound are independently cleavable, and the method further includes adding a cleavage agent to the system to cleave the elements L1, L2 or both; or wherein element A of the compound is derived from psoralen, and the method further includes applying ultraviolet light with a wavelength of about 230 nm to cleave the element A; thereby generating a fingerprint of cross-linked proteins adjacent to nucleic acids in the system.

[0251] Embodiment 22A. A method of preparing a compound of any one of Embodiments 12A-16A, the method comprising: Providing an azide derivative of a nucleic acid-binding photoreactive reagent, the reagent comprising psoralen, methyltrimethoxalen, benzophenone, 4',6-diamidino-2-phenylindole (DAPI), Hoechst dye, polyamide or G-quadruplex binding molecule, ethoprodol or a derivative thereof; Providing an azide derivative of a photoreactive reagent comprising a diaziridine moiety to obtain an azide-diaziridine bifunctional photoreactive reagent, and the photoreactive reagent optionally further comprises an alkyne group, or providing an aryl azide, the aryl azide optionally selected from phenyl azide, o-hydroxyphenyl azide, m-hydroxyphenyl azide, tetrafluorophenyl azide, o-nitrophenyl azide, m-nitrophenyl azide or azidomethylcoumarin; optionally providing an azide derivative of a detectable agent, the detectable agent comprising a fluorophore, biotin, a chromophore, a chromogen, a quantum dot, a fluorescent microsphere or a nanoparticle; providing a multi-arm reagent having at least three functional groups, each functional group independently comprising an alkyne; and Each azide derivative and aryl azide (if provided) are combined in one reaction vessel along with the multi-arm reagent to prepare the compound.

[0252] Embodiment 23A. A method as described in Embodiment 22A, wherein the multi-arm reagent has at least three functional groups, each functional group independently comprising a cyclooctyne group.

[0253] Embodiment 24A. A method as described in embodiment 22A or embodiment 23A, wherein the nucleic acid-binding photoreactive reagent comprises a first primary amine functional group, and providing an azide derivative of the nucleic acid-binding photoreactive reagent includes converting the first primary amine functional group into a first azide-containing portion, optionally by reacting the nucleic acid-binding photoreactive reagent with imidazole-1-sulfonyl azide; and / or wherein the photoreactive reagent comprising a diazirine portion further comprises a second primary amine functional group or is modified with a second primary amino functional group, and providing an azide derivative of the photoreactive reagent includes converting the second primary amine functional group into a second azide-containing portion, optionally by reacting the photoreactive reagent with imidazole-1-sulfonyl azide.

[0254] Embodiment 25A. A method for cross-linking nucleic acids to proteins in a system, the method comprising: providing a compound of any one of Embodiments 1A-16A; Providing a system, wherein the system comprises a nucleic acid and a protein; contacting the compound with the system; and The system and the compound are irradiated with ultraviolet light under conditions effective to cross-link the nucleic acid to the protein.

[0255] Embodiment 26A. The method of embodiment 25A, wherein the system is a living cell.

[0256] Embodiment 27A. The method of Embodiment 25A or Embodiment 26A, wherein the wavelength of the ultraviolet light is between 300 nm and 370 nm.

[0257] Embodiment 28A. The method of any one of embodiments 25A-27A, further comprising using the system to perform one or more of immunoprecipitation, chromatin precipitation, 3D chromatin conformation capture, mass spectrometry, and electrophoresis.

[0258] Embodiment 29A. A method as described in any of embodiments 25A-28A, wherein element L1, L2 or both of the compound are independently cleavable, and the method further includes adding a cleavage agent to the system to cleave the elements L1, L2 or both; or wherein element A of the compound is derived from psoralen, and the method further includes applying ultraviolet light with a wavelength of about 230 nm to cleave the element A; thereby generating a fingerprint of cross-linked proteins adjacent to nucleic acids in the system.

[0259] Additional embodiments include those listed below.

[0260] In some embodiments, the present invention provides a compound:

[0261] In some embodiments, the present invention provides a compound:

[0262] In some embodiments, the present invention provides a compound:

[0263] In some embodiments, the present invention provides a compound:

[0264] In some embodiments, the present invention provides a compound:

[0265] In some embodiments, the present invention provides a compound:

[0266] In some embodiments, the present invention provides a compound:

[0267] In some embodiments, the present invention provides a compound:

[0268] Additional embodiments include those listed below.

[0269] In some embodiments, the compound of the present invention is a compound of formula (I), a compound of formula (II), a compound of formula (III), a compound of formula (IIa), a compound of formula (IIb), a compound of formula (IIc), a compound of formula (IId), or a compound of formula (IIe), or any combination thereof. In some embodiments, the compound of the present invention is a compound of formula (II).

[0270] In various embodiments, the present invention provides a compound of formula (II): A-L1-B. In some embodiments, L1 is absent. In some embodiments, L1 is present. In some embodiments, L1 is absent or is the first linker.

[0271] Additional embodiments include those listed below.

[0272] In some embodiments, the compound of the present invention is selected from the group consisting of:

[0273] Additional embodiments include those listed below.

[0274] In various embodiments, the present invention provides a compound of formula (II): A-L1-B. In some embodiments, L1 is absent. In some embodiments, L1 is present. In some embodiments, L1 is absent or is the first linker. In some embodiments, the compound of formula (II) is selected from the group consisting of:

[0275] Other embodiments include those listed below.

[0276] In some embodiments, the compound of formula (I) is selected from the group consisting of:

[0277] Additional embodiments include those listed below.

[0278] In some embodiments, A is selected from the group consisting of: in: R 1 are independently H, halo, OH, optionally substituted alkoxy or optionally substituted alkyl; R 2 are independently H, halo, OH, optionally substituted alkoxy or optionally substituted alkyl; a is 0, 1, 2, 3, 4 or 5; and b is 0, 1, 2, 3, or 4; in: R 3 are independently H, halo, OH, optionally substituted alkoxy or optionally substituted alkyl; R 4 is H, halo, OH, optionally substituted alkoxy or optionally substituted alkyl; R 5 are independently H, halo, OH, optionally substituted alkoxy or optionally substituted alkyl; c is 0, 1, 2, 3 or 4; and d is 0, 1, 2, 3 or 4; in: R 6 is H, halo, OH, optionally substituted alkoxy or optionally substituted alkyl; R 7 is H, halo, OH, optionally substituted alkoxy or optionally substituted alkyl; R 8 is H, halo, OH, optionally substituted alkoxy or optionally substituted alkyl; and R 9 is H, halo, OH, optionally substituted alkoxy or optionally substituted alkyl; in: R 10 is H, halo, OH, optionally substituted alkoxy or optionally substituted alkyl; R 11 is H, halo, OH, optionally substituted alkoxy or optionally substituted alkyl; and R 12 is H, halo, OH, optionally substituted alkoxy or optionally substituted alkyl;

[0279] Additional embodiments include those listed below.

[0280] In some embodiments, A is: in: R 1 are independently H, halo, OH, optionally substituted alkoxy or optionally substituted alkyl; R 2 are independently H, halo, OH, optionally substituted alkoxy or optionally substituted alkyl; a is 0, 1, 2, 3, 4 or 5; and b is 0, 1, 2, 3 or 4.

[0281] In some embodiments, A is: in: R 3 are independently H, halo, OH, optionally substituted alkoxy or optionally substituted alkyl; R 4 is H, halo, OH, optionally substituted alkoxy or optionally substituted alkyl; R 5 are independently H, halo, OH, optionally substituted alkoxy or optionally substituted alkyl; c is 0, 1, 2, 3 or 4; and d is 0, 1, 2, 3 or 4.

[0282] In some embodiments, A is: in: R 6 is H, halo, OH, optionally substituted alkoxy or optionally substituted alkyl; R 7 is H, halo, OH, optionally substituted alkoxy or optionally substituted alkyl; R 8 is H, halo, OH, optionally substituted alkoxy or optionally substituted alkyl; and R 9 is H, halo, OH, optionally substituted alkoxy or optionally substituted alkyl.

[0283] In some embodiments, A is: in: R 10 is H, halo, OH, optionally substituted alkoxy or optionally substituted alkyl; R 11 is H, halo, OH, optionally substituted alkoxy or optionally substituted alkyl; and R 12 is H, halo, OH, optionally substituted alkoxy or optionally substituted alkyl.

[0284] In some embodiments, A is:

[0285] In some embodiments, A is:

[0286] Additional embodiments include those listed below.

[0287] In some embodiments, A is selected from the group consisting of: in: R 1 are independently H, halo, OH, OCH3 or CH3; R 2 are independently H, halo, OH, OCH3 or CH3; a is 0, 1, 2, 3, 4 or 5; and b is 0, 1, 2, 3, or 4; in: R 3 are independently H, halo, OH, OCH3 or CH3; R 4 is H, halo, OH, OCH3 or CH3; R 5 are independently H, halo, OH, OCH3 or CH3; c is 0, 1, 2, 3 or 4; and d is 0, 1, 2, 3 or 4; in: R 6 is H, halo, OH, OCH3 or CH3; R 7 is H, halo, OH, OCH3 or CH3; R 8 is H, halo, OH, OCH3 or CH3; and R 9 is H, halo, OH, OCH3 or CH3; in: R 10 is H, halo, OH, OCH3 or CH3; R 11 is H, halo, OH, OCH3 or CH3; and R 12 is H, halo, OH, OCH3 or CH3; as well as

[0288] Additional embodiments include those listed below.

[0289] In some embodiments, A is: in: R 1 are independently H, halo, OH, OCH3 or CH3; R 2 are independently H, halo, OH, OCH3 or CH3; a is 0, 1, 2, 3, 4 or 5; and b is 0, 1, 2, 3 or 4.

[0290] In some embodiments, A is: in: R 3 are independently H, halo, OH, OCH3 or CH3; R 4 is H, halo, OH, OCH3 or CH3; R 5 are independently H, halo, OH, OCH3 or CH3; c is 0, 1, 2, 3 or 4; and d is 0, 1, 2, 3 or 4.

[0291] In some embodiments, A is: in: R 6 is H, halo, OH, OCH3 or CH3; R 7 is H, halo, OH, OCH3 or CH3; R 8 is H, halo, OH, OCH3 or CH3; and R 9 It is H, halo, OH, OCH3 or CH3.

[0292] In some embodiments, A is: in: R 10 is H, halo, OH, OCH3 or CH3; R 11 is H, halo, OH, OCH3 or CH3; and R 12 It is H, halo, OH, OCH3 or CH3.

[0293] In some embodiments, A is:

[0294] In some embodiments, A is:

[0295] Additional embodiments include those listed below.

[0296] In some embodiments, A is selected from the group consisting of:

[0297] In some embodiments, A is selected from the group consisting of:

[0298] In some embodiments, A is:

[0299] In some embodiments, A is:

[0300] In some embodiments, A is:

[0301] In some embodiments, A is:

[0302] In some embodiments, A is:

[0303] In some embodiments, A is:

[0304] Additional embodiments include those listed below.

[0305] In some embodiments, L1 is selected from the group consisting of: in: q is 0, 1, 2, 3, or 4; in: p is 0, 1, 2, 3 or 4; in: R 13 is independently H, halo, OH, optionally substituted alkoxy, or optionally substituted alkyl; and e is 0, 1, 2, 3, or 4; in: r is 0, 1, 2, 3, or 4; in: s is 0, 1, 2, 3, or 4; in: t is 0, 1, 2, 3, or 4; in: u is 0, 1, 2, 3, or 4;

[0306] In some embodiments, L1 is: in: q is 0, 1, 2, 3 or 4.

[0307] In some embodiments, L1 is: in: p is 0, 1, 2, 3 or 4.

[0308] In some embodiments, L1 is: in: R 13 is independently H, halo, OH, optionally substituted alkoxy, or optionally substituted alkyl; and e is 0, 1, 2, 3, or 4.

[0309] In some embodiments, L1 is: in: r is 0, 1, 2, 3 or 4.

[0310] In some embodiments, L1 is: in: s is 0, 1, 2, 3, or 4.

[0311] In some embodiments, L1 is: in: t is 0, 1, 2, 3, or 4.

[0312] In some embodiments, L1 is: in: u is 0, 1, 2, 3, or 4

[0313] In some embodiments, L1 is:

[0314] In some embodiments, L1 is:

[0315] In some embodiments, L1 is selected from the group consisting of:

[0316] In some embodiments, L1 is:

[0317] In some embodiments, L1 is:

[0318] In some embodiments, L1 is:

[0319] In some embodiments, L1 is:

[0320] In some embodiments, L1 is:

[0321] In some embodiments, L1 is:

[0322] In some embodiments, L1 is:

[0323] Additional embodiments include those listed below.

[0324] In some embodiments, B is selected from the group consisting of:

[0325] In some embodiments, B is:

[0326] In some embodiments, B is:

[0327] In some embodiments, B is:

[0328] Additional embodiments include those listed below.

[0329] In some embodiments, the psoralen is:

[0330] In some embodiments, the benzophenone is:

[0331] In some embodiments, 4',6-diamidino-2-phenylindole (DAPI) is:

[0332] In some embodiments, non-limiting examples of Hoechst dyes and their salts include:

[0333] In some embodiments, non-limiting examples of G-quadruplex binding molecule amine derivatives include:

[0334] In some embodiments, 4'-aminomethyltrimethylsaren (4AMT) is:

[0335] In some embodiments, methyl trimesaren is:

[0336] Additional embodiments include those listed below.

[0337] In some embodiments, 4AMT-LC-SDA is:

[0338] In some embodiments, 4AMT adipate AAD is:

[0339] In some embodiments, SPB-Spermidine-AD is:

[0340] In some embodiments, SPB-PEG4-ADD is:

[0341] In some embodiments, 4AMT-SDAD is:

[0342] In some embodiments, 4AMT terephthalate AAD is:

[0343] In some embodiments, SPB-PEG3-AAD is:

[0344] In some embodiments, the SPB-AAD is:

[0345] Additional embodiments include those listed below.

[0346] In various embodiments, the present invention provides a method for cross-linking nucleic acids and proteins in a system, the method comprising: incubating a compound of the present invention with the system, and irradiating the system with ultraviolet light. In some embodiments, the system comprises nucleic acids and proteins. In some embodiments, the compound of the present invention is a compound of formula (I), a compound of formula (II), a compound of formula (III), a compound of formula (IIa), a compound of formula (IIb), a compound of formula (IIc), a compound of formula (IId), or a compound of formula (IIe), or any combination thereof. In some embodiments, the system comprises nucleic acids and proteins. In some embodiments, the nucleic acid is DNA, RNA, or a combination thereof. In some embodiments, the compound of the present invention is a compound of formula (II). In some embodiments, the method is performed in vitro, in vivo, or a combination thereof. In some embodiments, the method is performed in vitro or in vivo. In some embodiments, the method is performed in vitro. In some embodiments, the method is performed in vivo. In some embodiments, the nucleic acids and proteins are adjacent to each other in the system. In some embodiments, the system is a biological system. In some embodiments, the system is a living biological cell. In some embodiments, the living cell is a living biological cell. In some embodiments, the system is a living mammalian cell.

[0347] In various embodiments, the present invention provides a method for crosslinking nucleic acids and proteins in a system, the method comprising: providing a compound of the present invention; providing a system, wherein the system comprises nucleic acids and proteins; contacting the compound with nucleic acids and proteins in the system; and irradiating the system with ultraviolet light under conditions effective to crosslink nucleic acids and proteins. In some embodiments, the system is a living cell. In some embodiments, the system is an in vivo system. In some embodiments, the system is an in vitro system. In some embodiments, the system is an in vivo system or an in vitro system. In some embodiments, the compound of the present invention is a compound of formula (I), a compound of formula (II), a compound of formula (III), a compound of formula (IIa), a compound of formula (Iib), a compound of formula (Iic), a compound of formula (Iid), or a compound of formula (Iie), or any combination thereof. In some embodiments, the system comprises nucleic acids and proteins. In some embodiments, the nucleic acids are DNA, RNA, or a combination thereof. In some embodiments, the system is a sample. In some embodiments, the system is a biological sample. In some embodiments, the compound of the present invention is a compound of formula (II). In some embodiments, the method is performed in vitro, in vivo, or a combination thereof. In some embodiments, the method is performed in vitro or in vivo. In some embodiments, the method is performed in vitro. In some embodiments, the method is performed in vivo. In some embodiments, the nucleic acid and protein are adjacent to each other in the system. In some embodiments, the system is a living biological cell. In some embodiments, the living cell is a living biological cell. In some embodiments, the system is a living mammalian cell.

[0348] In various embodiments, the present invention provides a method for crosslinking nucleic acids and proteins in a sample, the method comprising: providing a compound of the present invention; providing a sample, wherein the sample comprises nucleic acids and proteins; contacting the compound with nucleic acids and proteins in the sample; and irradiating the sample with ultraviolet light under conditions effective to crosslink nucleic acids and proteins. In some embodiments, the sample is a living cell. In some embodiments, the sample is an in vivo sample. In some embodiments, the sample is an in vitro system. In some embodiments, the sample is an in vivo sample or an in vitro sample. In some embodiments, the sample is a biological sample. In some embodiments, the compound of the present invention is a compound of formula (I), a compound of formula (II), a compound of formula (III), a compound of formula (Iia), a compound of formula (Iib), a compound of formula (Iic), a compound of formula (Iid), or a compound of formula (Iie), or any combination thereof. In some embodiments, the sample comprises nucleic acids and proteins. In some embodiments, the nucleic acids are DNA, RNA, or a combination thereof. In some embodiments, the compound of the present invention is a compound of formula (II). In some embodiments, the method is performed in vitro, in vivo, or a combination thereof. In some embodiments, the method is performed in vitro or in vivo. In some embodiments, the method is performed in vitro. In some embodiments, the method is performed in vivo. In some embodiments, the nucleic acid and protein are adjacent to each other in the sample. In some embodiments, the sample is a living biological cell. In some embodiments, the living cell is a living biological cell. In some embodiments, the sample is a living mammalian cell. Example

[0349] The following examples are provided to better illustrate the claimed invention and should not be construed as limiting the scope of the invention. With regard to the specific materials mentioned, they are for illustrative purposes only and are not intended to limit the present invention. Those skilled in the art may develop equivalent means or reactants without exerting creativity and without departing from the scope of the present invention.

[0350] Example 1.

[0351] The disclosed probes have a wide range of purposes and applications, including the following examples: i) as a more efficient and effective cross-linking agent alternative to formaldehyde or similar aldehyde reagents for in vitro or in vivo cross-linking to explore the situation of nucleic acids to neighboring proteins, such as in immunoprecipitation (IP), chromatin immunoprecipitation (ChIP), 3D chromatin conformation capture technology (HiC / HiChIP); ii) as a cross-linking agent for exploring protein binding on nucleic acids in proteomic analysis (such as protein or peptide identification by mass spectrometry or electrophoresis), because the psoralen head from the probe can be released from the nucleic acid by irradiation with a different wavelength (about 230nm), while all other different nucleic acid binding heads can be reversed by the aforementioned cleavable azo linkage ( Figure 3A-3C ). The remaining part still attached to the protein side after cleavage can be used as a unique mass ID / fingerprint for spectrum comparison. Isotopic compounds with the same design can also be used as IDs.

[0352] Compared to traditional protein-nucleic acid crosslinkers such as formaldehyde or longer crosslinkers such as DSG (disuccinimidyl glutarate) or DSSP (3,3'-dithiobis(sulfonated succinimidyl propionate)), the probe disclosed herein has the following advantages: i) It provides direct, length-controlled (via different lengths of NHS or alkyne linker cores), flexible and water-soluble (using PEG core) crosslinking results between nucleic acids and proteins. In contrast, formaldehyde derivatives rely on the indirect binding of the two through proteins that bind tightly to nucleic acids, rather than directly binding to the nucleic acids themselves. ii) This direct binding mode will provide more precise protein binding locations on nucleic acids, whether genome-wide or locally labeled by fluorophores. In contrast, traditional formaldehyde-based methods rely on the distribution, density and location of auxiliary proteins (such as histones) that bind tightly to nucleic acids, which leads to imprecise evaluation of chromatin protein binding results. iii) Multi-arm PEG cores can endow probes with a variety of other desired functions and can be assembled in a relatively simple combinational manner, such as the aforementioned enrichment handles, reversibility by cleavage, length control, fluorescence, etc. iv) Compared to formaldehyde which can crosslink all other proteins at all positions and thereby distort access reality and denature antibody recognition epitopes, our designed probes (as demonstrated by direct DNA crosslinking and length control) do not cause any of the above. Due to the nucleic acid specific recognition end of the probe, protein-protein crosslinking cannot occur in large quantities, so epitopes from the rest of the protein are better retained in their native form. v) The photochemical groups selected for the new probes described herein have similar long wavelength UV spectra (about 300nm to 360nm), and our probes are less energetically damaging to crosslink biomolecular targets compared to shorter UV direct crosslinking without any probe. And the design of the probe included makes it more efficient in crosslinking between nucleic acids and proteins.

[0353] In vitro assays have demonstrated the success and efficiency of the demonstration probe (AMT-LC-SDA) in cross-linking various amounts of double-stranded DNA to the DNA-binding protein human myocyte enhancer factor 2A (MEF2A) (1-95aa) ( Figure 5C-5E ). Reliable cross-linking of the probes by UV light cross-linking conditions was confirmed by denaturation or digestion methods, such as 8.3M urea gels, SDS gel electrophoresis, combined with samples digested with proteinase K overnight at 65 degrees Celsius and denatured with SDS loading dye at high temperature at 95 degrees Celsius. Compared with all control conditions, cross-linking of proteins and nucleic acids only occurred when both probes and UV were present. Under these harsh evaluation conditions, proteins were thoroughly digested on the cross-linked complexes, but the remaining cross-linked amino acid residues still made the DNA larger in size and significantly migrated compared to other controls, as shown by the bold arrows. Agarose gels were also used to estimate cross-linking efficiency (UV 2hr) due to their larger pore size, and our probes have achieved an efficiency of approximately 30%-60% as estimated by the amount of free DNA remaining ( Figure 5E ), which is much higher than the UV-induced protein-DNA crosslinking (10-20%) without the use of a probe reported in Nature Communications, volume 11, Article number: 3019 (2020). In addition, the probe exhibited remarkable selectivity for nucleic acid-binding proteins: even in the presence of high concentrations of random non-nucleic acid-binding proteins (such as BSA), only the MEF2A sample showed DNA substrate depletion by crosslinking ( Figure 5C-5E ). This further supports the idea of ​​using this probe as a replacement for formaldehyde, which crosslinks proteins only in the vicinity of nucleic acids, thereby greatly reducing random protein-protein crosslinking noise and helping to preserve the native state of protein epitopes.

[0354] Example 2.

[0355] Compared with existing cross-linking technologies, the probes to be developed in the proposed study will have the following features: (1) By using DNA-binding molecules, we can achieve regioselectivity of the photochemical reaction at or near the DNA binding site, thereby reducing the background noise of nonspecific cross-linking. Unlike formaldehyde, which reacts with any protein in the cell and causes damage to the DNA binding surface of TF and antibody epitopes, our probes (e.g., psoralen 4AMT head) will preferentially bind only to nucleic acids ( Figure 2A , Figure 2B). (2) Through synthetic modification, we can introduce photoaffinity labeling groups that are highly effective but can be activated by long wavelength UV (350nm), thereby reducing the photodamage associated with the use of large doses of short UV (250nm) irradiation. Synthetically introduced photoaffinity labels (such as diazirine) will be more efficient photocrosslinking groups than endogenous protein residues and nucleoside bases. (3) The versatility of synthetically introduced functional heads for customized applications ( Figure 2C For example, multiple arms of protein photoaffinity labeling groups can be introduced to capture multiple TFs bound to complex DNA elements, thereby providing direct experimental evidence for the transcriptional synergy of multi-TF complexes (e.g. Figure 2C The linker length can be varied to capture either the proximal DNA binding domain or the more distal protein cofactor recruited by the TF. Furthermore, the linker can be made cleavable so that the captured peptide can be released (after protease digestion) for mass spectrometry analysis, allowing not only identification of DNA sites (as done in ChIP assays) but also structural mapping of the DNA binding surface on the protein (as in the XL-MS studies mentioned above).

[0356] Our initial attempt was to find a way to nonspecifically functionalize DNA with primary amines. To this end, we chose the cell-permeable, non-toxic 4'-aminomethyltrimethylsarin (4AMT), a psoralen derivative, which can bind and intercalate DNA almost nonspecifically (binding site preference 5'-TA>5'-AT>>5'-TG>5'-GT'), and covalently crosslink with DNA with high efficiency (up to 80%) under long-wavelength (~360nm) UV irradiation, thereby introducing a highly efficient nucleophile into DNA. Subsequently, any bifunctional amine-reactive group (such as DSS) can be used to crosslink DNA to nearby bound proteins. Although this design solves the problem of low DNA reactivity, amine-reactive crosslinkers may still nonspecifically modify other cellular proteins, thereby introducing high background noise and affecting antibody recognition of TF, and may also disrupt the DNA binding surface (see Figure 5G and described below). Therefore, we sought to use DNA-binding molecules to restrict the cross-linking reaction only to proteins that bind or are recruited to DNA. These considerations led to the design of bifunctional photocross-linking probes.

[0357] The present invention involves the design and custom synthesis of a new class of molecular tools that possess at least two photocrosslinking functionalities, hence termed bifunctional photocrosslinkers (BFPX) ( Figure 2C). One of the functional groups is responsible for binding and cross-linking to DNA upon UV irradiation, while the other functional group is responsible for capturing proteins bound or recruited to DNA by UV-activated carbenes or nitrenes. The two functional groups are connected by a linker that is engineered to have features that facilitate monitoring, separation, and analysis of cross-linked protein-DNA complexes.

[0358] For DNA binding and cross-linking heads, a variety of natural or synthetic DNA binding molecules that bind DNA almost non-specifically can be used to capture all protein / DNA complexes genome-wide, including psoralen or derivatives (Cimino et al., Annu. Rev. Biochem. 54, 1151-1193 (1985)), DAPI (Kapuscinski, Biotech. Histochem. 70, 220-233 (1995)), and Hoechst dye (Carrondo et al., Biochemistry 28, 7849-7859 (1989)). (Psoralen preferentially binds to nucleosome-free regions of the genome, which may be a desirable feature for studying active chromatin regions of protein complexes) Specific DNA binding molecules can be used to target and capture protein-DNA complexes bound to specific loci, such as sequence-specific DNA binding polyamides (Nickols et al., Proc. Natl. Acad. Sci. USA 104, 10418-10423 (2007)) or G-quadruplex binders. For protein capture heads, we can use the photoaffinity labeling group diazirine (Mackinnon et al., Curr. Protoc. Chem. Biol. 1, 55-73 (2009)). Diazirine has a long UVA wavelength activation spectrum (340-365nm), similar to psoralen. The protein capture head can be one or more arms of diazirine, which are used to cross-link one or more proteins bound to DNA near the psoralen insertion site. The two heads are connected by a synthetic connection of variable length. Shorter linkers will more efficiently capture the proximal DNA binding domain of TF, while longer linkers will more efficiently capture distal activation / inhibition domains or auxiliary factors. Changing the arm length can also help improve capture efficiency and selectivity by reducing cross-linking back to DNA or cross-linking to non-specific proteins at a distance. Two exemplary methods are to select different multi-arm cores or connect diamines of different lengths to the protein capture head. The linker core can be a commercially available multi-PEG arm ring octyne linker molecule (Creative PEG works). The arms composed of PEG are water-soluble. The copper-free clickable DBCO end makes it bioorthogonal and assembly efficient, which can be used to extract and enrich handles, fluorescent labeling appendages or peptides. Cuttable MS indexes. The flexibility of the linker can also be adjusted by introducing unsaturated parts (e.g., double bonds, triple bonds) to reduce cross-linking back to DNA, thereby facilitating cross-linking with proteins bound to DNA. Figure 3D Examples of double bond linkers are shown. Figure 3B The general design using commercially available multi-PEG-armed octyne-linked molecules (Creative PEG works) is depicted.

[0359] Initial studies: The main goal of developing BFPX is to allow direct crosslinking between DNA and its binding proteins. This feature was first evaluated by in vitro assembled TF / DNA complexes with purified TF protein and synthetic DNA substrates, which can simplify the analysis of crosslink products and help quantify the crosslink yield. This result can in turn promote the optimization of probe and protocol design.

[0360] In preliminary studies, we have successfully synthesized four exemplary psoralen-based BFPX probes: 4AMT-LC-SDA, 4AMT-SDAD, SPB-AAD, and SPB-PEG3-AAD ( Figure 4 ). Psoralen is a plant natural product with excellent cell permeability and tolerance. It has little reactivity to proteins, but preferentially binds (intercalates) double-stranded DNA and RNA with μM (Kd) affinity. After UVA activation (~360nm), it forms stable covalent adducts with DNA and RNA. These properties of psoralen are conducive to its use in in situ studies of chromatin structure and protein-nucleic acid interactions. In various embodiments of the present invention, we believe that psoralen preferentially binds to open chromatin regions, making it an excellent choice for designing BFPX probes for capturing transcription factors, as transcription factors are known to bind almost exclusively to open chromatin regions. The general strategy of BFPX is applicable to other DNA binding heads that bind non-specifically or specifically to the genome. Using commercially available psoralen and diazirine derivatives, we synthesized a number of BFPX probes with a variety of linker design features ( Figure 4 In addition to 4AMT, we also extended the BFPX probe synthesis to other psoralen derivatives, such as SPB (succinimidyl-[4-(psoralen-8-yloxy)]-butyrate).

[0361] 4AMT-LC-SDA

[0362] The BFPX probe 4AMT-LC-SDA ( Figure 5A , Figure 5B). To test its photocrosslinking ability, human transcription factor MEF2A (MADS-box / MEF2 domain) and double-stranded DNA containing MEF2 binding sites were used. We first used electrophoretic mobility shift assay (EMSA) to verify the activity of purified recombinant MEF2A DNA binding domain. After finding the optimal protein to DNA ratio, in vitro assembled MEF2 / DNA complexes were photocrosslinked in the presence or absence of 4AMT-LC-SDA probe, with or without UV irradiation. Multiple other control experiments were performed in parallel to test the specificity of photocrosslinking between DNA and MEF2. Specifically, parameters such as protein / DNA amount (1:1 to 20:1), ratio, UV time (10m to 4hr), wavelength (320nm to 360nm), UV source (mercury long lamp, LED) and its distance from the reaction mixture have been optimized in the process. MEF2 protein and DNA were first incubated to mimic the in vivo binding mode and stoichiometry, and then the mixture was incubated with our 4AMT-LC-SDA probe for UV irradiation. By running a denaturing gel ( Figure 5D SDS or Figure 5C The results were evaluated by adding urea in the sample and checking the corresponding protein (silver stain) or DNA (FAM label) signals. In order to further exclude any other possible false positives during the loading process, all samples (including negative controls) were digested with proteinase K at 65 degrees and subjected to SDS.

[0363] The results showed that we could obtain successful photo-crosslinked complexes between proteins and DNA only in the presence of MEF2A, DNA, BFPX probe (4AMT-LC-SDA), and UV irradiation. In addition, under the same conditions, the non-DNA binding control protein BSA did not show any signs of forming crosslinks with DNA, demonstrating the high selectivity of our 4AMT-LC-SDA probe for targeting only DNA-binding proteins. By using agarose gel to separate the macrocomplexes, we compared the intensities of the bands corresponding to various molecular species, and the photo-crosslinking efficiency of the MEF2-DNA complex was estimated to be about 50% under the current conditions.

[0364] We further tested the activity of the probe using in vitro assembled transcription factor / DNA complexes, in which two human transcription factors, myocyte enhancer factor 2 (MEF2; here we used MEF2A) and nuclear factor of activated T cells (NFAT; here we used NFAT1), and their respective double-stranded DNA substrates (labeled with fluorescein, 5'6-FAM) were used for in vitro assays.

[0365] We used electrophoretic mobility shift assay (EMSA) to examine the effects of formaldehyde (FA), BFPX probe (4AMT-LC-SDA), and UVA (365 nm) on DNA binding of NFAT1. Figure 5G Obtain multiple important observations.First, although formaldehyde does not show any effect on single DNA (compare lane 2 and lane 1), it seriously weakens the DNA binding of NFAT1 (compare lane 4 and lane 3) (the band showing streaking upwards may represent the residual NFAT / DNA complex modified by formaldehyde outside the DNA binding surface). This observation is consistent with the view that formaldehyde can modify the lysine-rich DNA binding surface of TF and inhibit its DNA binding activity.By contrast, 4AMT-LC-SDS or single UV365 with and without UV365 do not affect NFAT1 protein (data not shown).Secondly, UVA (365nm), 4AMT-LC-SDA and combination thereof do not affect DNA mobility (compare lanes 5, 6 and 7) and the binding of NFAT1 to DNA (compare lanes 8, 9 and 10).The presence of both DNA and protein in the same shifted band can be checked by FAM fluorescence signal (top) and Coomassie brilliant blue staining (bottom) respectively (free NFAT is positively charged and cannot be observed under EMSA conditions). Finally, when the sample in lane 10 was run on a denaturing SDS gel, a covalent complex containing both protein and NFAT could be observed (see Figure 5H , lanes #1, 2, 3).

[0366] Through preliminary studies similar to these, we found that the psoralen-based BFPX probe ( Figure 4 ) did not show any detectable effect on NFAT1 / DNA binding at concentrations up to 150 μM, which is consistent with previous in vivo studies showing that psoralens do not affect endogenous transcription complexes at similar concentrations.

[0367] 4AMT-SDAD

[0368] The BFPX probe 4AMT-SDAD with a cleavable linker can greatly facilitate the mass spectrometric analysis of peptides released from cross-linked protein-DNA complexes after digestion with specific proteases. Fig. 6A , Figure 6B .

[0369] SPB-AAD

[0370] Next, we used SDS denaturing gels to examine the formation of covalent protein-DNA complexes and determine the efficiency of cross-linking, which was verified using the BFPX probe SPB-AAD. Figure 7As shown, for the MEF2A / DNA complex and NFAT1 / DNA complex assembled in vitro, the addition of the BFPX probe (SPB-AAD) in the presence of UVA (365nm) leads to the formation of complexes that run as larger molecular species than both free protein and DNA; More importantly, these larger complexes contain both DNA and protein, as shown by FAM signal and Coomassie Brilliant Blue staining (MEF2 / DNA complex, lane 4-lane 7; NFAT1 / DNA complex, lane 9-lane 12). The robust stability of these complexes under strong denaturing conditions (2% SDS and boiling for 10min) is consistent with the covalent nature of the cross-linking reaction designed in BFPX. In contrast, the control (untreated, lane 1; UV365 only, lane 2; and SPB-AAD only: lanes 3 and 8) only shows free protein and dissociated DNA under denaturing conditions. Surprisingly, even at a low probe concentration of 0.125 μM, significant amounts of cross-linked protein complexes were observed for MEF2 / DNA (lane 7) and NFAT1 / DNA (lane 12). As the probe concentration increased, more cross-linked protein / DNA complexes were formed, and the multiple bands likely represent protein-DNA complexes with varying degrees of cross-linking, since multiple SPB-AAD molecules can bind to the flanking regions of the DNA binding site and undergo cross-linking reactions with the protein. At high probe concentrations (lanes 5 and 10), the complexes became more uniform, likely due to saturation binding of the BFPX probe to DNA. The reduced cross-linking efficiency at the highest probe concentration (lanes 4 and 9) was due to precipitation of the probe in the aqueous solution.

[0371] In order to estimate cross-linking efficiency, we compared the fluorescence intensity of free DNA and the DNA complexed with protein.We found that in the presence of SPB-AAD and other BFPX probes we tested, DNA (both MEF2 DNA and NFAT1 DNA) did not show significant mobility changes in SDS gel after UV365 cross-linking (data not shown).In addition, Coomassie brilliant blue staining shows that almost all DNA bands that move up include protein.Under the protein excess conditions that most DNA is bound to protein, we estimate the cross-linking efficiency (percentage of covalently cross-linked natural protein-DNA complexes) according to the free DNA FAM signal of lane 5 (relative to lanes 1-3) and lane 10 (relative to lane 8) to be about 80%-90%.

[0372] We performed similar experiments with different BFPX probes and observed that the 4AMT-LC-SDA probe ( Figure 5H , Lanes #1, 2, 3) and 4AMT-LC-SDAD probe ( Figure 5H , Lanes #4, 5, 6; Figure 4 ) and SPB-PEG3-AAD probes ( Figure 8C ) when MEF2 / DNA complex and NFAT1 / DNA complex were continuously and efficiently cross-linked.

[0373] like Figure 5H As shown, 4AMT-LC-SDA efficiently crosslinks NFAT to its DNA substrate (compare lanes 2 and 3 with lane 1). Similarly, 4AMT-LC-SDAD efficiently crosslinks NFAT to its DNA substrate (compare lanes 5 and 6 with lane 4). DTT treatment has almost no effect on the NFAT-DNA complex crosslinked by 4AMT-LC-SDA (compare lanes 2 and lane 3). In contrast, DTT treatment significantly reduces the NFAT / DNA complex crosslinked by 4AMT-LC-SDAD (compare lanes 6 and lane 5). However, not all covalent NFAT-DNA complexes are dissociated by DTT treatment. The reducing conditions tested may not be sufficient to reduce the disulfide bonds in 4AMT-LC-SDAD, and stronger reducing agents and / or conditions may completely release the crosslinked covalent NFAT / DNA complex. Alternatively, the photocrosslinking reaction based on carbene may produce a covalent connection between protein and DNA, which cannot be cut by a reducing agent.

[0374] SPB-PEG3-AAD

[0375] The BFPX probe SPB-PEG3-AAD has a longer linker arm that facilitates cross-linking to protein domains that are further away from the protein-DNA binding interface. Fig. 8A , Figure 8B .

[0376] For the BFPX probe SPB-PEG3-AAD, when the probe concentration is high enough ( Figure 8C3 and 7 in the 5FAM gel), based on the DNA signal, the cross-linking efficiency for the MEF2 / DNA complex is over 90% (compare lane 3 and lane 1 in the 5FAM gel, the free DNA in lane 3 almost completely disappears compared to the strong band in lane 1). This is consistent with the protein signal (compare lane 3 and lane 1 in the CCB-250 stained gel, the free protein in lane 3 is very weak compared to the strong band in lane 1). For the NFAT / DNA complex, similarly high cross-linking efficiency was observed (compare lane 7 and lane 5 in the 5FAM gel, the free DNA in lane 7 almost completely disappears compared to the strong band in lane 5). Similarly, the protein signal shows consistent results (compare lane 7 and lane 5 in the CCB-250 stained gel, the free protein in lane 5 almost completely disappears compared to the strong band in lane 5). When the probe concentration is too high (lanes 2 and 6), the probe precipitates out of the solution, resulting in lower cross-linking efficiency.

[0377] These results demonstrate the possibility of engineering tunable molecular tools to efficiently and robustly capture protein-DNA complexes.

[0378] Identification of covalent attachment sites:

[0379] To further characterize the BFPX-mediated cross-linking reaction between protein and DNA, we identified the covalent attachment sites on the protein. We first expanded the Figure 7 The scale of the cross-linking reaction of lane 11 of .Then, we used mono-Q column to purify covalent NFAT / DNA complex on FPLC (Fig. 9, sub-figures a and b).The purified NFAT / DNA complex was digested with trypsin, and Mono-Q was used to perform another round of FPLC purification.The DNA attached with peptides was eluted like free DNA.The DNA-peptide conjugate was sent for Edman sequencing, and unique sequencing motifs Ile, Thr and Gly were produced, which corresponded to I479, T480 and G481 (Fig. 9, sub-figure c) of the human NFAT1 used in this study.I479 is after R478, which is consistent with the fact that the peptide is produced by trypsin digestion cut at the C-terminus of Lys and Arg. Surprisingly, when this result is compared with the crystal structure of the NFAT / DNA complex, assuming that the psoralen head inserts into the 5'TpA-3' of the DNA substrate used in the experiment and assuming that the linker in SPB-AAD is in an extended conformation, the diazirine head group is precisely pointed toward the loop containing I479, T480, and G481 ( Figure 9 , panel d).

[0380] Although it is generally believed that the carbenes generated by diazirine react nonspecifically with solvent molecules and DNA in addition to proteins, given the uncertainty about the DNA-intercalating activity of modified psoralen heads, our results strongly suggest that the psoralen head in the BFPX probe has a strong preference for binding to the TpA site despite synthetic modifications, and that the carbenes generated by UV activation with diazirine have a strong preference for cross-linking to proteins. If most of the carbenes react with solvent molecules and DNA, then we would not observe high cross-linking efficiencies. If the high cross-linking efficiencies are due to many BFPX probes binding to DNA that is cross-linked to proteins, then we would not see unique attachment sites on NFAT. These observations suggest that BFPX can be used as a mechanistically clear, accurate, and quantitative cross-linking technique for studying protein-nucleic acid interactions.

[0381] In various embodiments of the present invention, we further envision studying: (i) Development of BFPX protocols using in vitro TF complex model systems: We will extend the above studies to other transcription factors that have been purified in the laboratory and studied by crystallography. These TFs include p53, FOS, Jun, TonEBP, FOXP2, FOXP3, GATA3, NF-κB p50 and NF-κB p65, NKX2.5. This group of TFs represents a variety of DNA binding domain families that can help us determine the general applicability of the BFPX approach. In addition to direct DNA binding proteins, we will also test whether transcription cofactors recruited to DNA can be cross-linked to DNA using probes with longer linker lengths. To this end, we will use ternary TF complexes of MEF2 / Cabin1 / DNA complex, MEF2 / HDAC / DNA complex, and MEF2 / p300 / DNA complex. These are classic higher-order TF complexes in which transcription repressors (such as Cabin1 and class IIa HDACs) or activators (such as p300) are also recruited. Our structural studies of the above TF / DNA complexes provide a suitable model system to develop and optimize BFPX probes and related protocols. In addition, we will also use the XL-MS method recently described for protein-DNA complexes captured by short UV (~250nm) to map the atomic connection between protein and DNA and compare with our crystal structure. These studies will not only help cross-validate BFPX at the structural level, but also establish a general method for analyzing protein-DNA interactions using BFPX-based XL-MS. (ii) Further development of BFPX probes: multi-arm BFPX probes with variable linkers containing multifunctional features including length, cleavage, enrichment handles and fluorescent reporters.

[0382] Different linker lengths can be optimized for different experimental applications: shorter linkers are used for direct DNA binding domains, while longer linkers are used for recruited cofactors or distant functional domains. It is also possible to add cleavable enrichment handles. The advantage is that after purification of the cross-linked protein-DNA complexes, the proteins or their protease-digested peptide fragments can be released by cleavage of the linker to facilitate enrichment and subsequent protein analysis (e.g., LC / MS). We will also test whether we can improve the cross-linking efficiency and enable capture of multiple nearby proteins by introducing multiple diazirine protein capture arms. With the clickable core, various fluorescent tags can also be added to facilitate probe position tracking.

[0383] Another advantage of introducing multiple bis-aziridine arms is that it avoids the need to introduce photoaffinity labels on DNA-binding heads that do not have the inherent ability to photocrosslink to DNA like psoralen, since one of the bis-aziridine arms can serve as a crosslinking group for DNA while the other serves the DNA-binding protein. A general protocol for developing such multi-armed BFPX probes is outlined in Figure 3B Since most compounds are commercially available or can be gently azidated from amines (e.g., the primary amine of DAPI), the assembly should be chemically easy to carry out. Given that it is a multi-step assembly, the desired compounds will be purified (HPLC / silica gel chromatography) and confirmed by NMR in addition to mass spectrometry. The photocrosslinking test of these multi-arm probes will be similar to the experiments of the BFPX probe in the preliminary study described above.

[0384] One such protocol for developing multi-armed BFPX probes is described ( Figure 3B ). Since most compounds can be mildly azidated from amines (e.g., the primary amine of DAPI), the assembly should be chemically easy to perform. Given that it is a multi-step assembly, the desired compounds will be purified (HPLC / silica gel chromatography) and confirmed by NMR in addition to MS. Photocrosslinking testing of these multi-arm probes will be similar to the experiments with the 4AMT-LC-SDA probes described above.

[0385] Some alternative considerations: Psoralen is known to preferentially bind to nucleosome-free regions of the genome. While this may be a desirable feature for studying protein complexes in active chromatin regions, it may also be a limitation of such probes if one wants to study protein-DNA interactions in other genomic regions, such as dense heterochromatin regions. Therefore, we will also extend future BFPX probe designs to other DNA-binding molecules (such as DAPI and Hoechst dyes) that have different DNA-binding properties than psoralen ( Figure 3B, upper part). When designing BFPX probes using DAPI and Hoechst dyes as DNA-binding heads, we will use the crystal structures of their complexes bound to DNA to select the site for introducing the linker with the diaziridine photoaffinity labeling group ( Figure 3E ).

[0386] Application of BFPX probe in cells

[0387] To this end, we tested the prototype BFPX probe 4AMT-LC-SDA in the suspension GM12878 cell line. The results showed that BFPX exhibited good cross-linking effects compared to no-probe or no-UV controls ( Fig. 5F ).

[0388] We tested the capture of MEF2A / DNA complexes in the GM12878 cell line, a Tier 1 ENCODE cell line with abundant genomic data. Fig. 10A As shown, the MEF2A antibody detected two bands in untreated GM12878 (lane 4), which is similar to what the manufacturer (SCBT) showed for this antibody. While cells treated with SPB-PEG3-AAD alone (lane 1) or UV365 alone (lane 2) showed similar two bands for free MEF2A, cells treated with SPB-PEG3-AAD showed two upper bands and a diffuse band, indicating that UV365-induced crosslinking with SPB-PEG3-AAD generated a mixture of larger MEF2A complexes. This preliminary test indicates that UV365 and SPB-PEG3-AAD can capture MEF2 complexes in the nucleus.

[0389] We performed similar experiments using a different system (Hela cells transfected with FOXP3 labeled with AVI-TEV-FLAG). Again, we observed that the BFPX probe SPB-AAD could capture higher molecular weight FOXP3 complexes (presumably covalent FOXP3-DNA complexes) under 365 nm UV irradiation ( Fig. 10B ). This experiment shows that higher molecular weight FOXP3 complexes can only be observed when both SPB-AAD and UV irradiation are present (lanes 5 and 7) (compare lanes 5 and 7 with lanes 2, 3, 4, and 6). In addition, the formation of higher molecular weight FOXP3 complexes is dose-dependent on the SPB-AAD concentration (compare lane 7 with lane 5). Interestingly, the signal from the FA control lane (lane 1) is very weak, which may be due to FOXP3 being trapped in a large protein complex or the antibody epitope being modified and / or masked.

[0390] We also analyzed (i) the cellular permeability and subcellular distribution of the BFPX probe. We validated the cellular and nuclear permeability of the BFPX probe and monitored their subcellular distribution using azide-containing fluorescent tags via click reactions with alkyne groups (as in SPB-AAD and SPB-PEG3-AAD). Briefly, HeLa cells were incubated with SPB-AAD (25 μM) in the dark for 30 min to allow DNA binding and intercalation of the psoralen moiety of the BFPX probe. After washing away excess BFPX probe, the cells were irradiated with 365 nm UV for 5 min. Alexa Fluor 647 picolyl azide molecules (from Click-IT TM Plus ALEXAFLUOR TM Alexa Fluor 647 Picolyl Azide Toolkit) reacts with the alkyne group on SPB-AAD in DNA through a click reaction. After click chemistry labeling, unreacted fluorescent molecules are removed by washing. The cell nucleus is then analyzed using fluorescence imaging (Figure 11). These experiments show that Alexa Fluor 647 picolyl azide molecules are fixed in cells through a click reaction with the BFPX probe only when the BFPX probe is present and activated by UV cross-linking. It also shows that the BFPX probe can enter cells and nuclei with high efficiency. This approach has been successfully used to monitor the nuclear binding and distribution of psoralen derivatives. We have previously used this approach to monitor the subcellular location of alkyne-containing molecular tools (via STORM).

[0391] In various embodiments of the present invention, we also envision (ii) fine-tuning the BFPX-based cross-linking and extraction of TF / DNA complexes in cells. In addition to endogenous MEF2A in GM12878, we have selected two additional model systems, namely MDA-MB-231 cells expressing N-3xFLAG-GATA3 (doxycycline inducible) and HEK 293T cells expressing FOXP3 labeled with AVI-TEV-FLAG (transient transfection). These additional systems will allow us to verify the general applicability of BFPX in capturing different TF / DNA complexes in cells with different antibodies (FLAG) or tags (biotin) and different abundance levels. The formation of denaturation-resistant TF / DNA complexes will be monitored by SDS gel / western blot and further analyzed by nuclease (DNase I) and mass spectrometry. By monitoring the yield of covalent TF / DNA complexes from a given amount of cells (e.g., 1 million), we can fine-tune the BFPX in vivo cross-linking protocol by varying the probe concentration and UV dose. To facilitate various subsequent analyses (e.g., ChIP-seq and proteomics), we will develop a protocol for the preparative extraction of covalent TF / DNA complexes from free protein and DNA by modifying existing phenol / toluene / chloroform extraction methods for covalent protein / DNA or covalent protein / RNA complexes.

[0392] In various embodiments of the present invention, we also envision placing the BFPX probe into various (iv) test applications. We will apply BFPX in chromatin immunoprecipitation (ChIP) assays with sequencing (ChIP-seq) and Hi-C genomic analysis techniques to compare four different TFs (CTCF, GATA1, PU.1, and MEF2A) with formaldehyde in two suspension (K562, GM12878) cell lines and an adherent (Hela) cell line. These protein targets represent different classes of DNA binding domains, and their formaldehyde-based ChIP-seq data are available in all three cell lines, which will enable us to systematically compare formaldehyde-based and BFPX-based ChIP-seq datasets. Multiple metrics will be used to evaluate the performance of BFPX-based ChIP-seq. We will first check the reproducibility of the protocol using data correlation between technical replicates. We will also compare ChIP-seq peaks for the same TF between different cells to see if significant differences can be observed (i.e. above background noise defined by technical replicates), and if so, compare such differences to known signaling and gene expression differences between different cells. Such mechanistic analyses are often challenging (if not impossible) for formaldehyde ChIP-seq data of TFs due to high background noise. As mentioned earlier, a major and important contributor to this noise is DNA fragments co-captured with the target TF target through non-specific cross-linking of protein-protein complexes. In contrast, BFPX is designed to only (or at least preferentially) capture proteins bound or recruited to DNA. We propose to use the percentage of isolated DNA fragments containing the expected binding site of the targeted TF to assess the accuracy of formaldehyde-based and BFPA-based ChIP-seq datasets. This will provide another metric for quantitative evaluation of data quality.

[0393] In various embodiments of the invention, we also envision studying (v) adapting tag fragmentation. We will adapt BFPX to a new version of the ChIP-seq protocol that combines chromatin immunoprecipitation with sequencing library preparation by Tn5 transposase ("tag fragmentation"), also known as ChIPmentation. In the initial development of ChIPmentation, it was found that tag fragmentation of purified ChIP DNA led to the same problems of data and protocol reproducibility between samples and between antibodies mentioned previously. Part of the problem is that tag fragmentation is particularly sensitive to the ratio of DNA to transposase, which is highly variable and often too low to be quantified in standard ChIP protocols. In addition, purified ChIP DNA is already fragmented, and excess transposase can result in small fragments that are difficult to sequence. It turns out that tag fragmentation directly on immunoprecipitated bead-bound chromatin can produce improved results in terms of data quality and reproducibility, presumably because other chromatin proteins bound to the beads can protect the DNA from excessive tag fragmentation. These observations suggest that BFPX would be a perfect match for ChIPmentation, as BFPX is able to separate covalent protein-DNA complexes from a background of other protein cross-linked complexes. Protein-DNA complexes captured by BFPX can be anchored to streptavidin beads using an azido-biotin enrichment tag and fragmented by Tn5. By replacing formaldehyde, BFPX could greatly improve the performance of ChIPmentation in protein-DNA interaction studies.

[0394] In various embodiments of the present invention, we can also compare ChIP datasets obtained with different BFPX probes (with different DNA binding heads to avoid DNA binding bias; with different linker lengths to sample different DNA binding complexes). These analyses will reveal further insights into the BFPX protocol and guide its optimal application in the study of protein-DNA interactions in cells.

[0395] Other embodiments

[0396] Unless otherwise indicated, all reagents used for chemical syntheses were purchased from Sigma-Aldrich, Alfa Aesar, or EMD Millipore and used without further purification. All anhydrous reactions were performed under argon or nitrogen atmosphere. All reactions, purifications, and manipulations were performed in the dark, avoiding direct exposure to natural or artificial light. Analytical thin layer chromatography (TLC) was performed on EMD silica gel 60F254 plates with detection by cerium ammonium molybdate (CAM), anisaldehyde, or UV. For flash chromatography, HPLC was used. Silica gel (EMD). Acquired on a Varian spectrometer Mercury 400, VNMRS-500 or VNMRS-600 at 400 MHz, 500 MHz or 600 MHz 1 H spectra. Chemical shifts are reported in ppm (δ) relative to the solvent. Coupling constants (J) are reported in Hz. 13 C spectra were acquired on the same instrument at 100 MHz, 125 MHz or 150 MHz. Abbreviations for proton spectrum multiplicities are: s, singlet; b, broad; d, doublet; t, triplet; q, quartet; m, multiplet.

[0397] Biotage Isolera Spektra FLASH system (solvent A, 0.1% TFA in water; solvent B, 0.1% TFA in acetonitrile) or Agilent 1200 series HPLC (solvent A: 0.1% TFA in water; solvent B: 0.1% TFA and 90% acetonitrile in water) system was used for reverse phase high performance liquid chromatography (RP-HPLC). Mass spectra were recorded on an Agilent HPLC / Q TOFMS / MS spectrometer.

[0398] Example 3

[0399] Preparation of Compound 2 (SPB-AAD) N-(2-(3-(but-3-yn-1-yl)-3H-bis(aziridin-3-yl)ethyl)-4-((7-oxo-7H-furo[3,2-g]chromen-9-yl)oxy)butyramide.

[0400] 2-(3-(But-3-yn-1-yl)-3H-diaziridin-3-yl)ethan-1-amine 1 (25 mg, 0.18 mmol) was added to a solution of succinimidyl-[4-(psoralen-8-yloxy)]butyrate SPB-NHS (50 mg, 0.13 mmol) in 0.2 mL DMSO. The reaction was stirred at room temperature for 16 h, and after evaporation of the solvent, the crude reaction mixture was purified by reverse phase C-18 column chromatography (0.1% TFAH2O:ACN, 95:5 to 0:100) within 30 min to give compound 1 (20 mg, 38%).

[0401] 1H NMR (400MHz, CDCl3) δ7.8 (dd, J=9.6, 1.2Hz, 1H), 7.7-7.7 (m, 1H), 7.4 (s, 1H), 6.8-6.8 (m, 1H), 6.4 (d, J=1.2Hz, 1 H), 4.5-4.4(m, 2H), 3.2(q, J=7.0Hz, 2H), 2.7(t, J=7.0Hz, 2H), 2.2-2.1(m, 2H), 2.0-1.9(m, 3H), 1.7-1.6(m, 4H).

[0402] 13 C NMR (101MHz, CDCl3) δ173.1, 161.0, 146.9, 144.8, 131.5, 126.2, 116.4, 114.5, 114.0, 106.9, 73.3, 69.2, 34.5, 33.0, 32.5, 31.9, 26.3, 13.2.

[0403] HRMS(ESI):C 22 H 21 N3O5(M+H + )Calculated value 408.1559, measured value 408.1589.

[0404] Example 4

[0405] Preparation of Compound 4 tert-butyl (18-(3-(but-3-yn-1-yl)-3H-diaziridin-3-yl)-15-oxo-3,6,9,12-tetraoxa-16-azaoctadecyl)carbamate.

[0406] A mixture of 2,2-dimethyl-4-oxo-3,8,11,14,17-pentaoxa-5-azaeicosane-20-oic acid 3 (133 mg, 364 μmol), 1-ethyl-3-(3-dimethylaminopropyl)carbodiimide (70 mg, 364 μmol), N-hydroxysuccinimide (42 mg, 364 μmol) in DMF (600 μL) was stirred at room temperature for 45 min.

[0407] To this mixture was added a solution of 2-(3-(but-3-yn-1-yl)-3H-bis-aziridine-3-yl)ethyl-1-amine 1 (25 mg, 182 μmol) in DMF (300 μL). The reaction was stirred at room temperature for 18 h. The reaction was stirred at room temperature for 18 h, and after evaporation of the solvent, the crude reaction mixture was purified by reverse phase C-18 column chromatography (H2O: MeOH, 100: 0 to 0: 100) over 15 column volumes (CV) to give compound 4 as a colorless oil (50 mg, 57%).

[0408] 1 H NMR (400MHz, CDCl3) δ6.7 (s, 1H), 5.1 (s, 1H), 3.7 (t, J=5.8Hz, 2H), 3.6-3.6 (m, 11H), 3.5 (t, J=5.2Hz, 2H), 3.3-3.2 (m, 2H), 3.1-3.0 (m, 2H), 2.4 (t, J=5.7Hz, 2H), 2.1 (s, 1H), 2.0-1.9 (m, 3H), 1.6 (td, J=7.2, 2.3Hz, 4H), 1.4 (s, 9H).

[0409] 13 C NMR (101MHz, CDCl3) δ171.7, 155.9, 125.4, 121.7, 82.6, 70.5, 70.46, 70.4, 70.24, 70.22, 70.17, 70.14, 40.3, 36.8, 34.2, 32.5, 32.1, 28.4, 26.9, 13.2.

[0410] HRMS(ESI):C 23 H 41 N4O7(M+H + )Calculated value 485.2975, measured value 485.2990.

[0411] Example 5

[0412] Preparation of N-(2-(3-(but-3-yn-1-yl)-3H-bis(aziridin-3-yl)ethyl)-1-(4-((7-oxo-7H-furo[3,2-g]chromen-9-yl)oxy)butyranamido)-3,6,9,12-tetraoxapentadecan-15-amide (6) (SPB-PEG4-AAD).

[0413] Compound 4 (25 mg, 51 μmol) was dissolved in TFA / DCM [1:1 (v / v) 1.5 mL] solution and stirred at room temperature for 30 min. The reaction was concentrated under vacuum and co-evaporated with toluene three times to give crude compound 5. N,N-diisopropylethylamine (18 μL, 103 μmol) was added to a solution of crude amine 5 and SPB-NHS (25 mg, 51 μmol) in 500 μL of anhydrous DMF. The reaction was stirred at room temperature for 16 h, and then the solvent was evaporated.

[0414] The reaction mixture was purified by reverse phase C-18 column chromatography (H2O:MeOH, 100:0 to 0:100) over 15 column volumes (CV) to give compound 6 as a colorless oil (18 mg, 53% for 2 steps).

[0415] 1 H NMR (400MHz, CDCl3) δ7.8-7.8 (m, 1H), 7.7 (d, J = 2.2Hz, 1H), 7.4 (s, 1H), 6.8 (d, J = 2.2Hz, 1H), 6.7-6.7 (m, 1H), 6.4 (d, J=9.5Hz, 1H), 4.5 (t, J=5.8Hz, 2H), 3.7 (t, J=5.5Hz, 2H), 3.6 -3.6 (m, 11H), 3.6-3.5 (m, 2H), 3.5-3.4 (m, 2H), 3.1 (q, J=5.8Hz, 2H), 2.6 (t, J=6.8Hz, 2H) , 2.5(t, J=5.8Hz, 2H), 2.2-2.1(m, 2H), 2.0-2.0(m, 2H), 1.7-1.6(m, 3H), 1.3-1.2(m, 2H).

[0416] 13 C NMR (101MHz, CDCl3) δ172.7, 171.7, 160.7, 146.8, 144.6, 131.6, 126.1, 116.4, 114.6, 113.5, 106.8, 82.7, 73.2, 70.5, 70.5, 70.4, 70.3, 70.2, 70.1, 69.9, 69.4, 67.1, 39.2, 36.9, 34.2, 32.7, 32.1, 26.9, 26.1, 13.2.

[0417] HRMS(ESI):C 33 H 43 N4O 10 (M+H + )Calculated value: 655.2979, measured value: 655.2964.

[0418] Example 6

[0419] Preparation of compound 9

[0420] tert-Butyl(3-((tert-butoxycarbonyl)amino)propyl)(4-(6-(3-(3-methyl-3H-diaziridin-3-yl)propionamido)hexanamido)butyl)carbamate.

[0421] Triethylamine (32 μL, 246 μmol) was added to a solution of 2,5-dioxopyrrolidin-1-yl 6-(3-(3-methyl-3H-bis(aziridin-3-yl)propionamido)hexanoate (8) (40 μL, 118 μmol) and tert-butyl(4-aminobutyl)(3-((tert-butoxycarbonyl)amino)propyl)carbamate (7) (49 mg, 141 μmol) in DMF (2.1 mL).

[0422] The reaction was concentrated in vacuo and purified by reverse phase C-18 column chromatography (H2O:MeOH, 100:0 to 0:100) over 15 column volumes (CV) to afford compound 9 as a colorless oil (62 mg, 92%).

[0423] 1 H NMR (400MHz, CDCl3) δ6.8-6.5 (m, 1H), 3.6-3.3 (m, 10H), 2.4 (t, J=7.4Hz, 2H), 2.3 (dd, J=9.0, 6.7Hz, 2H), 2.0 (dd, J=8.8 , 6.7Hz, 2H), 2.0-1.9 (m, 4H), 1.8 (dd, J=15.0, 7.5Hz, 7H), 1.7 (s, 9H), 1.7 (s, 9H), 1.6-1.6 (m, 2H), 1.3 (d, J=0.7Hz, 2H).

[0424] 13 C NMR (101MHz, CDCl3) δ171.8, 39.5, 36.6, 30.9, 30.4, 29.3, 28.7, 26.6, 25.8, 25.4, 20.2.

[0425] HRMS(ESI):C 28 H 53 N6O6(M+H + )Calculated value: 569.4027, measured value: 569.4057.

[0426] Example 7

[0427] Preparation of compound 11 (SPB-spermidine-AD)

[0428] 6-(3-(3-methyl-3H-bis(aziridin-3-yl)propionamido)-N-(4-((3-(4-((7-oxo-7H-furo[3,2-g]chromen-9-yl)oxy)butanamido)propyl)amino)butyl)hexanamide.

[0429] Compound 9 (62 mg, 109 μmol) was dissolved in TFA / DCM [1:1 (v / v) 1.5 mL] solution and stirred at room temperature for 30 min. The reaction was concentrated in vacuo and co-evaporated with toluene three times to obtain a crude compound 10.

[0430] N,N-diisopropylethylamine (19 μL, 109 μmol) was added to a solution of crude amine 10 and SPB-NHS (42 mg, 109 μmol) in 500 μL of anhydrous DMF. The reaction was stirred at room temperature for 16 h, and then the solvent was evaporated.

[0431] The reaction mixture was purified by reverse phase C-18 column chromatography (H2O:MeOH, 100:0 to 0:100) over 15 column volumes (CV) to give compound 11 as a colorless oil (18 mg, 53% over 2 steps).

[0432] 1 H NMR (400MHz, CDCl3) δ7.8 (dd, J=9.6, 0.8Hz, 1H), 7.7 (dd, J=2.3, 0.8Hz, 1H), 7.4 (s, 1H), 6.8 (dd, J =2.3, 0.8Hz, 1H), 6.4 (dd, J=9.6, 0.8Hz, 1H), 4.5 (t, J=5.6Hz, 2H), 3.5 (d, J=0.8Hz, 2H), 3.3 (q, J=6 .2Hz, 2H), 3.3-3.2(m, 4H), 2.6(t, J=6.5Hz, 2H), 2.6-2.6(m, 4H), 2.2-2.1(m, 5H), 2.0(dd, J=8.7, 6 .8Hz, 2H), 1.7-1.7(m, 4H), 1.6-1.6(m, 3H), 1.6-1.4(m, 5H), 1.4-1.3(m, 2H), 1.0(d, J=0.9Hz, 3H).

[0433] 13C NMR (100MHz, CDCl3) δ173.1, 173.0, 171.5, 161.0, 146.9, 144.9, 126.3, 116.4, 114.4, 113.7, 106.8, 73.2, 49.0, 47.3, 39.2, 39.1, 37.9, 36.3, 33.0, 30.6, 30.1, 29.0, 28.9, 27.2, 27.0, 26.2, 25.0, 19.9.

[0434] HRMS(ESI):C 33 H 47 N6O7(M+H + )Calculated value: 639.3506, measured value: 639.3602.

[0435] Various embodiments of the present invention are described in the above detailed description. Although these descriptions directly describe the above embodiments, it should be understood that modifications and / or variations to the specific embodiments shown and described herein may be conceived by those skilled in the art. Any such modifications or variations falling within the scope of this specification are also intended to be included therein. Unless otherwise specified, it is the intention of the inventor to give the words and phrases in the specification and claims the common and customary meanings to those of ordinary skill in the applicable field.

[0436] The foregoing description of various embodiments of the present invention known to the applicant at the time of filing the application has been presented and is intended for the purpose of illustration and description. This specification is not intended to be exhaustive or to limit the present invention to the precise form disclosed, and many modifications and variations are possible in accordance with the above teachings. The described embodiments are used to explain the principles of the present invention and its practical application, and to enable other technical personnel in the field to utilize the present invention in various embodiments with various modifications adapted to the intended specific use. Therefore, it is not intended to limit the present invention to the specific embodiments disclosed for the implementation of the present invention.

[0437] Although specific embodiments of the present invention have been shown and described, it will be apparent to those skilled in the art that changes and modifications may be made without departing from the present invention and its broader aspects based on the teachings herein, and therefore, the appended claims will cover within their scope all such changes and modifications that are within the true spirit and scope of the present invention. It will be understood by those skilled in the art that, in general, the terms used herein are generally intended to be "open" terms (e.g., the term "including" should be interpreted as "including but not limited to", the term "having" should be interpreted as "at least having", the term "includes" should be interpreted as "including but not limited to", etc.).

Claims

1. Compounds of formula (I): A-L1-(C) n -((L2) n -B) m Formula (I) in: A represents a nucleic acid binding functional group derived from psoralen, methyl trimethoxalan, benzophenone, 4',6-diamidino-2-phenylindole (DAPI), Hoechst dye, polyamide or G-quadruplex binding molecule, ethopropaldehyde or its derivatives; L1 does not exist or represents the first linker; When n=1, C represents a core moiety having at least two functional groups, each functional group being used to attach to L1 and to attach to at least one arm represented by L2-B, respectively; or when n=0, C is absent; When n=1, L2 independently represents the second linker of each arm represented by L2-B; or when n=0, L2 does not exist; For each of the arms, B independently represents: A photoreactive functional group, wherein the photoreactive functional group comprises diaziridine or a derivative thereof or an aryl azide or a derivative thereof, and optionally, the aryl azide or a derivative thereof is selected from phenyl azide, o-hydroxyphenyl azide, m-hydroxyphenyl azide, tetrafluorophenyl azide, o-nitrophenyl azide, m-nitrophenyl azide or azidomethylcoumarin; or Detectable functional groups; wherein, in at least one of the arms, B represents a photoreactive functional group; n = 0 or 1; m represents (L2) n -B represents the number of arms, wherein when n=1, m is an integer of 1 or more, or when n=0, m=1.

2. The compound according to claim 1, wherein At least one of L1 and L2 is not absent, and the at least one of L1 and L2 is cleavable.

3. The compound according to claim 2, wherein L1, L2 or both independently comprise one or more of a sulfoxide-containing mass spectrometry (MS) cleavable bond, an acid cleavable CS bond, a disulfide group and an azo group.

4. The compound according to any one of claims 1 to 3, wherein n=0, m=1, and the compound is represented by formula (II): A-L1-B formula (II), in, L1 is absent or is the first linker.

5. The compound according to claim 4, wherein: A is an amine-containing or amine-reactive derivative of psoralen, an amine-containing or amine-reactive derivative of methyltrimethsalen, an amine-containing or amine-reactive derivative of benzophenone, an amine-containing or amine-reactive derivative of 4',6-diamidino-2-phenylindole (DAPI), an amine-containing or amine-reactive derivative of Hoechst dye, an amine-containing or amine-reactive derivative of polyamide, or an amine-containing or amine-reactive derivative of a G-quadruplex binding molecule, or an amine-containing or amine-reactive derivative of ethoxybutyraldehyde, optionally A is derived from succinimidyl-[4-(psoralen-8-yloxy)]-butyrate (SPB) or 4'-aminomethyltrimethsalen (4AMT); B comprises diaziridine or diaziridine alkyne, optionally aminodiaziridine alkyne (AAD); and L1 is absent or is a first linker, wherein the first linker comprises one or more of the following: (i) a cleavable bond, (ii) an oligomer or polymer having -OCH2CH2- repeating units, and (iii) an unsaturated portion, the unsaturated portion being optionally selected from a carbon-carbon double bond, a carbon-carbon triple bond, or an aromatic group.

6. The compound according to claim 5, wherein L1-B is derived from succinimidyl 6-(4,4'-azidopentanamido)hexanoate (NHS-LC-SDA), succinimidyl 2-((4,4'-azidopentanamido)ethyl)-1,3'dithiopropionate (NHS-SS-diaziridine) or 2-(3-(but-3-yn-1-yl)-3H-diaziridine-3-yl)ethane-1-amine (AAD); and / or wherein A is derived from 4'-aminomethyltrimethsalin (4AMT) or succinimidyl-[4-(psoralen-8-yloxy)]-butyrate (SPB); and wherein optionally, the photocrosslinking molecule is represented by formula (IIa) or formula (IIc):

7. The compound according to claim 5, wherein: A is derived from succinimidyl-[4-(psoralen-8-yloxy)]-butyrate (SPB) or 4'-aminomethyltrimethylsalen (4AMT); B comprises diaziridine or diaziridine alkyne, optionally aminodiaziridine alkyne (AAD); and L1 is the first linker, the first linker comprising one or more of the following: (i) a cleavable bond, (ii) an oligomer or polymer having a -OCH2CH2- repeating unit, and / or (iii) an unsaturated portion, the unsaturated portion being optionally selected from a carbon-carbon double bond or an aromatic group; and wherein optionally, the photo-crosslinking molecule is represented by formula (IIb), formula (IId), formula (IIe) or formula (IIf):

8. The compound according to claim 4, wherein: A is selected from the group consisting of: Among them, R 1 R is independently H, halo, OH, optionally substituted alkoxy or optionally substituted alkyl; 2 are independently H, halo, OH, optionally substituted alkoxy, or optionally substituted alkyl; a is 0, 1, 2, 3, 4, or 5; and b is 0, 1, 2, 3, or 4; Among them, R 3 R is independently H, halo, OH, optionally substituted alkoxy or optionally substituted alkyl; 4 is H, halo, OH, optionally substituted alkoxy or optionally substituted alkyl; R 5 are independently H, halo, OH, optionally substituted alkoxy, or optionally substituted alkyl; c is 0, 1, 2, 3, or 4; and d is 0, 1, 2, 3, or 4; Among them, R 6 is H, halo, OH, optionally substituted alkoxy or optionally substituted alkyl; R 7 is H, halo, OH, optionally substituted alkoxy or optionally substituted alkyl; R 8 is H, halo, OH, optionally substituted alkoxy or optionally substituted alkyl; and R 9 is H, halo, OH, optionally substituted alkoxy or optionally substituted alkyl; Among them, R 10 is H, halo, OH, optionally substituted alkoxy or optionally substituted alkyl; R 11 is H, halo, OH, optionally substituted alkoxy or optionally substituted alkyl; and R 12 is H, halo, OH, optionally substituted alkoxy or optionally substituted alkyl; as well as L1 is absent or L is selected from the group consisting of: Where q is 0, 1, 2, 3 or 4; Where p is 0, 1, 2, 3 or 4; Among them, R 13 is independently H, halo, OH, optionally substituted alkoxy or optionally substituted alkyl; and e is 0, 1, 2, 3 or 4; Wherein, r is 0, 1, 2, 3 or 4; Where s is 0, 1, 2, 3 or 4; Where t is 0, 1, 2, 3 or 4; u is 0, 1, 2, 3, or 4; as well as and B is selected from the group consisting of:

9. The compound according to claim 4, wherein: A is selected from the group consisting of: L1 is absent or L1 is selected from the group consisting of: and B is selected from the group consisting of:

10. The compound according to claim 1 or claim 4, wherein The compound is:

11. The compound according to claim 5, wherein L1 comprises from 2 to 20 carbons or from 20 to 100 carbons in length.

12. The compound according to any one of claims 1 to 3, wherein n=1, m is an integer of 2 or greater, and C represents a core portion having at least three functional groups, each functional group being used to attach to L1 and to attach to at least two arms each represented by (L2-B), such that the compound is represented by formula (III):

13. The compound according to claim 12, wherein In one of the at least two arms, B comprises diaziridine or diaziridine azide, and in the other of the at least two arms, B represents a detectable functional group, which comprises a fluorophore, biotin, a chromophore, a chromogen, a quantum dot, a fluorescent microsphere or a nanoparticle.

14. The compound according to claim 12 or claim 13, wherein L1, L2 or both independently comprise one or more of the following: (i) a cleavable bond, (ii) an oligomer or polymer having repeating units of -OCH2CH2-, and (iii) an unsaturated moiety.

15. A compound according to any one of claims 9 to 14, wherein C represents a dendritic core moiety comprising at least three surface functional groups, each surface functional group being used to attach to L1 and to the at least two arms each represented by L2-B, respectively.

16. A compound according to any one of claims 9 to 15, wherein L1, L2 or both independently comprise a triazole bonded to A.

17. A method for cross-linking nucleic acids and proteins in a system, the method comprising: Providing a compound according to any one of claims 1-16; Providing a system, wherein the system comprises a nucleic acid and a protein; contacting the compound with the system; and The system and the compound are irradiated with ultraviolet light under conditions effective to cross-link the nucleic acid to the protein.

18. The method according to claim 17, wherein: The system is a living cell.

19. The method according to claim 17 or claim 18, wherein: The wavelength of the ultraviolet light is between 300nm and 370nm.

20. The method of any one of claims 17-19, further comprising using the system to perform one or more of immunoprecipitation, chromatin precipitation, 3D chromatin conformation capture, mass spectrometry, and electrophoresis.

21. The method according to any one of claims 17 to 20, wherein: Element L1, L2 or both of the compound are independently cleavable, and the method further includes adding a cleavage agent to the system to cleave the element L1, L2 or both; or wherein element A of the compound is derived from psoralen, and the method further includes applying ultraviolet light with a wavelength of about 230 nm to cleave the element A; thereby generating a fingerprint of cross-linked proteins adjacent to nucleic acids in the system.

22. A method for preparing a compound according to any one of claims 12 to 16, comprising: Providing azide derivatives of nucleic acid-binding photoreactive agents, the agents comprising psoralen, methyltrimethoxalen, benzophenone, 4',6-diamidino-2-phenylindole (DAPI), Hoechst dye, polyamide or G-quadruplex binding molecules, ethopropaldehyde or derivatives thereof; Providing an azide derivative of a photoreactive reagent comprising a diaziridine moiety to obtain an azide-diaziridine bifunctional photoreactive reagent, and the photoreactive reagent optionally further comprises an alkyne group, or providing an aryl azide, the aryl azide optionally selected from phenyl azide, o-hydroxyphenyl azide, m-hydroxyphenyl azide, tetrafluorophenyl azide, o-nitrophenyl azide, m-nitrophenyl azide or azide-methylcoumarin; optionally providing an azide derivative of a detectable agent, the detectable agent comprising a fluorophore, biotin, a chromophore, a chromogen, a quantum dot, a fluorescent microsphere or a nanoparticle; providing a multi-arm reagent having at least three functional groups, each functional group independently comprising an alkyne; as well as Each azide derivative and, if provided, an aryl azide are combined with the multi-arm reagent in one reaction vessel to prepare the compound.

23. The method according to claim 22, wherein: The multi-arm reagent has at least three functional groups, each of which independently comprises a cyclooctyne group.

24. A method according to claim 22 or claim 23, wherein: The nucleic acid-binding photoreactive reagent comprises a first primary amine functional group, and providing an azide derivative of the nucleic acid-binding photoreactive reagent comprises converting the first primary amine functional group to a first azide-containing moiety, optionally by reacting the nucleic acid-binding photoreactive reagent with imidazole-1-sulfonyl azide; and / or wherein, The photoreactive reagent comprising a diaziridine moiety further comprises a second primary amine functional group or is modified with a second primary amino functional group, and providing an azide derivative of the photoreactive reagent comprises converting the second primary amine functional group into a second azide-containing moiety, optionally by reacting the photoreactive reagent with imidazole-1-sulfonyl azide.

Citation Information

Patent Citations

  • Electro-magnetic instructional and amusement device

    US3231988A

  • Single polypeptide chain binding molecules

    US4946778A

  • Humanized immunoglobulins

    US5585089A