Arginine reactive probes and uses thereof
Patent Information
- Application Number
- PCT/US2026/021282
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2025-03-28
- Filing Date
- 2026-03-27
- Publication Date
- 2026-10-01
Smart Images

Figure US2026021282_01102026_PF_FP_ABST
Abstract
Description
Atty. Docket: UCSF-865WOARGININE REACTIVE PROBES AND USES THEREOFCROSS-REFERENCE TO RELATED APPLICATION
[0001] Pursuant to 35 U.S.C. § 119(e), this application claims priority to the filing date of United States Provisional Patent Application Serial No. 63 / 780,066 filed March 28, 2025, the disclosure of which is herein incorporated by reference in its entirety.BACKGROUND OF THE INVENTION
[0002] Covalent therapeutics have revolutionized drug discovery by enabling precise and irreversible engagement of protein targets. While cysteine has been the predominant residue leveraged for covalent modifications, selectively targeting amino acid side chains beyond cysteine is of high interest but remains an ongoing challenge. Lysine reactivity has been explored but pan-specific probes targeting lysine are often reversible, requiring the introduction of reducing agents in order to render engagement irreversible and characterize reactivity through traditional chemical proteomic workflows.
[0003] Arginine, with its positively charged guanidinium moiety, plays indispensable roles in enzymatic catalysis, molecular recognition, and nucleic acid interactions. Its guanidinium group engages in a diverse array of noncovalent interactions, including hydrogen bonding and cation-7t contacts, enabling arginine to stabilize transition states, mediate protein-protein interfaces, and recognize aromatic moieties within biomolecular targets. However, its high pKa. resonance-stabilized charge distribution, and limited nucleophilicity complicate the development of covalent modification strategies. Although previous approaches, such as a,p-unsaturated amides and phenyl glyoxal derivatives, have been explored, they often suffer from poor biocompatibility and moderate selectivity.
[0004] The development of a selective and robust warhead for covalent arginine modification would provide a powerful tool for studying arginine reactivity and function in complex biological contexts and facilitate the design of new therapeutic strategies.Atty. Docket: UCSF-865WOGENERAL INFORMATION
[0005] Before the present methods and uses are described, it is to be understood that this invention is not limited to particular steps, devices and compounds described, as such may, of course, vary. It is also to be understood that the terminology used herein is for the purpose of describing particular embodiments only, and is not intended to be limiting, since the scope of the present invention will be limited only by the appended claims.
[0006] Where a range of values is provided, it is understood that each intervening value, to the tenth of the unit of the lower limit unless the context clearly dictates otherwise, between the upper and lower limits of that range is also specifically disclosed. Each smaller range between any stated value or intervening value in a stated range and any other stated or intervening value in that stated range is encompassed within the invention. The upper and lower limits of these smaller ranges may independently be included or excluded in the range, and each range where either, neither or both limits are included in the smaller ranges is also encompassed within the invention, subject to any specifically excluded limit in the stated range. Where the stated range includes one or both of the limits, ranges excluding either or both of those included limits are also included in the invention.
[0007] Unless defined otherwise, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention belongs. Although any methods and materials similar or equivalent to those described herein can be used in the practice or testing of the present invention, some potential and preferred methods and materials are now described. All publications mentioned herein are incorporated herein by reference to disclose and describe the methods and / or materials in connection with which the publications are cited. It is understood that the present disclosure supersedes any disclosure of an incorporated publication to the extent there is a contradiction.
[0008] It must be noted that as used herein and in the appended claims, the singular forms "a", "an", and "the" include plural referents unless the context clearly dictates otherwise. Thus, for example, reference to "a scan" includes a plurality of such scans and reference to "the Biomarker of Response" includes reference to one or more such Biomarkers of Response and equivalents thereof known to those skilled in the art, and so forth.Atty. Docket: UCSF-865WO
[0009] The publications discussed herein are provided solely for their disclosure prior to the filing date of the present application. Nothing herein is to be construed as an admission that the present invention is not entitled to antedate such publication by virtue of prior invention.Further, the dates of publication provided may be different from the actual publication dates which may need to be independently confirmed.
[0010] All patents and publications, including all sequences disclosed within such patents and publications, referred to herein are expressly incorporated by reference.BRIEF SUMMARY OF THE INVENTION
[0011] In another aspect, the invention provides a modified arginine-containing protein comprising: a small molecule fragment moiety, covalently bonded to an arginine residue of an arginine-containing protein, wherein a covalent bond is formed by reaction with a non-naturally occurring small molecule probe having a structure described herein, such as Formula (I): F1— Q (I), wherein F1is a small molecule fragment moiety comprising an azide moiety, an alkyne moiety, a fluorophore moiety, a labeling group, or a combination thereof; and Q comprises a cyclic trione moiety, a cyclic dihydroxy dione moiety, a cyclic dialkoxy dione moiety, or a cyclic hydroxy alkoxy dione moiety. In an exemplary embodiment, wherein Q is ninhydrin. In an exemplary embodiment, wherein the arginine residue is attached to the small molecule fragment through an imidazolidinimine moiety. In an exemplary embodiment, wherein F1is an alkyne moiety. In an exemplary embodiment, wherein F1is a labeling group which comprises desthiobiotin. In an exemplary embodiment, wherein the arginine-containing protein is described herein.
[0012] In another aspect, the invention provides a modified arginine-containing protein comprising: a small molecule fragment moiety, covalently bonded to an arginine residue of an arginine-containing protein, wherein a covalent bond is formed by reaction with a non-naturally occurring ligand electrophile having a structure of Formula (X): F2— Q (X), wherein F2is a small molecule fragment moiety; and Q comprises a cyclic trione moiety. In an exemplary embodiment, wherein Q is ninhydrin. In an exemplary embodiment, wherein the arginine residue is attached to the small molecule fragment through an imidazolidinimine moiety. In an exemplary embodiment, wherein F2comprises Ci-Ce alkyl, Ci-Ce fluoroalkyl, Ci-Ce heteroalkyl,Atty. Docket: UCSF-865WOa substituted or unsubstituted C3-C6 cycloalkyl, a substituted or unsubstituted C2-C6 heterocycloalkyl, a substituted or unsubstituted aryl, or a substituted or unsubstituted heteroaryl. In an exemplary embodiment, wherein the arginine-containing protein is described herein.
[0013] In another aspect, the invention provides a compound having a structure according to onFormula (la): O (la), wherein F1is a small molecule fragment moiety comprising an azide moiety, an alkyne moiety, a fluorophore moiety, a labeling group, or a combination thereof. In an exemplary embodiment, the compound has a structure according to:OZ-Y-XYXVA / 0RAO (le), wherein Z is an alkyne, a fluorophore, or a labeling group; Y is substituted or unsubstituted Ci-Ce alkylene; X is -CH2-, O, -C(O)NH, -NH-, -OC(O)O-, triazole, sulfonate, and phosphonate; and Raand Rbare independently selected from unsubstituted methyl, ethyl, propyl, isopropyl, butyl, isobutyl, and sec-butyl.
[0014] In another aspect, the invention provides a method of identifying a reactive arginine of a protein, comprising: (a) providing a protein sample comprising isolated proteins, living cells (such as primary immune cells), cell lysate, or tissue (such as blood); (b) contacting the protein sample with a probe compound of Formula (I) at a first concentration for a time sufficient for the probe compound to react with the reactive arginine of the protein sample; and (c) analyzing the proteins of the protein sample to identify the reactive arginine that bound with the probe compound at the first concentration; wherein the probe compound has a structure represented by Formula (I): F1— Q (I), wherein F1is a small molecule fragment moiety comprising an alkyne moiety, a fluorophore moiety, a labeling group, or a combination thereof; and Q comprises a cyclic trione moiety.
[0015] In another aspect, the invention provides a method of identifying a reactive arginine of a protein, comprising: (a) providing a protein sample comprising isolated proteins, living cells, or a cell lysate and separating the protein sample into a first protein sample and a second protein sample; (b) contacting the first protein sample with a probe compound of Formula I at a firstAtty. Docket: UCSF-865WOconcentration for a time sufficient for the probe compound to react with a reactive arginine of the first protein sample, and contacting the second protein sample with the probe compound of Formula (I) at a second concentration for a sufficient time for the probe compound to react with a reactive arginine of the second protein sample; (c) analyzing the proteins of the first protein sample and the second protein sample of step (b) to identify the reactive arginines that bound with the probe compound; (d) comparing the identity of the reactive arginines of step (c) from the first protein sample at the first concentration of probe compound to the reactive arginines from the second protein sample at the second concentration of probe compound; and (e) based on step (d), determining a reactive arginine of a protein; wherein the probe compound has a structure represented by Formula (I): F1— Q (I), wherein F1is a small molecule fragment moiety comprising an alkyne moiety, a fluorophore moiety, a labeling group, or a combination thereof; and Q comprises a cyclic trione moiety.
[0016] In another aspect, the invention provides a method of identifying a protein that interacts with a ligand of interest, comprising: (a) providing a protein sample comprising isolated proteins, living cells, or a cell lysate and separating the protein sample into a first protein sample and a second protein sample; (b) contacting the first protein sample with a ligand for a sufficient time for the ligand to react with a reactive arginine of the first protein sample; (c) contacting the first protein sample and the second protein sample with a probe compound of Formula (I) for a sufficient time for the probe compound to react with the reactive arginines of the first and second protein samples; (d) analyzing the proteins of the first and second protein samples to identify the reactive arginines that bound with the probe compound; (e) comparing the reactivity of the reactive arginine from the first protein sample to the reactivity of the reactive arginine from the second protein sample, wherein a decrease in the reactivity of the reactive arginine of the first protein sample relative to the reactive arginine of the second protein sample indicates interaction of the ligand with the reactive arginine of the first protein sample; and (f) determining the protein comprising the reactive arginine of the first protein sample that interacts with the ligand; wherein the probe compound has a structure represented by Formula (I): F1— Q (I), wherein F1is a small molecule fragment moiety comprising an alkyne moiety, a fluorophore moiety, a labeling group, or a combination thereof; and Q comprises a cyclic trione moiety.
[0017] Covalent molecules have emerged as powerful tools and next-generation therapeutics, promising improved specificity and sustained target engagement. Chemical proteomic methodsAtty. Docket: UCSF-865WOto screen for reactive cysteine and lysine residues on proteins amenable to this approach have also transformed small molecule discovery pipelines. These molecules and platforms have improved our understanding of fundamental biological processes and changed the way we treat human disease, but their application remains limited to sites where reactive cysteines and lysines are present. To expand the utility of covalent molecules, we developed a ninhydrin-based warhead that selectively modifies arginine residues. We developed alkyne-functionalized variants of ninhydrin to establish an arginine-specific chemical proteomics platform, enabling the classification of over 10,000 arginines. These studies uncovered potential modification sites on disease-relevant proteins. In one case, we identified a reactive arginine within a protein’s catalytic site, essential for its function. By endowing a reversible small molecule inhibitor with our ninhydrin warhead, we achieved selective, covalent engagement, highlighting the potential for targeting arginines in therapeutic development. These findings establish ninhydrin as a warhead for studying arginine reactivity and modulating protein function, enabling new avenues for covalently engaging proteins.
[0018] The foregoing summary is illustrative only and is not intended to be in any way limiting. In addition to the illustrative embodiments and features described herein, further aspects, embodiments, objects and features of the disclosure will become fully apparent from the drawings and the detailed description and the claims.BRIEF DESCRIPTION OF THE DRAWINGS
[0019] The features of the present disclosure are set forth with particularity in the appended claims. A better understanding of the features and advantages of the present disclosure will be obtained by reference to the following detailed description that sets forth illustrative embodiments, in which the principles of the disclosure are utilized, and the accompanying drawings.
[0020] FIG 1A involves profiling the reactivity of ninhydrin toward arginine in vitro; including a proposed reaction scheme between phenylglyoxal or ninhydrin and arginine, showing formation of stable adducts. Calculated equilibrium constants (Keq) were obtained by density functional theory (DFT) at the B3LYP / 6-311G** level with implicit water solvation.Atty. Docket: UCSF-865WO
[0021] FIG IB shows a kinetic analysis of pH- and time-dependent consumption of Fmoc-Arg (500 pM) by ninhydrin (15 mM, 30x) in PBS at 20 °C. Arginine consumption was quantified by LC-MS over 2 to 60 minutes.
[0022] FIG 1C shows a surface rendering of bovine serum albumin (BSA. PDB: 3V03) showing the distribution of 23 arginine residues, highlighting potential reactive sites.
[0023] FIG ID shows intact protein mass spectrometry analysis of BSA (1.5 pM) treated with 150 pM ninhydrin at 37 °C for 90 minutes across pH 7.5 to 11. Reactions were quenched by dilution into 0.1% aqueous TFA and analyzed by LC-MS on a Xevo G2-XS system.
[0024] FIG IE shows a heatmap depicting the second order reaction kinetics of ninhydrin or phenylglyoxal with Fmoc-arginine determined by UHPLC analysis compared to the reaction of iodoacetamide with cysteine determined by Ellman's reagent.
[0025] FIG IF shows intact protein mass spectrometry analysis of BSA (1.5 pM) treated with 150 pM ninhydrin at 37 °C for 30 minutes across pH 6.5 to 9.5. Reactions were quenched by dilution into 0.1% aqueous TFA and analyzed by LC-MS on a Agilent 6230 TOF System. Error bars represent SEM. n = 3.
[0026] FIG 1G shows a kinetic analysis of pH- and time-dependent consumption of Fmoc-Arg (500 pM) by PGO (15 mM, 30x) in PBS at 20 °C. Arginine consumption was quantified by LC-MS over 2 to 60 minutes. Error bars represent SD. n = 3
[0027] FIG 1H shows UpSet and waterfall plots for the closed search of the non-oxidized PGO-engaged and Nin-Alk engaged arginines detected by RAP. High overlap was observed between PGO and Nin-Alk-engaged arginines (n = 2,387), with the majority preferentially engaged with a log2 ratio > 1 by Nin-Alk (n = 3,249) relative to PGO (n = 31).
[0028] FIG II, 1J, and IK show a corresponding analysis for the closed search including the oxidized PGO species (+ox). Similar trends were observed (n = 3,306 total), with a higher preference for Nin-Alk (n = 3,187) relative to PGO (31).
[0029] FIG 2A shows a comparative profiling and biochemical characterization of argininereactive electrophiles, including chemical structures of electrophilic probes used in this study, including lA-Alk (iodoacetamide alkyne, cysteine-reactive control), phenylglyoxal (PGO), and two alkyne-functionalized ninhydrin derivatives, 4-Nin-Alk and 5-Nin-Alk.Atty. Docket: UCSF-865WO
[0030] FIG 2B shows qualitative, gel-based assessment of proteome-wide reactivity of alkyne-bearing covalent probes in Mino cell lysate. Soluble proteome (1.5 pg pL1) was treated with each probe at increasing concentrations (1, 10, or 100 pM) for 1 h, followed by CuAAC conjugation with TAMRA-N3, separation by SDS-PAGE, and in-gel fluorescence scanning.
[0031] FIG 2C shows qualitative time-dependent reactivity of 4-Nin-Alk, 5-Nin-Alk, and PGO in live cells. A549 cells were treated with 10 pM of indicated probe for 0, 30, or 60 min followed by lysis, CuAAC conjugation with TAMRA-N3, separation by SDS-PAGE, and in-gel fluorescence scanning.
[0032] FIG 2D shows a schematic of the RisoLFG workflow. Lysates are treated with 5-Nin-Alk or DMSO, followed by CuAAC conjugation to isotopically light or heavy isoDTB tags, sample combination (1:1), enrichment, digestion, and LC-MS / MS. Arginine residues with a heavy-to-light (H / L) ratio > 3 were classified as probe-reactive sites.
[0033] FIG 2E shows an open-mass search analysis of modified peptides reveals a predominant Amass of 701.3014 Da, corresponding to covalent modification by 5-Nin-Alk and isoDTB tag conjugation.
[0034] FIG 2F shows a glutathione consumption assay using Ellman’s reagent. Free glutathione (250 pM) was incubated with 500 pM propargyl PGO, propargyl 5-Nin-Alk, or DMSO over 60 min. Propargyl PGO induces substantial glutathione depletion, whereas 5-Nin-Alk shows minimal consumption. Data are presented as mean ± s.d. (n = 3).
[0035] FIG 2G shows a qualitative, gel-based assessment of proteome-wide reactivity of alkyne-bearing covalent probes in Mino cell lysate. Soluble proteome (1.5 pg pL1) was treated with each probe at the indicated concentration for 1 h, followed by CuAAC conjugation with TAMRA-N3, separation by SDS-PAGE, and in-gel fluorescence scanning.
[0036] FIG 2H shows a qualitative time-dependent reactivity of Nin-Alk, and PGO in live cells. Mino cells were treated with 100 pM of the indicated probe for 0, 30, or 60 min, followed by lysis, CuAAC conjugation with TAMRA-N3, separation by SDS-PAGE, and in-gel fluorescence scanning. Uncropped gel images available in SI.
[0037] FIG 21 shows a schematic of the Reactive Arginine Profiling (RAP) workflow to assess arginine reactivity. Cell lysates were treated with an arginine-reactive electrophile at high (1Atty. Docket: UCSF-865WOmM) or low (100 pM) concentrations and subjected to copper-catalyzed azide-alkyne cycloaddition (“click” chemistry) with isotopically light or heavy isoDTB tags. Samples were combined 1:1, enriched, digested, and analyzed by LC-MS / MS. Peptides from hyper-reactive arginine residues show similar heavy / light (H / L) intensities, whereas less reactive sites are preferentially labeled at higher probe concentrations, yielding elevated H / L ratios.
[0038] FIG 2J shows an open-mass search analysis of modified peptides reveals a predominant Amass of 701.308600 and 695.310118 Da, corresponding to covalent modification of arginine by Nin-Alk and heavy or light isoDTB tag conjugation, respectively. Error bar represents SD.
[0039] FIG 2K shows pLogo motif analysis of the +10-residue sequence window surrounding all modified arginines (foreground: 6,913 modified sites: background: 773,893 human arginines.
[0040] FIG 2L shows comparative profiling of phenylglyoxal (PGO) and Nin-Alk demonstrates increased saturation and promiscuity for Nin-Alk. UpSet plot demonstrating high overlap between PGO-engaged arginines and Nin-Alk-engaged arginines at 1 mM (1 h), with Nin-Alk engaging more targets.
[0041] FIG 2M shows waterfall plot showing median log? ratios of 1 mM Nin-Alk to 1 mM PGO isoDTB enrichment (n = 6). Peptides with log2(Nin-Alk / PGO) > 1 were classified as preferring Nin-Alk, whereas those with log2(Nin-Alk / PGO) < -1 preferred PGO. Because PGO generates a distinct oxidation side product, quantification was performed by merging closed searches corresponding to oxidized and unoxidized species; intensities for both light channels (± ox) were summed for each peptide, and the heavy channel was averaged across searches to prevent double-counting.
[0042] FIG 2N shows a reactivity analysis of pH- and time-dependent consumption of Fmoc-Arg (500 pM) by ninhydrin (Nin, 15 mM, 30x) or phenylglyoxal (PGO, 15 mM, 30x) in PBS at 20 °C. Arginine consumption was quantified by UHPLC analysis of reactions quenched at 15 or 60 minutes. Error bars represent SD n = 4.
[0043] FIG 3A shows quantitative profiling of arginine reactivity in the human proteome, including a schematic of the isoDTB-based chemoproteomic workflow to assess arginine reactivity. Cell lysates were treated with an arginine-reactive electrophile at high (lOx) or low (lx) concentrations, followed by copper-catalyzed azide-alkyne cycloaddition (“click”Atty. Docket: UCSF-865WOchemistry) to isotopically light or heavy isoDTB tags. Samples were combined 1:1, enriched, digested, and analyzed by LC-MS / MS. Peptides from hyper-reactive arginine residues show similar heavy / light (H / L) intensities, whereas less reactive sites are preferentially labeled at higher probe concentrations, yielding elevated H / L ratios.
[0044] FIG 3B shows quantified peptides containing reactive arginines based on log2enrichment ratios (1 mM / 100 pM). A threshold of log2 ratio < 1 defines hyper-reactive sites. A total of 10,642 arginines were quantified.
[0045] FIG 3C shows relative solvent accessibility of all reactive (purple) and hyper-reactive (green) arginine residues, indicating a shared preference for surface exposure.
[0046] FIG 3D shows a distribution of predicted pKavalues for reactive and hyper-reactive sites, indicating a shared distribution of pKa’s.
[0047] FIG 3E shows a structural and functional context of arginines identified by RAP in Mino cell lysates including a total number of unique arginines and corresponding proteins identified using the RAP workflow (see FIG 21).
[0048] FIG 3F shows a bubble plot of the top 10 most significantly enriched PANTHER protein classes. Statistical significance was calculated using Fisher’s Exact test and adjusted using the Benjamini-Hochberg procedure to control the false discovery rate at a = 0.05; only terms with adjusted p-value < 109are shown. Bubble size reflects the number of proteins in each class.
[0049] FIG 3G shows a target development level of proteins containing reactive arginines, based on Target Central Resource Database classification, where Tclin = Approved Drug, Tchem = Chemical Evidence, Tbio = Biological Evidence, Tdark = Understudied Proteins.
[0050] FIG 3H shows a number of proteins detected in our RAP dataset that are annotated as targeted in the DrugBank.
[0051] FIG 31 shows a distribution of secondary structure assignments for reactive arginines, as annotated using DSSP on PDB structures.
[0052] FIG 3J shows a histogram of predicted intrinsic disorder scores at modified arginine sites, calculated using IUPred2A. Scores > 0.5 indicate disordered regions, (n = 3630).Atty. Docket: UCSF-865WO
[0053] FIG 3K shows a quantification of cation-7t interactions between reactive arginines and aromatic residues in PDB deposited structures. Interaction counts are shown at distance thresholds of 8.0 or 5.5 A, measured from the guanidinium CZ atom to the centroid of the aromatic ring. Geometric constraints were applied for inclusion of parallel interactions (face-on) when the angle between the CZ-to-centroid vector and the normal vector of the aromatic plane was < 30°, and T-shaped interactions (edge-on) when deviation from 90° was < 20°.
[0054] FIG 3L shows a violin plot showing the distribution of salt bridge distances between guanidinium nitrogen atoms of reactive arginines and carboxylate oxygens of acidic residues (Asp / Glu) with discrepancies between PDB (n = 1594) and AlphaFold (n = 1445) models.Welch’s t-test (two-tailed, parametric).
[0055] FIG 3M shows a predicted relative solvent accessibility (RSA) of reactive arginines, computed using the Shrake-Rupley algorithm on PDB (n = 1649) and predicted AlphaFold (n = 4309) structures, indicating an over-representation of arginine surface burial in AlphaFold models compared to PDB models. Welch’s t-test (two-tailed, parametric). Inspired by the unique roles arginines play in protein structure and scaffolding, we sought to characterize the broader structural and functional context of Nin-Alk enriched residues. Through secondary-structure mapping across the PDB, we found reactive arginines are enriched in helices and more flexible loops compared to -sheets (FIG 31). To gauge local flexibility, we used IUPred2A. an energybased predictor of intrinsic disorder, to compute per-residue scores from 0 (ordered) to 1 (disordered). Reactive arginines peaked in the moderate-disorder regime (0.3 - 0.4) yet - 40% reside in well-ordered contexts (scores < 0.2), demonstrating that both structural support and some degree of flexibility facilitate efficient covalent engagement (FIG 3J).
[0056] FIG 3N shows biophysical and functional features between reactive arginine residues. Waterfall plot showing the distribution of median RAP enrichment scores (log2 heavy-to-light [H / L] ratios) for all quantified arginines across 4 biological replicates. Arginine residues with log2(H / L) <1 were classified as hyper-reactive (n = 410), representing a subset of the total reactive arginines identified (n = 6.888).
[0057] FIG 30 shows a distribution of reactive and hyper-reactive arginines per unique protein.Atty. Docket: UCSF-865WO
[0058] FIG 3P shows a distribution of predicted pKa values for hyper-reactive (n = 85) and all reactive arginine residues (n = 1570), calculated using PROPKA3 and the PDB. Welch’s t-test (two-tailed, parametric).
[0059] FIG 3Q shows a predicted relative solvent accessibility (RSA) of hyper-reactive (n = 93) and reactive arginines (n = 1649), computed using the Shrake-Rupley algorithm on PDB structures, indicating a shared preference for surface exposure.
[0060] FIG 3R shows a violin plot showing the distribution of salt bridge distances between guanidinium nitrogen atoms of hyper-reactive (n = 101) or all reactive arginines (n = 1445) and carboxylate oxygens of acidic residues (Asp / Glu) in PDB models. Welch’s t-test (two-tailed, parametric).
[0061] FIG 3S shows mean pathogenicity score of mutations to each hyper-reactive (n = 220) and reactive arginine (n = 4865), predicted using AlphaMissense.
[0062] FIG 3T shows live-cell RAP profiling in Mino cells, a. Waterfall plot showing the distribution of median RAP enrichment ratios (Iog2100 pM Nin-Alk / 10 pM Nin-Alk) obtained from live-cell labeling in Mino cells. Arginines with log2 > 1 were characterized as reactive (n = 80), while those with a log2 < 1 were designated hyper-reactive (n = 4). Representative proteins with their annotated subcellular localizations (UniProt) are highlighted, demonstrated labeling of arginines from distinct cellular compartments including the nucleus / cytoplasm (TCL1 A, PPIA), endoplasmic reticulum (PDIA3), and mitochondria (EFTU).
[0063] FIG 3U shows a waterfall plot showing distribution of median RAP enrichment ratios for lysate labeling of Mino cells at 100 pM and 10 pM of Nin-Alk. Despite the absence of cellular barriers and growth media, overall coverage remained modest (n = 60 reactive arginines), suggesting that the lower number of labeled sites observed in live-cell experiments is not solely attributable to probe permeability or intracellular accessibility.
[0064] FIG 3V shows a waterfall plot comparing distribution of RAP enrichment ratios for Nin-Alk (100 pM) and PGO (100 pM) labeling in live Mino cells. Peptides with log2 > 1 were classified as preferring Nin-Alk (n = 45), whereas those with log2 < -1 preferred PGO (n = 5). These results mirror the trends observed in lysate.Atty. Docket: UCSF-865WO
[0065] FIG 3W shows an UpSet plot showing overlap between arginine-modified peptides detected in lysate and live-cell experiments (n = 116 total). 28 peptides were common to both datasets, suggesting potential context-dependent factors driving arginine accessibility and reactivity.
[0066] FIG 4A shows a structure-guided development of arginine- selective ligands, including a crystal structure of mitochondrial aconitase (ACO2) bound to isocitrate and a [4Fe-4S] cluster, highlighting the proximity of R479 to the active site (PDB: 1B0J).
[0067] FIG 4B shows covalent docking model of 5-Nin-Alk to R479 of ACO2, showing spatial alignment with the native isocitrate ligand.
[0068] FIG 4C shows a schematic of the in-house structure-based filtering pipeline used to identify candidate ligands in proximity to modified arginines. Input UniProt IDs and modified residue positions were filtered by PDB alignment, ligand proximity (<3 A), and drug-likeness of the ligand to prioritize hits for follow-up.
[0069] FIG 4D shows chemical structures of the tri- vector ligand and tri- vector alkyne, a clickable analog previously shown to bind cyclophilin A (PP1A). The tri-vector alkyne contains an alkyne moiety that enabled CuA AC-based enrichment and inspired the synthesis of 5-CA-Alk and 4-CA-Alk.
[0070] FIG 4E shows published co-crystal structure of PPIA bound to the tri-vector ligand (PDB: 6GJI), showing R55 positioned at the periphery of the ligand binding site.
[0071] FIG 4F shows covalent docking model of 4-CA-Alk to R55 of PPIA, showing similar binding orientation as the tri-vector ligand.
[0072] FIG 4G shows chemoproteomic enrichment of peptides labeled by 4-CA-Alk (10 pM, Ih) in a RisoLFG experiment. Each data point represents a distinct modified peptide; covalent labeling of R55 in cyclophilin A (CYPA) is highlighted in purple.
[0073] FIG 4H shows relative enrichment of CYPA R55 between 4-CA-Alk and 5 -C A- Aik as determined by RisoLFG. n = 4 replicates.
[0074] FIG 41 shows CypA activity measured by a coupled peptidyl-prolyl isomerase assay. Pre-incubation of 1 pM recombinant CypA with CA-Nin-Alk (4-CA-Alk) (10 pM, 25 pM; 1 h)Atty. Docket: UCSF-865WOdecreases CypA catalytic efficiency (kcat / Km) by 20.3 ± 8.1% and 56.4 + 4.8% , respectively compared to DMSO. Data represent mean ± s.d. (n = 3).
[0075] FIG 4J shows a chemoproteomic enrichment of peptides labeled by CA-Nin-Alk (4-CA-Alk) (10 pM, Ih) in a RAP experiment. Each data point represents a distinct modified peptide; covalent labeling of R55 in cyclophilin A (CypA) is highlighted in blue.
[0076] FIG 5 shows the relative frequency percentage of all reactives and hyper reactives by active site distance.DETAILED DESCRIPTION OF THE INVENTIONI. Definitions
[0077] The symbol r xrun?whether utilized as a bond or displayed perpendicular to a bond, indicates the point at which the displayed moiety is attached to the remainder of the compound.
[0078] As used herein, the term “alkyl” by itself or as part of another substituent refers to a saturated branched or straight-chain monovalent hydrocarbon radical derived by the removal of one hydrogen atom from a single carbon atom of a parent alkane. Typical alkyl groups include, but are not limited to, methyl: ethyl, propyls such as propan- 1-yl or propan-2-yl; and butyls such as butan-l-yl, butan-2-yl, 2-methyl-propan-l-yl or 2-methyl-propan-2-yl. In some embodiments, an alkyl group comprises from 1 to 20 carbon atoms. In other embodiments, an alkyl group comprises from 1 to 10 carbon atoms. In still other embodiments, an alkyl group comprises from 1 to 6 carbon atoms, such as from 1 to 4 carbon atoms, such as from 1 to 3 carbon atoms.
[0079] "Alkanyl" by itself or as part of another substituent refers to a saturated branched, straight-chain or cyclic alkyl radical derived by the removal of one hydrogen atom from a single carbon atom of an alkane. Typical alkanyl groups include, but are not limited to, methanyl; ethanyl; propanyls such as propan-l-yl, propan-2-yl (isopropyl), cyclopropan-l-yl, etc.; butanyls such as butan-l-yl, butan-2-yl (sec-butyl), 2-methyl-propan-l-yl (isobutyl), 2-methyl-propan-2-yl (t-butyl), cyclobutan-l-yl, etc.; and the like.
[0080] "Alkylene" refers to a branched or unbranched saturated hydrocarbon chain, usually having from 1 to 20 carbon atoms, more usually 1 to 10 carbon atoms and even more usually 1 to 6 carbon atoms, such as from 1 to 4 carbon atoms, such as from 1 to 3 carbon atoms. This term isAtty. Docket: UCSF-865WOexemplified by groups such as methylene (-CH2-), ethylene (-CH2CH2-), the propylene isomers (e.g„ -CH2CH2CH2- and -CH(CH3)CH2-) and the like.
[0081] "Alkenyl" by itself or as part of another substituent refers to an unsaturated branched, straight-chain or cyclic alkyl radical having at least one carbon-carbon double bond derived by the removal of one hydrogen atom from a single carbon atom of an alkene. The group may be in either the cis or trans conformation about the double bond(s). Typical alkenyl groups include, but are not limited to, ethenyl; propenyls such as prop-l-en-l-yl, prop-l-en-2-yl, prop-2-en-l-yl (allyl), prop-2-en-2-yl, cycloprop- 1-en-l-yl; cycloprop-2-en-l-yl; butenyls such as but-l-en-l-yl, but-l-en-2-yl, 2-methyl-prop- 1-en-l-yl, but-2-en-l-yl, but-2-en-l-yl, but-2-en-2-yl, buta-1,3-dien-l-yl, buta-l,3-dien-2-yl, cyclobut- 1-en-l-yl, cyclobut-l-en-3-yl, cyclobuta-l,3-dien-l-yl, etc.; and the like.
[0082] "Alkynyl" by itself or as part of another substituent refers to an unsaturated branched, straight-chain or cyclic alkyl radical having at least one carbon-carbon triple bond derived by the removal of one hydrogen atom from a single carbon atom of an alkyne. Typical alkynyl groups include, but are not limited to, ethynyl; propynyls such as prop-l-yn-l-yl, prop-2-yn-l-yl, etc.; butynyls such as but-l-yn-l-yl, but-l-yn-3-yl, but-3-yn-l-yl, etc.; and the like.
[0083] "Acyl" by itself or as part of another substituent refers to a radical -C(O)R30, where R30is hydrogen, alkyl, cycloalkyl, heterocycloalkyl, aryl, arylalkyl, heteroalkyl, heteroaryl, heteroarylalkyl as defined herein and substituted versions thereof. Representative examples include, but are not limited to formyl, acetyl, cyclohexylcarbonyl, cyclohexylmethylcarbonyl, benzoyl, benzylcarbonyl, piperonyl, propionyl, succinyl, and malonyl, and the like.
[0084] The term "aminoacyl" refers to the group -C(O)NR21R22, wherein R21and R22independently are selected from the group consisting of hydrogen, alkyl, substituted alkyl, alkenyl, substituted alkenyl, alkynyl, substituted alkynyl, aryl, substituted aryl, cycloalkyl, substituted cycloalkyl, cycloalkenyl, substituted cycloalkenyl, heteroaryl, substituted heteroaryl, heterocyclic, and substituted heterocyclic and where R21and R22are optionally joined together with the nitrogen bound thereto to form a heterocyclic or substituted heterocyclic group, and wherein alkyl, substituted alkyl, alkenyl, substituted alkenyl, alkynyl, substituted alkynyl, cycloalkyl, substituted cycloalkyl, cycloalkenyl, substituted cycloalkenyl, aryl, substituted aryl,Atty. Docket: UCSF-865WOheteroaryl, substituted heteroaryl, heterocyclic, and substituted heterocyclic are as defined herein.
[0085] "Alkoxy" by itself or as part of another substituent refers to a radical -OR31where R31represents an alkyl group as defined herein. Representative examples include, but are not limited to, methoxy, ethoxy, propoxy, butoxy, and the like.
[0086] "Cycloalkoxy" by itself or as part of another substituent refers to a radical -OR31where R31represents a cycloalkyl group as defined herein. Representative examples include, but are not limited to, cyclopropoxy, cyclobutoxy, cyclohexyloxy and the like.
[0087] "Alkoxycarbonyl" by itself or as part of another substituent refers to a radical -C(O)OR31where R31represents an alkyl or cycloalkyl group as defined herein. Representative examples include, but are not limited to, methoxycarbonyl, ethoxycarbonyl, propoxycarbonyl, butoxycarbonyl, cyclohexyloxycarbonyl and the like.
[0088] "Aryl" by itself or as part of another substituent refers to a monovalent aromatic hydrocarbon radical derived by the removal of one hydrogen atom from a single carbon atom of an aromatic ring system. Typical aryl groups include, but are not limited to, groups derived from aceanthrylene, acenaphthylene, acephenanthrylene, anthracene, azulene, benzene, chrysene, coronene, fluoranthene, fluorene, hexacene, hexaphene, hexalene, as-indacene. s-indacene, indane, indene, naphthalene, octacene, octaphene, octalene, ovalene, penta-2,4-diene, pentacene, pentalene, pentaphene, perylene, phenalene, phenanthrene, picene, pleiadene, pyrene, pyranthrene, rubicene. triphenylene, trinaphthalene and the like. In certain embodiments, an aryl group comprises from 6 to 20 carbon atoms. In certain embodiments, an aryl group comprises from 6 to 12 carbon atoms. Examples of an aryl group are phenyl and naphthyl.
[0089] "Arylalkyl" by itself or as part of another substituent refers to an acyclic alkyl radical in which one of the hydrogen atoms bonded to a carbon atom, typically a terminal or sp3carbon atom, is replaced with an aryl group. Typical arylalkyl groups include, but are not limited to, benzyl, 2-phenylethan-l-yl, 2-phenylethen-l-yl, naphthylmethyl, 2-naphthylethan-l-yl, 2-naphthylethen-l-yl, naphthobenzyl, 2-naphthophenylethan-l-yl and the like. Where specific alkyl moieties are intended, the nomenclature arylalkanyl, arylalkenyl and / or arylalkynyl is used. In certain embodiments, an arylalkyl group is (C7-C30) arylalkyl, e.g., the alkanyl, alkenyl or alkynyl moiety of the arylalkyl group is (C1-C10) and the aryl moiety is (C6-C20). In certainAtty. Docket: UCSF-865WOembodiments, an arylalkyl group is (C7-C20) arylalkyl, e.g., the alkanyl, alkenyl or alkynyl moiety of the arylalkyl group is (Ci-Cs) and the aryl moiety is (C6-C12).
[0090] "Arylaryl" by itself or as part of another substituent, refers to a monovalent hydrocarbon group derived by the removal of one hydrogen atom from a single carbon atom of a ring system in which two or more identical or non-identical aromatic ring systems are joined directly together by a single bond, where the number of such direct ring junctions is one less than the number of aromatic ring systems involved. Typical arylaryl groups include, but are not limited to, biphenyl, triphenyl, phenyl-napthyl, binaphthyl, biphenyl-napthyl, and the like. When the number of carbon atoms in an arylaryl group are specified, the numbers refer to the carbon atoms comprising each aromatic ring. For example, (C5-C14) arylaryl is an arylaryl group in which each aromatic ring comprises from 5 to 14 carbons, e.g., biphenyl, triphenyl, binaphthyl, phenylnapthyl, etc. In certain embodiments, each aromatic ring system of an arylaryl group is independently a (C5-C14) aromatic. In certain embodiments, each aromatic ring system of an arylaryl group is independently a (C5-C10) aromatic. In certain embodiments, each aromatic ring system is identical, e.g.. biphenyl, triphenyl, binaphthyl, trinaphthyl, etc.
[0091] "Cycloalkyl" by itself or as part of another substituent refers to a saturated or unsaturated cyclic alkyl radical derived by the removal of one hydrogen atom from a single carbon atom of the parent. Where a specific level of saturation is intended, the nomenclature "cycloalkanyl" or "cycloalkenyl" is used. Typical cycloalkyl groups include, but are not limited to, groups derived from cyclopropane, cyclobutane, cyclopentane, cyclohexane and the like. In certain embodiments, the cycloalkyl group is (C3-C10) cycloalkyl. In certain embodiments, the cycloalkyl group is (Cb-Cg) cycloalkyl.
[0092] "Heterocycloalkyl" or "heterocyclyl" by itself or as part of another substituent, refers to a saturated or unsaturated cyclic alkyl radical in which one or more carbon atoms (and any associated hydrogen atoms) are independently replaced with the same or different heteroatom. Typical heteroatoms to replace the carbon atom(s) include, but are not limited to, N, P, O, S, Si, etc. Where a specific level of saturation is intended, the nomenclature "heterocycloalkanyl" or "heterocyclo alkenyl" is used. Typical heterocycloalkyl groups include, but are not limited to, groups derived from epoxides, azirines, thiiranes, imidazolidine, morpholine, piperazine, piperidine, pyrazolidine, pyrrolidine, quinuclidine and the like.Atty. Docket: UCSF-865WO
[0093] "Heteroalkyl, Heteroal kanyl, Heteroalkenyl and Heteroal kynyl" by themselves or as part of another substituent refer to alkyl, alkanyl, alkenyl and alkynyl groups, respectively, in which one or more of the carbon atoms (and any associated hydrogen atoms) are independently replaced with the same or different heteroatomic groups. Typical heteroatomic groups which can be included in these groups include, but are not limited to, -O-, -S-, -S-S-, -O-S-, -NR37R38-, =N-N=, -N=N-, -N=N-NR39R40, -PR41-, -P(O)2-, -POR42-, -O-P(O)2-, -S-O-, -S-(O)-. -SO2-. -SnR43R44- and the like, where R37, R38, R39, R40, R41, R42, R43and R44are independently hydrogen, alkyl, substituted alkyl, aryl, substituted aryl, arylalkyl, substituted arylalkyl, cycloalkyl, substituted cycloalkyl, heterocycloalkyl, substituted heterocycloalkyl, heteroalkyl, substituted heteroalkyl, heteroaryl, substituted heteroaryl, heteroarylalkyl or substituted heteroarylalkyl.
[0094] "Heteroaryl" by itself or as part of another substituent, refers to a monovalent heteroaromatic radical derived by the removal of one hydrogen atom from a single atom of a heteroaromatic ring system. Typical heteroaryl groups include, but are not limited to, groups derived from acridine, arsindole, carbazole, P-carboline, chromane, chromene, cinnoline, furan, imidazole, indazole, indole, indoline, indolizine, isobenzofuran, isochromene, isoindole, isoindoline, isoquinoline, isothiazole, isoxazole, naphthyridine, oxadiazole, oxazole, perimidine, phenanthridine, phenanthroline, phenazine, phthalazine, pteridine, purine, pyran, pyrazine, pyrazole, pyridazine, pyridine, pyrimidine, pyrrole, pyrrolizine, quinazoline, quinoline, quinolizine, quinoxaline, tetrazole, thiadiazole, thiazole, thiophene, triazole, xanthene, benzodioxole and the like. In certain embodiments, the heteroaryl group is from 5-20 membered heteroaryl. In certain embodiments, the heteroaryl group is from 5-10 membered heteroaryl. In certain embodiments, heteroaryl groups are those derived from thiophene, pyrrole, benzothiophene, benzofuran, indole, pyridine, quinoline, imidazole, oxazole and pyrazine.
[0095] "Heteroarylalkyl" by itself or as part of another substituent, refers to an acyclic alkyl radical in which one of the hydrogen atoms bonded to a carbon atom, typically a terminal or sp3carbon atom, is replaced with a heteroaryl group. Where specific alkyl moieties are intended, the nomenclature heteroarylalkanyl, heteroarylalkenyl and / or heterorylalkynyl is used. In certain embodiments, the hetero arylalkyl group is a 6-30 membered heteroarylalkyl, e.g., the alkanyl, alkenyl or alkynyl moiety of the hetero arylalkyl is 1-10 membered and the heteroaryl moiety is a 5-20-membered heteroaryl. In certain embodiments, the heteroarylalkyl group is 6-20 memberedAtty. Docket: UCSF-865WOheteroarylalkyl, e.g., the alkanyl, alkenyl or alkynyl moiety of the heteroaryl alkyl is 1-8 membered and the heteroaryl moiety is a 5-12-membered heteroaryl.
[0096] "Aromatic Ring System" by itself or as part of another substituent, refers to an unsaturated cyclic or polycyclic ring system having a conjugated it electron system. Specifically included within the definition of "aromatic ring system" are fused ring systems in which one or more of the rings are aromatic and one or more of the rings are saturated or unsaturated, such as, for example, fluorene, indane, indene, phenalene, etc. Typical aromatic ring systems include, but are not limited to, aceanthrylene, acenaphthylene, acephenanthrylene, anthracene, azulene, benzene, chrysene, coronene, fluoranthene, fluorene, hexacene, hexaphene, hexalene, as-indacene, s-indacene, indane, indene, naphthalene, octacene, octaphene, octalene, ovalene, penta-2,4-diene, pentacene, pentalene, pentaphene, perylene, phenalene, phenanthrene, picene, pleiadene, pyrene, pyranthrene. rubicene, triphenylene, trinaphthalene and the like.
[0097] "Heteroaromatic Ring System" by itself or as part of another substituent, refers to an aromatic ring system in which one or more carbon atoms (and any associated hydrogen atoms) are independently replaced with the same or different heteroatom. Typical heteroatoms to replace the carbon atoms include, but are not limited to, N, P, O, S, Si, etc. Specifically included within the definition of "heteroaromatic ring systems" are fused ring systems in which one or more of the rings are aromatic and one or more of the rings are saturated or unsaturated, such as. for example, arsindole, benzodioxan, benzofuran, chromane, chromene, indole, indoline, xanthene, etc. Typical heteroaromatic ring systems include, but are not limited to, arsindole, carbazole, 0-carboline, chromane, chromene, cinnoline, furan, imidazole, indazole, indole, indoline, indolizine, isobenzofuran, isochromene, isoindole, isoindoline, isoquinoline, isothiazole, isoxazole, naphthyridine, oxadiazole, oxazole, perimidine, phenanthridine, phenanthroline, phenazine, phthalazine, pteridine, purine, pyran, pyrazine, pyrazole. pyridazine, pyridine, pyrimidine, pyrrole, pyrrolizine, quinazoline, quinoline, quinolizine, quinoxaline, tetrazole, thiadiazole, thiazole, thiophene, triazole, xanthene and the like.
[0098] “Substituted” refers to a group in which one or more hydrogen atoms are independently replaced with the same or different substituent(s). Typical substituents include, but are not limited to, alkylenedioxy (such as methylenedioxy), -M, -R60, -O’, =0, -OR60, -SR60, -S’, =S, -NR60R61, =NR60. -CF3, -CN, -OCN, -SCN, -NO, -NO2. =N2, -N3, -S(O)2O . -S(O)2OH, -Atty. Docket: UCSF-865WOS(O)2R60, -0S(0)20', -OS(O)2R60, -P(0)(0')2, -P(O)(OR60)(O ), -OP(O)(OR60)(OR61), -C(O)R60, -C(S)R60, -C(O)OR60, -C(O)NR60R61, -C(0)0', -C(S)OR60, -NR62C(O)NR60R61, -NR62C(S)NR60R61. -NR62C(NR63)NR60R61and -C(NR62)NR60R61where M is halogen: R60, R61, R62and R63are independently hydrogen, alkyl, substituted alkyl, alkoxy, substituted alkoxy, cycloalkyl, substituted cycloalkyl, heterocycloalkyl, substituted heterocycloalkyl, aryl, substituted aryl, heteroaryl or substituted heteroaryl, or optionally R60and R61together with the nitrogen atom to which they are bonded form a heterocycloalkyl or substituted heterocycloalkyl ring; and R64and R65are independently hydrogen, alkyl, substituted alkyl, aryl, cycloalkyl, substituted cycloalkyl, heterocycloalkyl, substituted heterocycloalkyl, aryl, substituted aryl, heteroaryl or substituted heteroaryl, or optionally R64and R65together with the nitrogen atom to which they are bonded form a heterocycloalkyl or substituted heterocycloalkyl ring. In certain embodiments, substituents include -M, -R60, =0, -OR60, -SR60, -S', =S, -NR60R61, =R60, -CF3, -CN, -OCN, -SCN, -NO, -N02, =N2, -N3, -S(O)2R60. -0S(0)20', -OS(O)2R60, -P(0)(0')2, -P(O)(OR60)(O ), -OP(O)(OR60)(OR61). -C(O)R60, -C(S)R60, -C(O)OR60, -C(O)NR60R61. -C(0)0-, -NR62C(O)NR60R61. In certain embodiments, substituents include -M, -R60, =0, -OR60, -SR60, -NR60R61, -CF3, -CN, -N02, -S(O)2R60, -P(0)(OR60)(0 ), -OP(O)(OR60)(OR61), -C(O)R60, -C(O)OR60, -C(O)NR60R61, -C(0)0 . In certain embodiments, substituents include -M, -R60, =0, -OR60, -SR60, -NR60R61, -CF3, -CN, -N02, -S(O)2R60, -OP(O)(OR60)(OR61), -C(O)R60, -C(O)OR60, -C(0)0', where R60, R61and R62are as defined above. For example, a substituted group may bear a methylenedioxy substituent or one, two, or three substituents selected from a halogen atom, a (Ci-4)alkyl group and a (Ci-4)alkoxy group.
[0099] “Amino" refers to the group -NRXRYwherein Rxand RYare each independently H or a non-hydrogen substituent. Exemplary non-hydrogen substituents include alkyl groups (e.g. methyl, ethyl, and isopropyl).
[0100] “Ether” refers to a diradical group of formula -O-. For instance, if the ether group is connected to an alkyl group, then the overall group is an alkoxy group (e.g. -OCH3or methoxy). If the ether is connected to a carbonyl group, then the overall group is an ester group of formula -OC(O)-.
[0101] “Halo” and “halogen” refer to the chloro, bromo, fluoro, and iodo groups.
[0102] “Nitro” refers to the group of formula -NO2.Atty. Docket: UCSF-865WO
[0103] The terms "patient," "subject," and "human subject" are used interchangeably herein.
[0104] As to any of the groups disclosed herein which contain one or more substituents, it is understood, of course, that such groups do not contain any substitution or substitution patterns which are sterically impractical and / or synthetically non-feasible. In addition, the subject compounds include all stereochemical isomers arising from the substitution of these compounds.
[0105] In certain embodiments, a substituent may contribute to optical isomerism and / or stereo isomerism of a compound. Salts, solvates, hydrates, and prodrug forms of a compound are also of interest. All such forms are embraced by the present disclosure. Thus, the compounds described herein include salts, solvates, hydrates, prodrug and isomer forms thereof, including the pharmaceutically acceptable salts, solvates, hydrates, prodrugs and isomers thereof. In certain embodiments, a compound may be metabolized into a pharmaceutically active derivative.II. Introduction
[0106] Ninhydrin, a tri-keto indane, has long been utilized for detecting amino acids through its chromogenic reaction with primary amines. Despite its established role in analytical chemistry, its potential as a selective covalent modifier of arginine has remained underexplored. Given its electrophilic properties and ability to form stabilized cyclic adducts with guanidine, we hypothesized that ninhydrin could be repurposed for targeted covalent modification of arginine in biological systems. Furthermore, we postulated that if ninhydrin were reactive enough across proteinacious arginine residues, further functionalization with a bioorthogonal handle, such as an alkyne, would allow for the development of a pan-arginine chemical proteomics platform akin to the iodoacetamide-alkyne probe and the isoTOPP-ABPP workflow, as developed by Cravatt and co-workers for cysteine.
[0107] Small molecules serve as versatile probes for perturbing the functions of proteins in biological systems. In some instances, a plurality of human proteins lack selective chemical ligands. In some cases, several classes of proteins are further considered as undruggable.Covalent ligands offer a strategy to expand the landscape of proteins amenable to targeting by small molecules. In some instances, covalent ligands combine features of recognition and reactivity, thereby enabling targeting sites on proteins that are difficult to address by reversible binding interactions alone.Atty. Docket: UCSF-865WO
[0108] Described herein are small molecule probes that interact with a reactive arginine residue of an arginine-containing protein and methods of identifying a protein that contains such a reactive arginine residue (e.g., a druggable arginine residue). In some instances, also described herein are methods of profiling a ligand that interacts with one or more arginine-containing proteins comprising reactive arginines.
[0109] Described herein are modified arginine-containing proteins that are formed by reaction of an arginine-containing protein with one or more probes, ligands, ligand-electrophiles, or other moiety comprising a chemical group capable of reacting with an arginine residue. Further described herein are modified-arginine-containing proteins covalently attached to a small molecule fragment moiety via a moiety created with the side chain of an arginine residue, such as an imidazolidinimine moiety. Further described herein are kits for generating modified arginine-containing proteins.III. Small Molecule Probe Compounds
[0110] In an exemplary embodiment, the small molecule probe compound described herein comprises a reactive moiety which interacts with the guanidinium group of an arginine residue of an arginine containing protein. In an exemplary embodiment, small molecule probes react with arginine residues to form covalent bonds. Often, small molecule probes are non-naturally occurring, or form non-naturally occurring products after reaction with the guanidinium group of an arginine residue of an arginine containing protein. In an exemplary embodiment, the guanidinium group of the arginine containing protein is connected to a small molecule fragment moiety via a moiety, such as an imidazolidinimine moiety, after reaction with a small molecule probe.
[0111] In an exemplary embodiment, a small molecule probe compound described herein is a small molecule compound that has a structure represented by Formula (I): F1— Q (I), wherein F1is a small molecule fragment moiety comprising an azide, an alkyne, a fluorophore, a labeling group, or a combination thereof; and Q comprises a cyclic trione moiety, a cyclic dihydroxy dione moiety, a cyclic dialkoxy dione moiety, or a cyclic hydroxy alkoxy dione moiety. In an exemplary embodiment, the Q is ninhydrin. In an exemplary embodiment the small moleculeAtty. Docket: UCSF-865WO oprobe compound has a structure represented by:0(lb),oO (Id), wherein Ra, when present, is unsubstituted C1-C4 alkyl and Rb, when present, is unsubstituted C1-C4 alkyl. In an exemplary embodiment the small molecule probe compound has a structure represented by:F1Oo (lid). In an exemplary embodiment the small molecule probe compound has a structure represented by:when present, is unsubstituted C1-C4 alkyl and Rb, when present, is unsubstituted C1-C4 alkyl.
[0112] In an exemplary embodiment the small molecule probe compound has a structureo orepresented by:o (IVa), o (ivb),Atty. Docket: UCSF-865WOo oORaORaOH ORbO (IVc), or O (IVd), Raand Rb, when present, X, and Y are as described herein. In an exemplary embodiment, the small molecule probe compound has a structure represented by Formula (IVa), (IVb), (IVc), or (IVd), wherein X is -O- or -C(O)NH-. In an exemplary embodiment, the small molecule probe compound has a structure represented by Formula (IVa), (IVb), (IVc), or (IVd). wherein Y is methylene or ethylene. In an exemplary embodiment, the small molecule probe compound has a structure represented by Formula (IVa), (IVb), (IVc), or (IVd), wherein X is -O- or -C(O)NH-, and Y is methylene or ethylene.
[0113] In an exemplary embodiment the small molecule probe compound has a structure represented by Formula (la), (lb), (Ic), or (Id), wherein F1comprises a moiety with a stereocenter, and Raand Rb, when present, are as described herein. In an exemplary embodiment the small molecule probe compound has a structure represented by formula (la), (lb), (Ic), or (Id), wherein the moiety with a stereocenter is a carbon with a (R) orientation, or a (S) orientation. In an exemplary embodiment the small molecule probe compound has a structurewhen present, X, and Y are as described herein, and Xqcomprises a moiety with a stereocenter, such as a carbon with a (R) orientation, or a (S) orientation. In an exemplary embodiment the small molecule probe compound has a structure represented by:(VIb),Atty. Docket: UCSF-865WO(Vid), wherein Raand Rb, when present, X, and Y are as described herein, and A, including the N to which it is attached, is a 4 to 9 membered ring, and the bond between X and A is in a (R) orientation, or in a (S) orientation. In an exemplary embodiment the small molecule probe compound has a structure represented by:and Rb, when present, X, and Y are as described herein, and A, including the N to which it is attached, is a 4 to 9 membered ring. In an exemplary embodiment the small molecule probe compound has a structure represented by:and Rb. when present, X, and Y are as described herein, and A, including the N to which it is attached, is a 4 to 9 membered ring. For any of the exemplary embodiments of this paragraph, A is a 5 membered ring. For any of the exemplary embodiments of this paragraph, A is a 6 membered ring.Atty. Docket: UCSF-865WO
[0114] In an exemplary embodiment, the small molecule probe compound has a structurePH OH(IXb),herein Z is azide, alkyne, fluorophore, labeling group, or a combination thereof; Y is absent or substituted or unsubstituted Ci-C6alkylene; X is -CH2-, O, -C(O)NH, -NH-, -OC(O)O-, triazole, sulfonate, and phosphonate; and Raand Rbare independently selected from unsubstituted methyl, ethyl, propyl, isopropyl, butyl, isobutyl, and sec-butyl. In an exemplary embodiment, the small molecule probe compound has a structure according to formula (IXa), (IXb), (IXc), (IXd), wherein Z is a fluorophore; Y is absent or substituted or unsubstituted Ci-Ce alkylene; X is triazole; and Raand Rhare independently selected from unsubstituted methyl, ethyl, propyl, isopropyl, butyl, isobutyl, and sec-butyl.
[0115] In an exemplary embodiment, the fluorophore comprises rhodamine, rhodol, fluorescein, thiofluorescein, aminofluorescein, carboxyfluorescein, chlorofluorescein, methylfluorescein, sulfofluorescein, aminorhodol, carboxyrhodol, chlororhodol, methylrhodol, sulforhodol, aminorhodamine, carboxyrhodamine, chlororhodamine, methylrhodamine, sulforhodamine, thiorhodamine, cyanine, indocarbocyanine, oxacarbocyanine, thiacarbocyanine, merocyanine, cyanine 2, cyanine 3, cyanine 3.5, cyanine 5, cyanine 5.5, cyanine 7, oxadiazole derivatives, pyridyloxazole, nitrobenzoxadiazole, benzoxadiazole, pyren derivatives, cascade blue, oxazine derivatives, Nile red, Nile blue, cresyl violet, oxazine 170, acridine derivatives, proflavin, acridine orange, acridine yellow, arylmethine derivatives, auramine, crystal violet, malachite green, tetrapyrrole derivatives, porphin, phthalocyanine, bilirubin l-dimethylaminonaphthyl-5-sulfonate, l-anilino-8-naphthalene sulfonate, 2-p-toluidinyl-6-naphthalene sulfonate, 3-phenyl-7-isocyanatocoumarin, N-(p-(2-benzoxazolyl)phenyl)maleimide, stilbenes, pyrenes, 6-FAM (Fluorescein), 6-FAM (NHS Ester), 5(6)-FAM, 5-FAM, Fluorescein dT, 5-TAMRA-cadavarine, 2-aminoacridone, HEX, JOE (NHS Ester), MAX, TET, ROX, TAMRA, TARMA™ (NHSAtty. Docket: UCSF-865WOEster), TEX 615, ATTO™ 488, ATTO™ 532, ATTO™ 550, ATTO™ 565, ATTO™ RholOl, ATTO™ 590, ATTO™ 633, ATTO™ 647N, TYE™ 563, TYE™ 665, or TYE™ 705.
[0116] In an exemplary embodiment, F1comprises a fluorophore moiety. In some cases, F1is obtained from a compound library. In some cases, the compound library comprises ChemBridge fragment library, Pyramid Platform Fragment-Based Drug Discovery, Maybridge fragment library, FRGx from AnalytiCon, TCI-Frag from AnCoreX, Bio Building Blocks from ASINEX, BioFocus 3D from Charles River, and / or Fragments of Life (FOL) from Emerald Bio.
[0117] In an exemplary embodiment, the labeling group is biotin moiety, streptavidin moiety, bead, resin, a solid support, or a combination thereof.IV. Ligand
[0118] In an exemplary embodiment, a ligand competes with a probe compound described herein for binding with a reactive arginine residue. In an exemplary embodiment, a ligand comprises a small molecule compound, a polynucleotide, a polypeptide or its fragments thereof, or a peptidomimetic. In an exemplary embodiment, the ligand comprises a small molecule compound. In an exemplary embodiment, a small molecule compound comprises a fragment moiety that facilitates interaction of the compound with a reactive arginine residue. In an exemplary embodiment, a small molecule compound comprises a small molecule fragment that facilitates hydrophobic interaction, hydrogen bonding, or a combination thereof. Often, ligands are non-naturally occurring, or form non-naturally occurring products after reaction with the guanidinium group of an arginine residue of an arginine containing protein. In an exemplary embodiment, a ligand comprises a small-molecule compound. In an exemplary embodiment, a small molecule compound comprises a ligand-electrophile. Such ligand-electrophiles often react with the guanidinium group of an arginine residue of an arginine-containing protein.
[0119] In an exemplary embodiment, a ligand comprises a polynucleotide. In an exemplary embodiment, the polynucleotide comprises an endogenous substrate that interacts with an arginine-containing protein. In an exemplary embodiment, the polynucleotide comprises modified and / or synthetic substrate. In an exemplary embodiment, the polynucleotide comprises natural nucleotides. In an exemplary embodiment, the polynucleotide comprises artificial nucleotides.Atty. Docket: UCSF-865WO
[0120] In an exemplary embodiment, a polynucleotide comprises from about 8 to about 50 bases in length. In some cases, a polynucleotide comprises from about 12 to about 45, from about 15 to about 40, from about 20 to about 40, or from about 25 to about 300 bases in length. In an exemplary embodiment, a polynucleotide comprises 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 30, 35, 40, 45, or 50 bases in length.
[0121] In an exemplary embodiment, a ligand comprises a polypeptide or its fragments thereof. In an exemplary embodiment, the polypeptide comprises a wild-type functional protein, protein variants, or mutants that are substrates for an arginine-containing protein of interest. In an exemplary embodiment, fragments of the polypeptide comprise truncated functional proteins that interact with the arginine-containing protein of interest.
[0122] In an exemplary embodiment, a functional fragment of a polypeptide comprises from about 10 to about 80 amino acid residues in length. In an exemplary embodiment, the functional fragment comprises from about 15 to about 70, from about 20 to about 60, from about 30 to about 50, or from about 40 to about 80 amino acid residues in length. In an exemplary embodiment, the functional fragment comprises about 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60.70, 80, or more amino acid residues in length.
[0123] In an exemplary embodiment, a polypeptide or its fragments thereof comprise natural amino acids, unnatural amino acids, or a combination thereof. In an exemplary embodiment, the polypeptide or its fragments thereof comprise L-amino acids, D-amino acids, or a combination thereof.
[0124] In an exemplary embodiment, a ligand comprises a peptidomimetic. Peptidomimetic is a small protein-like chain that mimics a peptide. Exemplary peptidomimetic s include, but are not limited to, peptoids, P-peptides, or foldamers. Peptoids, also known as poly-N-substituted glycines, are a class of peptidomimetics in which the side chains are appended to the nitrogen atom of the peptide backbone instead of the a-carbon. P-peptides are P-amino acids in which the guanidinium groups are bonded to the P-carbon rather than the a-carbon. A foldamer is a discrete chain molecule or oligomer that folds into an ordered conformation such as helices and P-sheets.
[0125] In an exemplary embodiment, unnatural amino acid residues comprise a racemic mixture of amino acid analogs. For example, in some instances, the D isomer of the amino acid analog is used. In some cases, the L isomer of the amino acid analog is used. In an exemplaryAtty. Docket: UCSF-865WOembodiment, the amino acid analog comprises chiral centers that are in the R or S configuration. In an exemplary embodiment, the amino group(s) of a a-amino acid analog is substituted with a protecting group, e.g., tert-butyloxycarbonyl (BOC group), 9-fluorenylmethyloxycarbonyl (FMOC), tosyl, and the like. In an exemplary embodiment, the carboxylic acid functional group of a P-amino acid analog is protected, e.g., as its ester derivative. In an exemplary embodiment, the salt of the amino acid analog is used.
[0126] In an exemplary embodiment, unnatural amino acid residues comprise analogs of amino acid residue alanine, valine, glycine, leucine, arginine, lysine, aspartic acid, glutamic acid, cysteine, methionine, tyrosine, phenylalanine, tryptophane, serine, threonine, or proline.
[0127] In an exemplary embodiment, an artificial nucleotide comprises, for example, modifications at one or more of ribose moiety, phosphate moiety, nucleoside moiety, or a combination thereof. In an exemplary embodiment, an artificial nucleotide comprises a nucleic acid with a modification at a 2' hydroxyl group of the ribose moiety. In an exemplary embodiment, the modification is a 2’-O-methyl modification or a 2’-O-methoxyethyl (2’-O-MOE) modification. The 2’-O-methyl modification is added a methyl group to the 2' hydroxyl group of the ribose moiety whereas the 2'-O-methoxyethyl modification is added a methoxyethyl group to the 2' hydroxyl group of the ribose moiety. In some cases, the 2' hydroxyl group includes a 2-0-aminopropyl sugar conformation which can involve an extended amine group comprising a propyl linker that binds the amine group to the 2' oxygen. In some cases, the 2' hydroxyl group includes a locked or bridged ribose conformation (e.g., locked nucleic acid or LNA) where the 4' ribose position can also be involved. In this modification, the oxygen molecule bound at the 2' carbon is linked to the 4' carbon by a methylene group, thus forming a 2'-C, 4'-C-oxy-methylene-linked bicyclic ribonucleotide monomer. In an exemplary embodiment, the 2' hydroxyl group comprises ethylene nucleic acids (ENA) such as for example 2' -4' -ethylene-bridged nucleic acid, which locks the sugar conformation into a C3 '-endo sugar puckering conformation. In an exemplary embodiment, the 2' hydroxyl group includes 2'-deoxy, T-deoxy-2'-fluoro. 2'-O-aminopropyl (2'-O-AP), 2'-O-dimethylaminoethyl (2’-0-DMA0E), 2’-O-dimethylaminopropyl (2'-0-DMAP), T-O-dimethylaminoethyloxyethyl (2'-0-DMAE0E), or 2'-O-N-methylacetamido (2'-0-NMA).Atty. Docket: UCSF-865WO
[0128] In an exemplary embodiment, a nucleotide analogue further comprises a morpholino, a peptide nucleic acid (PNA), a methylphosphonate nucleotide, a thiolphosphonate nucleotide, 2'-fluoro N3-P5'-phosphoramidite, 1’. 5'-anhydrohexitol nucleic acid (HNA), or a combination thereof.
[0129] In an exemplary embodiment, a ligand described herein comprises a small molecule ligand electrophile compound.V. Small Molecule Ligand-Electrophile Compounds
[0130] In an exemplary embodiment, a ligand-electrophile compound described herein is a small molecule compound that has a structure represented by Formula (X): F2— Q (X), wherein F2is a small molecule fragment moiety; and Q comprises a cyclic trione moiety, a cyclic dihydroxy dione moiety, a cyclic dialkoxy dione moiety, or a cyclic hydroxy alkoxy dione moiety. In some embodiments, the Q is ninhydrin. In an exemplary embodiment the small molecule probecompound has a structure represented by:O (Xc), or O (Xd), wherein Ra, when present, is unsubstituted C1-C4 alkyl and Rb, when present, is unsubstituted C1-C4 alkyl. In an exemplary embodiment the smallmolecule probe compound has a structure represented by:F2Oz0RaORh(Xlb), (XIc), or o (Xld). In an exemplary embodiment the smallAtty. Docket: UCSF-865WOomolecule probe compound has a structure represented by:o (Xlla),O (Xllb), O (XIIc), or O (Xlld), wherein Ra, when present, is unsubstituted C1-C4 alkyl and Rb, when present, is unsubstituted C1-C4 alkyl. In an exemplary embodiment, F2comprises Ci-Ce alkyl, Ci-Ce fluoroalkyl, Ci-Ce heteroalkyl, a substituted or unsubstituted C3-C6 cycloalkyl, a substituted or unsubstituted C2-C6 heterocycloalkyl, a substituted or unsubstituted aryl, or a substituted or unsubstituted heteroaryl. In an exemplary embodiment the small molecule probe compound has a structure represented by Formula (Xa), (Xb). (Xc). (Xd), (Xia). (Xlb), (XIc), (Xld), (Xlla), (Xllb), (XIIc), or (Xlld), wherein F2is -X-Y-Z”, wherein X is -CH2-, O, -C(O)NH, -NH-, -OC(O)O-, triazole, sulfonate, or phosphonate; Y is absent or substituted or unsubstituted Ci-Ce alkylene; and Z” comprises a small molecule inhibitor, or portion thereof, of cyclophilin A or PHGDH. In an exemplary embodiment the small molecule probe compound has a structure represented by Formula (Xa), (Xb), (Xc), (Xd), (Xia), (Xlb), (XIc), (Xld), (Xlla). (Xllb), (XIIc), or (Xlld). Y & X are asdescribed herein, Z’ ’ has a structure whichisF3CIn an exemplary embodiment the small molecule probe compound has a structure represented by Formula (Xa), (Xb), (Xc), (Xd), (Xia), (Xlb), (XIc), (Xld). (Xlla). (Xllb), (XIIc), or (Xlld). F2isIn an exemplary embodiment the small molecule probe compound has a structure represented by Formula (Xa), (Xb), (Xc), (Xd), (Xia), (Xlb), (XIc), (Xld), (Xlla), (Xllb), (XIIc), or (Xlld), Y & X are as described herein, Z” has a structure whichAtty. Docket: UCSF-865WOIn an exemplary embodiment the small molecule probe compound has a structure represented by Formula (Xa), (Xb), (Xc), (Xd), (Xia), (Xlb), (XIc),oF(Xld), (Xlla), (Xllb), (XIIc), or (Xlld),F2is
[0131] In an exemplary embodiment the small molecule fragment moiety has a structureRh, when present, is unsubstituted C1-C4 alkyl, X is absent, -CH2-, O, -C(O)NH, -NH-, - OC(O)O-, sulfonate, phosphonate, and triazole; Y is absent or substituted or unsubstituted Ci-Ce alkyl; Z’ is absent or a chemoselective group, such as a thioether (e.g. succinimidyl thioether) or an amide; SC is a side chain, or a portion thereof, of an amino acid residue; and M comprises oII HRz-C— C— RyI SC■vJv- wherein Ryis NH2 or an amino acid, a polypeptide (such as a wild-type functional protein or protein variants) or fragments of the polypeptide; and Rzis OH or an amino acid, a polypeptide (such as a protein (e.g. wild-type functional protein) such as tRNA or protein variants) or fragments of the polypeptide. In an exemplary embodiment the small molecule fragment moiety has a structure represented by formula (Xllla), (Xlllb), (XIIIc), or (Xllld), wherein Ra& Rb, when present, is unsubstituted C1-C4 alkyl, X is -C(O)NH-; Y is absent; Z’ isAtty. Docket: UCSF-865WOoII HRz-C— C— RyI SCabsent; SC is -(CHI)4-; and M comprises-vJv- wherein Ryis NH2 or an amino acid, a polypeptide (such as a wild-type functional protein or protein variants) or fragments of the polypeptide; and Rzis OH or an amino acid, a polypeptide (such as a protein (e.g. wild-type functional protein) such as tRNA or protein variants) or fragments of the polypeptide. In an exemplary embodiment the small molecule fragment moiety has a structure represented by formula (Xllla), (Xlllb), (XIIIc). or (Xllld). wherein Ra& Rb. when present, is unsubstituted Ci- o 11 H Rz-C- C— RyI SCC4 alkyl, X is absent; Y is absent; Z’ is absent; SC is methyl; and M comprises-vJv wherein Ryis NH2 or an amino acid, a polypeptide (such as a wild-type functional protein or protein variants) or fragments of the polypeptide; and Rzis OH or an amino acid, a polypeptide (such as a protein (e.g. wild-type functional protein) such as tRNA or protein variants) or fragments of the polypeptide.
[0132] In an exemplary embodiment, a small molecule ligand-electrophile compound of Formula (X) has a structure which is described herein.
[0133] In an exemplary embodiment, the ligand-electrophile compound has a structure which is described herein.
[0134] In an exemplary embodiment, F2is obtained from a compound library. In some cases, the compound library comprises ChemBridge fragment library, Pyramid Platform Fragment-Based Drug Discovery, Maybridge fragment library, FRGx from AnalytiCon, TCI-Frag from AnCoreX, Bio Building Blocks from ASINEX, BioFocus 3D from Charles River, and / or Fragments of Life (FOL) from Emerald Bio.
[0135] Often, a ligand-electrophile is a non-naturally occurring compound. In an exemplary embodiment, reaction of a ligand-electrophile with the guanidinium group of an arginine -containing protein results in nonnaturally occurring product. In an exemplary embodiment, the guanidinium group of the arginine-containing protein is connected to a small molecule fragment moiety via a moiety created with the side chain of an arginine residue, such as an imidazolidinimine moiety, after reaction with a ligand electrophile.Atty. Docket: UCSF-865WO VI. Further Forms of Compounds
[0136] I. General preparation of ninhydrin-comprising compounds described herein:
[0137] II. General preparation of ninhydrin-comprising compounds described herein:Amino-indanone functionalizationAtty. Docket: UCSF-865WO
[0138] III. General preparation of ninhydrin-comprising compounds described herein:
[0139] IV. General preparation of ninhydrin-comprising compounds described herein:
[0140] For the functionalizations of an indanone described herein, the final step to functionalized ninhydrins is this oxidation. Deprotections of protecting groups on R1can occur after oxidation under acidic conditions.Final ninhydrin oxidation
[0141] In an exemplary embodiment, the compound of a Formula described herein, possesses one or more stereocenters and each stereocenter exists independently in either the R or S configuration. The compounds presented herein include all diastereomeric, enantiomeric, and epimeric forms as well as the appropriate mixtures thereof. The compounds and methods provided herein include all cis, trans, syn, anti, entgegen (E), and zusammen (Z) isomers as wellAtty. Docket: UCSF-865WOas the appropriate mixtures thereof. In an exemplary embodiment, compounds described herein are prepared as their individual stereoisomers by reacting a racemic mixture of the compound with an optically active resolving agent to form a pair of diastereoisomeric compounds / salts, separating the diastereomers and recovering the optically pure enantiomers. In an exemplary embodiment, resolution of enantiomers is carried out using covalent diastereomeric derivatives of the compounds described herein. In an exemplary embodiment, diastereomers are separated by separation / resolution techniques based upon differences in solubility. In an exemplary embodiment, separation of stereoisomers is performed by chromatography or by the forming diastereomeric salts and separation by recrystallization, or chromatography, or any combination thereof. Jean Jacques, Andre Collet, Samuel H. Wilen, "Enantiomers, Racemates and Resolutions", John Wiley And Sons, Inc., 1981. In one aspect, stereoisomers are obtained by stereoselective synthesis.
[0142] In an exemplary embodiment, the compounds described herein are labeled isotopically (e.g. with a radioisotope) or by another other means, including, but not limited to, the use of chromophores or fluorescent moieties. bioluminescent labels, or chemiluminescent labels.
[0143] Compounds described herein include isotopically-labeled compounds, which are identical to those recited in the various formulae and structures presented herein, but for the fact that one or more atoms are replaced by an atom having an atomic mass or mass number different from the atomic mass or mass number usually found in nature. Examples of isotopes that can be incorporated into the present compounds include isotopes of hydrogen, carbon, nitrogen, oxygen, sulfur, fluorine and chlorine, such as, for example,2H ,3H,13C,14C,1:,N,170,180,35S,18F,36C1. In an exemplary embodiment, isotopically-labeled compounds described herein, for example those into which radioactive isotopes such as3H and14C are incorporated, are useful in drug and / or substrate tissue distribution assays. In an exemplary embodiment, substitution with isotopes such as deuterium affords certain therapeutic advantages resulting from greater metabolic stability, such as, for example, increased in vivo half-life or reduced dosage requirements.
[0144] Compounds described herein may be formed as, and / or used as, pharmaceutically acceptable salts. It should also be understood that a reference to a pharmaceutically acceptable salt includes the solvent addition forms, particularly solvates. Solvates contain eitherAtty. Docket: UCSF-865WOstoichiometric or nonstoichiometric amounts of a solvent, and may be formed during the process of crystallization with pharmaceutically acceptable solvents such as water, ethanol, and the like. Hydrates are formed when the solvent is water, or alcoholates are formed when the solvent is alcohol. Solvates of compounds described herein might be conveniently prepared or formed during the processes described herein. In addition, the compounds provided herein might exist in unsolvated as well as solvated forms. In general, the solvated forms are considered equivalent to the unsolvated forms for the purposes of the compounds and methods provided herein.VII. Arsinine-Containins Proteins
[0145] In an exemplary embodiment, disclosed herein are arginine-containing proteins that comprises one or more ligandable arginines. In an exemplary embodiment, the arginine-containing protein is a soluble protein. In an exemplary embodiment, the arginine-containing protein is a membrane protein. In an exemplary embodiment, the arginine-containing protein is associated with one or more of diseases such as cancer or one or more disorders or conditions such as immune, metabolic, developmental, reproductive, neurological, psychiatric, renal, cardiovascular, or hematological disorders or conditions.
[0146] In an exemplary embodiment, a ligandable arginine residue is located from 10A to 60 A away from an active site residue. In an exemplary embodiment, a ligandable arginine residue is located at least 10A, 12A, 15A, 20A, 25A, 30A, 35A, 40A, 45A, or 50A away from an active site residue. In an exemplary embodiment, a ligandable arginine residue is located about 10 A, 12A, 15A, 20A, 25A, 30A, 35A, 40A, 45A, or 50A away from an active site residue.
[0147] In an exemplary embodiment, the arginine-containing protein exists in an active form. In an exemplary embodiment, the arginine-containing protein exists in a pro-active form.
[0148] In an exemplary embodiment, the arginine-containing protein comprises one or more functions of an enzyme or a chaperone. In an exemplary embodiment, the arginine-containing protein comprises one or more functions of an isomerase, such as peptidyl prolyl isomerase or citrate-isocitrate isomerase; a kinase, such as tyrosine kinase; a dehydrogenase, such as phosphoglycerate dehydrogenase; and / or a chaperone, such as heat shock protein. In an exemplary embodiment, the arginine-containing protein is an enzyme or a chaperone. In an exemplary embodiment, the arginine-containing protein is an isomerase. In an exemplary embodiment, the arginine-containing protein is peptidyl prolyl isomerase or citrate-isocitrateAtty. Docket: UCSF-865WOisomerase. Tn an exemplary embodiment, the arginine-containing protein is cyclophilin A (CYPA) or aconitase. In an exemplary embodiment, the arginine-containing protein is a kinase. In an exemplary embodiment, the arginine-containing protein is a tyrosine kinase. In an exemplary embodiment, the arginine-containing protein is a receptor tyrosine kinase. In an exemplary embodiment, the arginine-containing protein is epidermal growth factor receptor (EGFR). In an exemplary embodiment, the arginine-containing protein is a non-receptor tyrosine kinase. In an exemplary embodiment, the arginine-containing protein is Bruton’s tyrosine kinase (BTK). In an exemplary embodiment, the arginine-containing protein is a dehydrogenase. In an exemplary embodiment, the arginine-containing protein is a phosphoglycerate dehydrogenase (PHGDH). In an exemplary embodiment, the arginine-containing protein is 3-phosphoglycerate dehydrogenase. In an exemplary embodiment, the arginine-containing protein is a chaperone. In an exemplary embodiment, the arginine-containing protein is a heat shock protein. In an exemplary embodiment, the arginine-containing protein is HSC70.
[0149] In an exemplary embodiment, disclosed herein is a modified arginine-containing protein which comprises a small molecule fragment moiety, covalently bonded to an arginine residue of an arginine-containing protein. In an exemplary embodiment, the arginine-containing protein comprises one or more functions of an enzyme or a chaperone. In an exemplary embodiment, the covalent bond is formed by reaction with a non-naturally occurring small molecule probe described herein. In an exemplary embodiment, the covalent bond is formed by reaction with a non-naturally occurring small molecule probe having a structure of Formula (I): F1— Q (I), wherein F1and Q are as described herein. In an exemplary embodiment, the covalent bond is formed by reaction with a non-naturally occurring ligand electrophile described herein, such as those having a structure of Formula (X): F2— Q (X), wherein F2and Q are as described herein.
[0150] In an exemplary embodiment, disclosed herein is a modified arginine-containing protein which comprises a small molecule fragment moiety, covalently bonded to an arginine residue of@HHM- -SC- -N^ N OH F1an arginine-containing protein, comprising a structure whichis °Atty. Docket: UCSF-865WO© oand M are as described herein. In an exemplary embodiment, M comprises an arginine-containing protein, such as one described herein. In an exemplary embodiment, disclosed herein is a modified arginine-containing protein which comprises a small molecule fragment moiety, covalently bonded to an arginine residue of an arginine-containing protein, comprising a, wherein F1, SC, and M are as described herein. In an exemplary embodiment, M comprises an arginine-containing protein, such as one described herein.
[0151] In an exemplary embodiment, disclosed herein is a modified arginine-containing protein which comprises a small molecule fragment moiety, covalently bonded to an arginine residue ofan arginine-containing protein, comprising a structure which is°and M are as described herein. In an exemplary embodiment, M comprises an arginine-containing protein, such as one described herein. In an exemplary embodiment, disclosed herein is a modified arginine-containing protein which comprises a small molecule fragment moiety, covalently bonded to an arginine residue of an arginine-containing protein, comprising aAtty. Docket: UCSF-865WO© H HM— SC— N R OH H2N HC M— SC— N HNstructure whichis wherein F2, SC, and M are as described herein. Tn an exemplary embodiment, M comprises an arginine-containing protein, such as one described herein.VIII. Cells, Analytical Techniques, and Instrumentation
[0152] In an exemplary embodiment, one or more of the methods disclosed herein comprise a sample (e.g., a cell sample, or a cell lysate sample). In an exemplary embodiment, the sample for use with the methods described herein is obtained from cells of an animal. In some instances, the animal cell includes a cell from a marine invertebrate, fish, insects, amphibian, reptile, or mammal. In an exemplary embodiment, the mammalian cell is a primate, ape, equine, bovine, porcine, canine, feline, or rodent. In an exemplary embodiment, the mammal is a primate, ape, dog. cat, rabbit, ferret, or the like. In an exemplary embodiment, the rodent is a mouse, rat, hamster, gerbil, hamster, chinchilla, or guinea pig. In an exemplary embodiment, the bird cell is from a canary, parakeet or parrots. In an exemplary embodiment, the reptile cell is from a turtle, lizard or snake. In some cases, the fish cell is from a tropical fish. In an exemplary embodiment, the fish cell is from a zebrafish (e.g. Danino rerio). In an exemplary embodiment, the worm cell is from a nematode (e.g. C. elegans). In an exemplary embodiment, the amphibian cell is from a frog. In an exemplary embodiment, the arthropod cell is from a tarantula or hermit crab.
[0153] In an exemplary embodiment, the sample for use with the methods described herein is obtained from a mammalian cell. In an exemplary embodiment, the mammalian cell is an epithelial cell, connective tissue cell, hormone secreting cell, a nerve cell, a skeletal muscle cell, a blood cell, or an immune system cell.
[0154] Exemplary mammalian cells include, but are not limited to, 293 A cell line, 293FT cell line, 293F cells, 293 H cells, HEK 293 cells, CHO DG44 cells, CHO-S cells, CHO-K1 cells, Expi293F™ cells, Flp-In™ T-REx™ 293 cell line, Flp-In™-293 cell line, Flp-In™-3T3 cell line, Flp-In™-BHK cell line, Flp-In™-CHO cell line, Flp-InTM-CV-l cell line, Flp-In™-Jurkat cell line, FreeStyle™ 293-F cells, FreeStyle™ CHO-S cells, GripTite™ 293 MSR cell line, GS-Atty. Docket: UCSF-865WOCHO cell line, HepaRG™ cells, T-REx™ Jurkat cell line, Per.C6 cells, T-REx™-293 cell line, T-REx™CHO cell line, T-REx™-HeLa cell line, NC-HIMT cell line, and PC 12 cell line.
[0155] In an exemplary embodiment, the sample for use with the methods described herein is obtained from cells of a tumor cell line. In an exemplary embodiment, the sample is obtained from cells of a solid tumor cell line. In an exemplary embodiment, the solid tumor cell line is a sarcoma cell line. In some instances, the solid tumor cell line is a carcinoma cell line. In an exemplary embodiment, the sarcoma cell line is obtained from a cell line of alveolar rhabdomyosarcoma, alveolar soft part sarcoma, ameloblastoma, angiosarcoma, chondrosarcoma, chordoma, clear cell sarcoma of soft tissue, dedifferentiated liposarcoma, desmoid, desmoplastic small round cell tumor, embryonal rhabdomyosarcoma, epithelioid fibrosarcoma, epithelioid hemangioendothelioma, epithelioid sarcoma, esthesioneuroblastoma, Ewing sarcoma, extrarenal rhabdoid tumor, extraskeletal myxoid chondrosarcoma, extraskeletal osteosarcoma, fibrosarcoma, giant cell tumor, hemangiopericytoma, infantile fibrosarcoma, inflammatory myofibroblastic tumor, Kaposi sarcoma, leiomyosarcoma of bone, liposarcoma, liposarcoma of bone, malignant fibrous histiocytoma (MFH), malignant fibrous histiocytoma (MFH) of bone, malignant mesenchymoma, malignant peripheral nerve sheath tumor, mesenchymal chondrosarcoma, myxofibrosarcoma, myxoid liposarcoma. myxoinflammatory fibroblastic sarcoma, neoplasms with perivascular epitheioid cell differentiation, osteosarcoma, parosteal osteosarcoma, neoplasm with perivascular epitheioid cell differentiation, periosteal osteosarcoma, pleomorphic liposarcoma, pleomorphic rhabdomyosarcoma. PNET / extraskeletal Ewing tumor, rhabdomyosarcoma, round cell liposarcoma, small cell osteosarcoma, solitary fibrous tumor, synovial sarcoma, telangiectatic osteosarcoma.
[0156] In an exemplary embodiment, the carcinoma cell line is obtained from a cell line of adenocarcinoma, squamous cell carcinoma, adenosquamous carcinoma, anaplastic carcinoma, large cell carcinoma, small cell carcinoma, anal cancer, appendix cancer, bile duct cancer (i.e., cholangiocarcinoma), bladder cancer, brain tumor, breast cancer, cervical cancer, colon cancer, cancer of Unknown Primary (CUP), esophageal cancer, eye cancer, fallopian tube cancer, gastroenterological cancer, kidney cancer, liver cancer, lung cancer, medulloblastoma, melanoma, oral cancer, ovarian cancer, pancreatic cancer, parathyroid disease, penile cancer, pituitary tumor, prostate cancer, rectal cancer, skin cancer, stomach cancer, testicular cancer, throat cancer, thyroid cancer, uterine cancer, vaginal cancer, or vulvar cancer.Atty. Docket: UCSF-865WO
[0157] In an exemplary embodiment, the sample is obtained from cells of a hematologic malignant cell line. In some instances, the hematologic malignant cell line is a T-cell cell line. In an exemplary embodiment, B-cell cell line. In an exemplary embodiment, the hematologic malignant cell line is obtained from a T-cell cell line of: peripheral T-cell lymphoma not otherwise specified (PTCL-NOS), anaplastic large cell lymphoma, angioimmunoblastic lymphoma, cutaneous T-cell lymphoma, adult T-cell leukemia / lymphoma (ATLL), blastic NK-cell lymphoma, enteropathy-type T-cell lymphoma, hematosplenic gamma-delta T-cell lymphoma, lymphoblastic lymphoma, nasal NK / T-cell lymphomas, or treatment-related T-cell lymphomas.
[0158] In an exemplary embodiment, the hematologic malignant cell line is obtained from a B-cell cell line of: acute lymphoblastic leukemia (ALL), acute myelogenous leukemia (AML), chronic myelogenous leukemia (CML), acute monocytic leukemia (AMoL), chronic lymphocytic leukemia (CLL), high-risk chronic lymphocytic leukemia (CLL), small lymphocytic lymphoma (SLL), high-risk small lymphocytic lymphoma (SLL), follicular lymphoma (FL), mantle cell lymphoma (MCL), Waldenstrom's macroglobulinemia. multiple myeloma, extranodal marginal zone B cell lymphoma, nodal marginal zone B cell lymphoma, Burkitt's lymphoma, non-Burkitt high grade B cell lymphoma, primary mediastinal B-cell lymphoma (PMBL), immunoblastic large cell lymphoma, precursor B-lymphoblastic lymphoma, B cell prolymphocytic leukemia, lymphoplasmacytic lymphoma, splenic marginal zone lymphoma, plasma cell myeloma, plasmacytoma, mediastinal (thymic) large B cell lymphoma, intravascular large B cell lymphoma, primary effusion lymphoma, or lymphomatoid granulomatosis.
[0159] In an exemplary embodiment, the sample for use with the methods described herein is obtained from a tumor cell line. Exemplary tumor cell line includes, but is not limited to, 600MPE. AU565, BT-20. BT-474. BT-483. BT-549. Evsa-T, Hs578T, MCF-7, MDA-MB-23I, SkBr3, T-47D, HeLa, DU145, PC3, LNCaP, A549, H1299, NCI-H460, A2780, SKOV-3 / Luc, Neuro2a, RKO, RKOAS45-1, HT-29, SW1417. SW948, DLD-1, SW480, Capan-I, MC / 9, B72.3, B25.2, B6.2, B38.1, DMS 153, SU.86.86, SNU-182, SNU-423, SNU-449, SNU-475, SNU-387, Hs 817.T, LMH, LMH / 2A, SNU-398, PLHC-1, HepG2 / SF, OCI-Lyl, OCI-Ly2, OCI-Ly3, OCI-Ly4. OCI-Ly6, OCI-Ly7, OCI-LylO, OCI-Lyl8, OCI-Lyl9, U2932. DB, HBL-1, RIVA, SUDHL2, TMD8, MECI, MEC2, 8E5, CCRF-CEM, MOLT-3, TALL-104, AML-193,Atty. Docket: UCSF-865WOTHP-1, BDCM, HL-60, Jurkat, RPMT 8226, MOLT-4, RS4, K-562, KASUMT-1 , Daudi, GA-10, Raji, JeKo-1, NK-92, and Mino.
[0160] In an exemplary embodiment, the sample for use in the methods is from any tissue or fluid from an individual. Samples include, but are not limited to, tissue (e.g. connective tissue, muscle tissue, nervous tissue, or epithelial tissue), whole blood, dissociated bone marrow, bone marrow aspirate, pleural fluid, peritoneal fluid, central spinal fluid, abdominal fluid, pancreatic fluid, cerebrospinal fluid, brain fluid, ascites, pericardial fluid, urine, saliva, bronchial lavage, sweat, tears, ear flow, sputum, hydrocele fluid, semen, vaginal flow, milk, amniotic fluid, and secretions of respiratory, intestinal or genitourinary tract. In an exemplary embodiment, the sample is a tissue sample, such as a sample obtained from a biopsy or a tumor tissue sample. In some embodiments, the sample is a blood serum sample. In an exemplary embodiment, the sample is a blood cell sample containing one or more peripheral blood mononuclear cells (PBMCs). In an exemplary embodiment, the sample contains one or more circulating tumor cells (CTCs). In an exemplary embodiment, the sample contains one or more disseminated tumor cells (DTC, e.g.. in a bone marrow aspirate sample).
[0161] In an exemplary embodiment, the samples are obtained from the individual by any suitable means of obtaining the sample using well-known and routine clinical methods.Procedures for obtaining tissue samples from an individual are well known. For example, procedures for drawing and processing tissue sample such as from a needle aspiration biopsy is well-known and is employed to obtain a sample for use in the methods provided. Typically, for collection of such a tissue sample, a thin hollow needle is inserted into a mass such as a tumor mass for sampling of cells that, after being stained, will be examined under a microscope.Sample Preparation and Analysis
[0162] In an exemplary embodiment, the sample (e.g., cell sample, cell lysate sample, or comprising isolated proteins) is a sample solution. In an exemplary embodiment, the sample solution comprises a solution such as a buffer (e.g. phosphate buffered saline) or a media. In an exemplary embodiment, the media is an isotopically labeled media. In an exemplary embodiment, the sample solution is a cell solution.
[0163] In an exemplary embodiment, the sample (e.g., cell sample, cell lysate sample, or comprising isolated proteins) is incubated with one or more compound probes for analysis ofAtty. Docket: UCSF-865WOprotein-probe interactions. In an exemplary embodiment, the sample (e.g., cell sample, cell lysate sample, or comprising isolated proteins) is further incubated in the presence of an additional compound probe prior to addition of the one or more probes. In an exemplary embodiment, the sample (e.g., cell sample, cell lysate sample, or comprising isolated proteins) is further incubated with a non-probe small molecule ligand, in which the non-probe small molecule ligand does not contain a photoreactive moiety and / or an alkyne group. In an exemplary embodiment, the sample is incubated with a probe and non-probe small molecule ligand for competitive protein profiling analysis.
[0164] In an exemplary embodiment, the sample is compared with a control. In an exemplary embodiment, a difference is observed between a set of probe protein interactions between the sample and the control. In an exemplary embodiment, the difference correlates to the interaction between the small molecule fragment and the proteins.
[0165] In an exemplary embodiment, one or more methods are utilized for labeling a sample (e.g. cell sample, cell lysate sample, or comprising isolated proteins) for analysis of probe protein interactions. In an exemplary embodiment, a method comprises labeling the sample (e.g. cell sample, cell lysate sample, or comprising isolated proteins) with an enriched media. In an exemplary embodiment, the sample (e.g. cell sample, cell lysate sample, or comprising isolated proteins) is labeled with isotope-labeled amino acids, such as13C or15N-labeled amino acids. In an exemplary embodiment, the labeled sample is further compared with a non-labeled sample to detect differences in probe protein interactions between the two samples. In an exemplary embodiment, this difference is a difference of a target protein and its interaction with a small molecule ligand in the labeled sample versus the non-labeled sample. In an exemplary embodiment, the difference is an increase, decrease or a lack of protein-probe interaction in the two samples. In an exemplary embodiment, the isotope-labeled method is termed SILAC. stable isotope labeling using amino acids in cell culture.
[0166] In some embodiments, a method comprises incubating a sample (e.g. cell sample, cell lysate sample, or comprising isolated proteins) with a labeling group (e.g., an isotopically labeled labeling group) to tag one or more proteins of interest for further analysis. In such cases, the labeling group comprises a biotin, a streptavidin, bead, resin, a solid support, or a combination thereof, and further comprises a linker that is optionally isotopically labeled. As described above.Atty. Docket: UCSF-865WOthe linker can be about 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15 or more residues in length and might further comprise a cleavage site, such as a protease cleavage site (e.g., TEV cleavage site). In an exemplary embodiment, the labeling group is a biotin-linker moiety, which is optionally isotopically labeled with13C and15N atoms at one or more amino acid residue positions within the linker. In an exemplary embodiment, the biotin-linker moiety is an isotopically-labeled TEV-tag as described in Weerapana, et al., "Quantitative reactivity profiling predicts functional cysteines in proteomes," Nature 468(7325): 790-795.
[0167] In an exemplary embodiment, isobaric tags for relative and absolute quantitation (iTRAQ) method is utilized for processing a sample. In an exemplary embodiment, the iTRAQ method is based on the covalent labeling of the N-terminus and side chain amines of peptides from a processed sample. In an exemplary embodiment, reagent such as 4-plex or 8-plex is used for labeling the peptides.
[0168] In an exemplary embodiment, tandem mass tags (TMT) are utilized for processing a sample. In an exemplary embodiment, the TMT is selected from TMTzero, TMTduplex, TMTsixplex, TMT 10-plex, TMTpro, and TMTpro Zero.
[0169] In an exemplary embodiment, the probe-protein complex is further conjugated to a chromophore, such as a fluorophore. In an exemplary embodiment, the probe-protein complex is separated and visualized utilizing an electrophoresis system, such as through a gel electrophoresis, or a capillary electrophoresis. Exemplary gel electrophoresis includes agarose based gels, polyacrylamide based gels, or starch based gels. In an exemplary embodiment, the probe-protein is subjected to a native electrophoresis condition. In an exemplary embodiment, the probe-protein is subjected to a denaturing electrophoresis condition.
[0170] In an exemplary embodiment, the probe-protein after harvesting is further fragmentized to generate protein fragments. In an exemplary embodiment, fragmentation is generated through mechanical stress, pressure, or chemical means. In an exemplary embodiment, the protein from the probe-protein complexes is fragmented by a chemical means. In an exemplary embodiment, the chemical means is a protease. Exemplary proteases include, but are not limited to, serine proteases such as chymotrypsin A, penicillin G acylase precursor, di peptidase E, DmpA aminopeptidase, subtilisin, prolyl oligopeptidase, D-Ala-D-Ala peptidase C, signal peptidase I, cytomegalovirus assemblin, Lon-A peptidase, peptidase Clp, Escherichia coli phage KIFAtty. Docket: UCSF-865WOendosialidase CIMCD self-cleaving protein, nucleoporin 145, lactoferrin, murein tetrapeptidase LD-carboxypeptidase, or rhomboid- 1; threonine proteases such as ornithine acetyltransferase; cysteine proteases such as TEV protease, amidophosphoribosyltransferase precursor, gammaglutamyl hydrolase (Rattus norvegicus), hedgehog protein, DmpA aminopeptidase, papain, bromelain, cathepsin K, calpain, caspase- 1, separase, adenain, pyroglutamyl-peptidase I, sortase A, hepatitis C virus peptidase 2, sindbis virustype nsP2 peptidase, dipeptidyl-peptidase VI, or DeSI-1 peptidase; aspartate proteases such as beta-secretase I (BACEI), beta-secretase 2 (BACE2), cathepsin D, cathepsin E, chymosin, napsin-A, nepenthesin, pepsin, plasmepsin, presenilin, or renin; glutamic acid proteases such as AfuGprA; and metalloproteases such as peptidase_M48.
[0171] In an exemplary embodiment, the fragmentation is a random fragmentation. In an exemplary embodiment, the fragmentation generates specific lengths of protein fragments, or the shearing occurs at particular sequence of amino acid regions.
[0172] In an exemplary embodiment, the protein fragments are further analyzed by a proteomic method such as by liquid chromatography (LC) (e.g. high performance liquid chromatography), liquid chromatography-mass spectrometry (LC-MS), matrix-assisted laser desorption / ionization (MALDITOF), gas chromatography-mass spectrometry (GC-MS), capillary electrophoresis-mass spectrometry (CE-MS), or nuclear magnetic resonance imaging (NMR).
[0173] In an exemplary embodiment, the LC method is any suitable LC methods well known in the art, for separation of a sample into its individual parts. This separation occurs based on the interaction of the sample with the mobile and stationary phases. Since there are many stationary / mobile phase combinations that are employed when separating a mixture, there are several different types of chromatography that are classified based on the physical states of those phases. In some embodiments, the LC is further classified as normal-phase chromatography, reverse-phase chromatography, size-exclusion chromatography, ion-exchange chromatography, affinity chromatography, displacement chromatography, partition chromatography, flash chromatography, chiral chromatography, and aqueous normal-phase chromatography.
[0174] In an exemplary embodiment, the LC method is a high performance liquid chromatography (HPLC) method. In an exemplary embodiment, the HPLC method is further categorized as normal-phase chromatography, reverse-phase chromatography, size-exclusionAtty. Docket: UCSF-865WOchromatography, ion-exchange chromatography, affinity chromatography, displacement chromatography, partition chromatography, chiral chromatography, and aqueous normal-phase chromatography.
[0175] In an exemplary embodiment, the HPLC method of the present disclosure is performed by any standard techniques well known in the art. Exemplary HPLC methods include hydrophilic interaction liquid chromatography (HILIC), electrostatic repulsion-hydrophilic interaction liquid chromatography (ERL1C) and reverse phase liquid chromatography (RPLC).
[0176] In an exemplary embodiment, the LC is coupled to a mass spectroscopy as a LC-MS method. In an exemplary embodiment, the LC-MS method includes ultra-performance liquid chromatography electrospray ionization quadrupole time-of-flight mass spectrometry (UPLC-ESLQTOF-MS), ultraperformance liquid chromatography-electrospray ionization tandem mass spectrometry (UPLCESI-MS / MS), reverse phase liquid chromatography-mass spectrometry (RPLC-MS). hydrophilic interaction liquid chromatography-mass spectrometry (HILIC-MS). hydrophilic interaction liquid chromatography-triple quadrupole tandem mass spectrometry (HILIC-QQQ), electrostatic repulsion-hydrophilic interaction liquid chromatography-mass spectrometry (ERLIC-MS), liquid chromatography time-of-flight mass spectrometry (LC-QTOF-MS), liquid chromatography-tandem mass spectrometry (LC-MS / MS), multidimensional liquid chromatography coupled with tandem mass spectrometry (LC / LC-MS / MS). In an exemplary embodiment, the LC-MS method is LC / LC-MS / MS. In an exemplary embodiment, the LC-MS methods of the present disclosure are performed by standard techniques well known in the art.
[0177] In an exemplary embodiment, the nuclear magnetic resonance (NMR) method is any suitable method well known in the art for the detection of one or more cysteine binding proteins or protein fragments disclosed herein. In an exemplary embodiment, the NMR method includes one dimensional (ID) NMR methods, two dimensional (2D) NMR methods, solid state NMR methods and NMR chromatography. Exemplary ID NMR methods include ’Hydrogen.13Carbon,15Nitrogen,17Oxygen,19Fluorine,31Phosphorus,39Potassium,23Sodium,33Sulfur,87Strontium,27Aluminum,43Calcium,35Chlorine,37Chlorine,63Copper,65Copper,57Iron,25Magnesium,199Mercury or67Zinc NMR method, distortionless enhancement by polarization transfer (DEPT) method, attached proton test (APT) method and ID-incredible natural abundance double quantum transition experiment (INADEQUATE) method. Exemplary 2D NMR methodsAtty. Docket: UCSF-865WOinclude correlation spectroscopy (COSY), total correlation spectroscopy (TOCSY), 2D-INADEQUATE, 2D-adequate double quantum transfer experiment (ADEQUATE), nuclear overhauser effect spectroscopy (NOSEY), rotating-frame NOE spectroscopy (ROESY), heteronuclear multiple-quantum correlation spectroscopy (HMQC), heteronuclear single quantum coherence spectroscopy (HSQC), short range coupling and long range coupling methods. Exemplary solid state NMR method include solid state13Carbon NMR, high resolution magic angle spinning (HR-MAS) and cross polarization magic angle spinning (CP-MAS) NMR methods. Exemplary NMR techniques include diffusion ordered spectroscopy (DOSY), DOSY-TOCSY and DOSY-HSQC.
[0178] In an exemplary embodiment, the protein fragments are analyzed by method as described in Weerapana et al., "Quantitative reactivity profiling predicts functional cysteines in proteomes," Nature, 468:790-795 (2010).
[0179] In an exemplary embodiment, the results from the mass spectroscopy method are analyzed by an algorithm for protein identification. In an exemplary embodiment, the algorithm combines the results from the mass spectroscopy method with a protein sequence database for protein identification. In some embodiments, the algorithm comprises ProLuCID algorithm, Probity, Scaffold, SEQUEST, or Mascot.
[0180] In an exemplary embodiment, a value is assigned to each of the protein from the probeprotein complex. In an exemplary embodiment, the value assigned to each of the protein from the probe-protein complex is obtained from the mass spectroscopy analysis. In an exemplary embodiment, the value is the area-under-the curve from a plot of signal intensity as a function of mass-to-charge ratio. In an exemplary embodiment, the value correlates with the reactivity of an Arg residue within a protein.
[0181] In an exemplary embodiment, a ratio between a first value obtained from a first protein sample and a second value obtained from a second protein sample is calculated. In some instances, the ratio is greater than 2.5, 3, 3.5, 4, 4.5, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20. In an exemplary embodiment, the ratio is at most 20.
[0182] In an exemplary embodiment, the ratio is calculated based on averaged values. In an exemplary embodiment, the averaged value is an average of at least two, three, or four values of the protein from each cell solution, or that the protein is observed at least two, three, or fourAtty. Docket: UCSF-865WOtimes in each cell solution and a value is assigned to each observed time. In an exemplary embodiment, the ratio further has a standard deviation of less than 12, 10, or 8.
[0183] In an exemplary embodiment, a value is not an averaged value. In an exemplary embodiment, the ratio is calculated based on value of a protein observed only once in a cell population. In an exemplary embodiment, the ratio is assigned with a value of 20.IX. Kits / Article of Manufacture
[0184] Disclosed herein, in certain embodiments, are kits and articles of manufacture for use with one or more methods described herein. In an exemplary embodiment, described herein is a kit for generating a protein comprising a photoreactive ligand. In an exemplary embodiment, such kit includes photoreactive small molecule ligands described herein, small molecule fragments or libraries and / or controls, and reagents suitable for carrying out one or more of the methods described herein. In an exemplary embodiment, the kit further comprises samples, such as a cell sample, and suitable solutions such as buffers or media. In some embodiments, the kit further comprises recombinant proteins for use in one or more of the methods described herein. In an exemplary embodiment, additional components of the kit comprises a carrier, package, or container that is compartmentalized to receive one or more containers such as vials, tubes, and the like, each of the container(s) comprising one of the separate elements to be used in a method described herein. Suitable containers include, for example, bottles, vials, plates, syringes, and test tubes. In an exemplary embodiment, the containers are formed from a variety of materials such as glass or plastic.
[0185] The articles of manufacture provided herein contain packaging materials. Examples of pharmaceutical packaging materials include, but are not limited to, bottles, tubes, bags, containers, and any packaging material suitable for a selected formulation and intended mode of use.
[0186] For example, the container(s) include probes, test compounds, and one or more reagents for use in a method disclosed herein. Such kits optionally include an identifying description or label or instructions relating to its use in the methods described herein.
[0187] A kit typically includes labels listing contents and / or instructions for use, and package inserts with instructions for use. A set of instructions will also typically be included.Atty. Docket: UCSF-865WO
[0188] In an exemplary embodiment, a label is on or associated with the container. In an exemplary embodiment, a label is on a container when letters, numbers or other characters forming the label are attached, molded or etched into the container itself; a label is associated with a container when it is present within a receptacle or carrier that also holds the container, e.g., as a package insert. In an exemplary embodiment, a label is used to indicate that the contents are to be used for a specific therapeutic application. The label also indicates directions for use of the contents, such as in the methods described herein.X. Certain Terminology
[0189] Unless defined otherwise, all technical and scientific terms used herein have the same meaning as is commonly understood by one of skill in the art to which the claimed subject matter belongs. It is to be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of any subject matter claimed. In this application, the use of the singular includes the plural unless specifically stated otherwise. It must be noted that, as used in the specification and the appended claims, the singular forms "a," "an" and "the" include plural referents unless the context clearly dictates otherwise. In this application, the use of "or" means "and / or" unless stated otherwise.Furthermore, use of the term "including" as well as other forms, such as "include", "includes," and "included," is not limiting.
[0190] As used herein, ranges and amounts can be expressed as "about" a particular value or range. About also includes the exact amount. Hence "about 5 pL" means "about 5 pL" and also "5 pL." Generally, the term "about" includes an amount that would be expected to be within experimental error.
[0191] The section headings used herein are for organizational purposes only and are not to be construed as limiting the subject matter described.EXAMPLES
[0192] The following examples are put forth to provide those of ordinary skill in the art with a complete disclosure and description of how to make and use the present invention and are not intended to limit the scope of what the inventors regard as their invention nor are they intended to represent that the experiments below are all or the only experiments performed. Efforts haveAtty. Docket: UCSF-865WObeen made to ensure accuracy with respect to numbers used (e.g. amounts, temperature, etc.) but some experimental errors and deviations should be accounted for. Unless indicated otherwise, parts are parts by weight, molecular weight is weight average molecular weight, temperature is in degrees Centigrade, and pressure is at or near atmospheric.EXAMPLE 1Materials and Methods:Density Functional Theory (DFT) Calculations
[0193] Initial molecular models were prepared using Avogadro. All calculations were performed using GAMESS on ChemCompute.org. Geometry optimization was performed with the PM3 basis set, restricted Hartree-Fock (RHF) molecular orbital method, and B3LYP DFT functional. Solvent effects were modeled using the water PCM solvent model on a single processor core. The optimized structures were subjected to a single-point energy calculation with the 6-311G** basis set. All output models were reviewed in Avogadro for realistic geometry, and the delta G free energy in solvent between the hydrated and dehydrated states (with an additional water molecule) for both phenyl glyoxal and ninhydrin was calculated to determine the relative equilibrium constants.Fmoc- Arginine and Lysine Kinetics Assay
[0194] In triplicate, 190 pL of PBS buffer (pH 6.0 - 10.0, in 0.5 pH unit increments) was allotted into 27 Eppendorf tubes. A 5 pL aliquot of a 20 mM stock solution of Fmoc-Arg-OH(or Fmoc-Lys-OH) in DMSO was added to each tube, mixed thoroughly, and briefly centrifuged to ensure complete pooling of the solution. A 600 mM Ninhydrin solution (1:1 DMSO / DI water) was prepared alongside a 2 M HC1 quenching solution. All tubes were pre-opened, and a timer, vortex mixer, waste container, and fresh pipette tips were arranged for efficient handling.Reactions were initiated by adding 5 pL of Ninhydrin solution to the first tube while starting the timer, followed by immediate capping and vortex mixing. This process was repeated every 10 seconds for the subsequent tubes, ensuring all reactions were initiated within 4.5 minutes. At designated time points, the reactions were quenched by sequentially adding 10 pL of 2 M HC1 to each tube, following the same 10-second interval sequence as for reaction initiation. Samples were then centrifuged, and 200 pL of the supernatant was carefully transferred using gel-loading tips into pre-labeled LC-MS vials with conical inserts, ensuring minimal air bubble formation.Atty. Docket: UCSF-865WOSamples were analyzed using a Waters Acquity UPLC SQD2 detector LC-MS system on a polar gradient method (20-55% acetonitrile / water + 0.1% formic acid). The UV detection window was set to 300 nm for Fmoc absorbance, and free arginine (SM) and adduct formation were identified based on the extracted ion chromatograms. Peak integrations at 300 nm were used to calculate the conversion ratio of reactants to products.Cysteine Reactivity Assay
[0195] L-cysteine was prepared at 250 pM in assay buffer (100 mM Tris. pH 9.0, 50% MeOH). In triplicate, 100 pL of the cysteine solution was dispensed into each well of a clear, flat-bottom 96-well plate (Costar). Inhibitors were added as 100X stocks (1 pL per well; 500 pM final concentration) or with DMSO vehicle control, and reactions were incubated at room temperature for the indicated times (0-60min). Following incubation, Ellman’s reagent (10 pL of a 50mM stock in DMSO; 5mM final concentration) was added, and plates were incubated for 10 minutes at room temperature. Absorbance was measured at 412nm in a plate reader (Spectramax M5). Free cysteine concentrations were quantified using a standard curve, and second-order rate constants were calculated as previously described.Intact Protein Mass Spectrometry
[0196] Reactions were performed with 1.5 pM BSA or Lysozyme in PBS, incubated with 150 pM ninhydrin at 37 °C for 30 minutes. Reactions were quenched by a dilution to 0.1% TFA (aq) and analyzed by intact protein LC / MS using an Agilent 6230 TOF LC / MS with a Dual AJS ESI ion source. Agilent 6230 TOF system equipped with an Agilent RLRP-S 1000A 5 pm 50x2 1MM column. The mobile phase was a linear gradient of 5-90% acetonitrile / water + 0.1% formic acid over 4 minutes with a flow rate of 300 pL / min.Glutathione reactivity assay
[0197] Reduced glutathione (GSH) was prepared at a concentration of 250 pM in assay buffer (100 mM Tris pH 9.0, 50% MeOH as co-solvent). In triplicate, 100 pL of GSH (250 pM) in assay buffer was added to each well of a clear, flat-bottom 96-well plate (Greiner Bio-one). To each well, 1 pL of lOOx inhibitor stocks in DMSO (final concentration 500 pM) or DMSO vehicle was added, and the plate was incubated at room temperature for the indicated time (0-60 min).Following incubation, Ellman’s reagent (Thermo) (10 pL, 50 mM stock in DMSO, final concentration 5 mM) was added, and the reaction was incubated at room temperature forAtty. Docket: UCSF-865WO10 minutes. Absorbance was measured at 412 nm using a plate reader (Cytation 5, BioTek). The remaining concentration of free thiols was quantified using a glutathione standard curve.Cell Culture
[0198] Mino and A549 cells were cultured under standard conditions (37°C, 5% CO2) in T175 flasks. Both cell lines were revived from frozen stocks stored in liquid nitrogen. Cryovials containing the cells were rapidly thawed in a 37°C water bath and immediately transferred to pre-warmed complete culture media. To remove residual DMSO, cells were centrifuged at 300 x g for 5 min, resuspended in fresh culture media, and plated in appropriate flasks. Cells were allowed to recover for 24-48 hours before the first media change and were subsequently maintained under routine culture conditions.
[0199] Mino cells were maintained in non-treated T175 flasks in Cytiva RPMI-1640 medium (no HEPES, stored at 4°C) supplemented with 10% fetal bovine serum (HyClone), L-glutamine (100 pg / mL), 0.05 mM p-mercaptoethanol, and penicillin- streptomycin (100 pg / mL). Cells were routinely passaged every 2-3 days when they reached a density of ~8 x 105cells / mL to maintain exponential growth. Passaging was performed by centrifuging the culture at 300 x g for 5 min, discarding the supernatant, and resuspending the cell pellet in fresh pre- warmed medium at a 1:3 dilution.
[0200] A549 cells were cultured in tissue culture-treated T175 flasks using Cytiva Dulbecco’s Modified Eagle Medium (DMEM) supplemented with 10% FBS, GlutaMAX, and Pen / Strep. Cells were passaged at -80-90% confluency using enzymatic detachment. The culture media was aspirated, and cells were washed once with phosphate-buffered saline (PBS) before treatment with 0.05% Trypsin-EDTA at 37 °C for 2-5 min. Detachment was confirmed under a microscope, and trypsinization was quenched by adding complete medium. The cells were then collected, pelleted at 300 x g for 5 min, resuspended in fresh media, and replated at a 1:3 dilution.
[0201] Both cell lines were harvested after three passages to ensure robust growth and adaptation to culture conditions. Mino cells were collected from suspension culture, pelleted at 300 x g for 5 min, and washed once with ice-cold PBS. A549 cells were harvested by aspirating the culture medium, washing twice with PBS, and detaching with Trypsin-EDTA as described above.Detached A549 cells were collected in fresh media, pelleted at 300 x g for 5 min, and washedAtty. Docket: UCSF-865WOonce with ice-cold PBS. Washed cell pellets were resuspended in ice-cold Dulbecco’s phosphate-buffered saline (DPBS) and sonicated on ice (2 x 20 pulses, 20% amplitude, 1 s pulse rate) using a tip sonicator. The lysed suspension was centrifuged at 21,000 x g for 20 min at 4°C. and the supernatant was transferred to a lobind tube. Protein concentrations were determined using the bicinchoninic acid (BCA) assay (Thermo Scientific), and the resulting lysate was used fresh for subsequent experiments.Gel-Based Probe Reactivity Profiling
[0202] 50 pL of freshly prepared cell lysates (2 mg / mL total protein) were treated as indicated in PBS buffer (pH 7.4) at 37°C. Lysates were dosed with a 100X probe in DMSO or an equal volume of DMSO as a control. Following probe treatment, CuAAC click chemistry was performed by sequentially adding 2 pL TAMRA-azide (5 mM, stock in DMSO), 2 pL tris(2-carboxyethyl)phosphine hydrochloride (TCEP) (50 mM, freshly prepared in water), 6 pL tris[( 1-benzyl-l-H-l,2,3-triazol-4-yl)methyl]amine (TBTA) (1.7 mM, 3:1 tert-butanol:DMSO), and 2 pL CUSO4*5H2O (50 mM in water) with vortexing between additions. The reaction was incubated at room temperature for 1 hour with gentle mixing.
[0203] For live-cell screening, 106cells were plated in 3 mL of complete cell growth medium (as described above) in a single well of a 12-well culture plate (Costar). Wells were treated with either 1000X of either probe or vehicle (0.1% v / v) for Ih under standard conditions (37°C, 5% CO2). Immediately following labeling, cells were washed 3X with 5 mL DPBS, pelleted, and flash-frozen in LN2 and stored at -80°C overnight. The following day, pellets were lysed and protein concentrations were normalized to 50 pg (1 mg / mL). Click reactions were then performed under the same conditions as above.
[0204] Reactions were quenched by adding 250 pL of ice-cold methanol, followed by incubation at -20°C for at least 2 hours. Precipitated proteins were pelleted by centrifugation (10,000 x g, 10 min, 4°C), and the supernatant was discarded. Tubes were inverted and allowed to dry for 30 minutes to ensure complete evaporation of methanol. Protein pellets were resuspended in 75 pL of 4% SDS in PBS (pH 7.5) and bath sonicated (35 kHz, 10 min) to achieve complete solubilization. Following resuspension, 25 pL of 4x Laemmli loading buffer (Bio-Rad, containing 10% P-mercaptoethanol) was added, and samples were vortexed gently before boiling at 95 °C for 5 minutes. After cooling to room temperature, 20 pg of protein per lane was loadedAtty. Docket: UCSF-865WOonto a 4-15% Tris-Glycine SDS-PAGE gel (Bio-Rad, Criterion TGX) for separation.Fluorescently labeled proteins were visualized using a Bio-Rad ChemiDoc fluorescence scanner, and total protein staining with Coomassie was performed for normalization.RisoLFG [RAP (Reactive Arginine Profiling)] SP3 Workflow
[0205] 250pL freshly prepared cell lysate at 2 pg / mL protein concentration was treated in lobind tubes with either lOOpM probe or l.OrnM probe for 1 hr at 37 °C in pairs. Following probe treatment, CuAAC click chemistry was performed by sequentially adding 25 pL isoDTB tags (light tags to the lOOpM treated, heavy tags to the l.OrnM treated, (Vector Labs, CCT-1565-L and CCT-1565-H, 10 mM, stock in DMSO), 25 pL tris(2-carboxyethyl)phosphine hydrochloride (TCEP) (50 mM, freshly prepared in water), 75 pL tris[(l-benzyl-l-H-l,2,3-triazol-4-yl)methyl] amine (TBTA) (1.7 mM, 3:1 tert-butanol:DMSO), and 25 pL CuSO4*5H2O (50 mM in water) with mixing between additions. The reaction was incubated at room temperature for 1 hour with gentle mixing. Protein samples were solubilized by adding 200 pL of 10% SDS in PBS, gently mixed and individually sonicated (35 kHz, 2 min) to dissolve any precipitate formed during the click chemistry reaction. Equal amounts of samples were combined in a 2 mL lobind Eppendorf tube and treated with 1 pL of benzonase, followed by incubation at 37°C for 30 minutes at 800 rpm.
[0206] For labeling of live cells, 107Mino cells were concentrated to 3 mL in complete RPML 1640 (Cytiva). Media was then supplemented with l,000X stocks of either probe or vehicle (final concentration 0.1% v / v). Cells were kept under standard conditions (37°C, 5% CO2) in T25 flasks for 1 hour. Following labeling, cells were immediately washed with 3X 10 mL DPBS, concentrated, flash-frozen in LN2, and stored at -80°C overnight. The following morning, cell pellets were then lysed in DPBS and centrifuged (10,000 x g; 10 min) and kept on ice. The soluble fraction was separated, and the remaining precipitated membrane fraction was resuspended in DPBS. Both fractions were quantified via BCA and concentrations of respective channels were then normalized for click reactions and solubilized under the same conditions as mentioned above.
[0207] For protein precipitation and cleanup, SP3 bead purification was employed. 200 pL Sera-Mag SpeedBeads (GE45152105050250, GE65152105050250), were combined in a 1:1 ratio, washed three times with ImL water, and resuspended in 250 pL of water before being added toAtty. Docket: UCSF-865WOthe sample. After incubation at room temperature with shaking at 1000 rpm for 5 minutes, 1 mL of absolute ethanol was added, and the sample was incubated under the same conditions for another 5 minutes. The beads were collected on a magnetic rack, and the supernatant was removed. Beads were washed three times with 1 mL of freshly prepared 80% ethanol.
[0208] To prepare proteins for digestion, the beads were resuspended in 500 pL of fresh 2 M urea in 0.5% SDS in PBS. Proteins were reduced with 10 mM DTT at 65°C for 15 minutes and then alkylated with 20 mM iodoacetamide at 37°C for 30 minutes at 500 rpm. The proteins were re-precipitated with ImL ethanol and subjected to the same washing steps as before.
[0209] Proteins were resuspended in 500 pL of fresh 2 M urea in PBS, and trypsin was added at a 1:100 enzyme-to-protein (w / w) ratio for overnight digestion at 37°C with shaking at 500 rpm. Following digestion, peptides were transferred to a 15mL falcon tube containing lOmL acetonitrile to achieve >95% final volume ACN. Samples were incubated at room temperature with shaking at 1000 rpm for 10 minutes before peptides were collected by centrifugation at 1600 ref for 5 minutes. The pellet was washed three times with 1 mL ACN, then transferred to a 1.5 mL LoBind Eppendorf tube. The samples were placed on a magnetic rack and washed once more in ImL fresh ACN. After magnetic separation, ACN was removed and peptides were eluted twice using 100 pL of 2% DMSO in water, incubated at 37°C for 30 minutes at 1000 rpm.
[0210] For enrichment, 250 pL of high-capacity NeutrA vidin resin (Thermo Fisher, 29204) was washed three times with 10 mL of IAP buffer (50 mM MOPS pH 7.4, 10 mM sodium phosphate, and 50 mM NaCl), then resuspended in a buffer volume adjusted to achieve a final peptide concentration of 0.2 pg / pL. Peptides were incubated with the resin for 2 hours at room temperature on a rotator, then washed twice with PBS and twice with water, with centrifugation at 1600 ref between each wash. The resin was transferred to a bio-spin column (Bio-Rad, 7326204), washed twice with 500 pL of water, and centrifuged at 800 ref for 30 seconds.
[0211] Enriched peptides were eluted in 60 pL of 80% acetonitrile in water containing 0.1% formic acid, incubated for 10 minutes at room temperature, then centrifuged at 1600 ref for 1 minute. This elution was repeated at 72°C for 10 minutes at 300 rpm, and a final elution was performed at room temperature, followed by centrifugation at 2600 ref for 2 minutes. The combined eluates were dried using a SpeedVac.Atty. Docket: UCSF-865WO
[0212] Crude peptides were then resuspended in 22 pL ACN and sonicated (35 kHz, 10 min).425 pL water and 2.3 pL trifluoroacetic acid were then added followed by another sonication. C18 resin (Thermo Fisher, 89870) was activated twice with 200 pL of 50% methanol in water, followed by three washes with 5% ACN in water containing 0.1% TFA. Resin was then placed into 2mL lobind tubes for sample loading. Samples were then added onto the resin over three 150 pL additions. After complete loading, 200 pL of flowthrough was re-added to the resin twice to ensure complete binding. Samples were then washed three times as mentioned before. All centrifugation steps were performed at 300 ref for 15 seconds. Samples were then eluted twice with 20 pL 70% ACN in water containing 0.1% formic acid after a 1 minute incubation (1600 ref, 1 min). A third elution was performed similarly, but at 2600 ref for 2 minutes. Eluted, desalted peptides were then dried using a SpeedVac.
[0213] Samples were resuspended in 20 pL LC-LOAD (PreOmics) followed by vortexing and sonication (35 kHz, 10 min). Samples were centrifuged (21000 ref, 10 min) and 15 pL were transferred into MS vials (Thermo Fisher, 200046) for analysis by mass spectrometry.Mass Spectrometry Data Acquisition
[0214] A nanoElute2 was attached in line to a timsTOF Pro2 equipped with a CaptiveSpray Source (Bruker, Hamburg, Germany). Chromatography was conducted at 40°C through a 25 cm reversed-phase C18 column (PepSep, 1893476) at a constant flowrate of 0.5 pL min-1. Mobile phase A was 98 / 2 / 0.1% water / MeCN / formic acid (v / v / v, Thermo Fisher, LS118, LS120) and phase B was MeCN with 0.1% formic acid (v / v). During a 120 min method, peptides were separated by a 3-step linear gradient (4% to 40% B over 95 min, 40% to 60% B over 10 min, 60% to 90% B over 10 min) followed by a 10 min isocratic flush at 95% before washing and a return to low organic conditions. Experiments were run as data-dependent acquisitions with ion mobility activated in PASEF mode. MS and MS / MS spectra were collected with m / z 100 to 1700 and ions with z = +1 were excluded.FragPipe Open Search
[0215] To comprehensively survey mass shifts present in peptides across the dataset, an Open Search was performed using MSFragger within FragPipe v22.0. The search was configured with the following parameters: precursor mass range -10 to +850 Da, fragment mass tolerance 40 amu, isotope error set to 0, and enzyme set to “trypsin” with cleavage after “KR” and not beforeAtty. Docket: UCSF-865WO“P” (C-terminal specificity). Enzymatic digestion was assumed, allowing up to 2 missed cleavages. Minimum and maximum peptide lengths were set to 7 and 50 residues, respectively, and peptide mass range from 500 to 5000 Da. N-terminal methionine clipping was enabled, and a minimum of 15 peaks and a 0.01 intensity ratio were required per spectrum. Fragment ions considered included b and y series, with a maximum fragment charge of 2+. Calibration and parameter optimization were disabled, and deisotoping and de-neutralloss filtering were enabled. No fixed modifications were applied, and variable modifications were included but not enforced in the initial search (used only for localization and PTM profiling downstream). Pep tideProphet was enabled with the following command-line options: -nonparam -expectscore -decoyprobs --masswidth 1000.0 -clevel -2, and results from multiple replicates were combined. PTMProphet was disabled. ProteinProphet was run with the setting -maxppmdiff 2000000. PTM-Shepherd was enabled for post-processing and mass shift profiling. Peak picking parameters were set to a smoothing factor of 2 bins, prominence ratio of 0.3, and peak picking width of 0.002 Da.Precursor mass tolerance was 0.01 Da. Fragment ion types considered for modification annotation were b and y, with a maximum fragment charge of 2+. PTM-Shepherd was run with default UniMod annotations disabled to allow unsupervised delta mass discovery, and glycan annotation was enabled but no glycan database was used. Glycan delta masses were removed before peak picking. Normalization of mass shifts was performed across PSMs (not scans). Report generation was enabled with the following options: -sequential — mapmods -prot 0.01. MS 1-based quantification, TMT-Integrator, and Spectral Eibrary generation were disabled. For downstream analysis, the global.modsummary.tsv file from PTM-Shepherd was used to extract the number of PSMs associated with each observed mass shift. These counts were plotted against the theoretical mass shift to visualize the landscape of modifications in the range of +200 to +750 Da.FragPipe Closed Search
[0216] Mass spectrometry data were analyzed using FragPipe v22.0. Closed Search was performed using MSFragger with the following parameters: precursor mass tolerance -50 to +50 ppm, fragment mass tolerance 50 mmu, isotope error set to 0 / 1 / 2, enzyme name “stricttrypsin”, cleavage at “KR” but not before “P” (C-terminal cleavage), enzymatic cleavage with up to 2 missed cleavages, peptide length between 7 and 50 residues, and peptide mass range from 500 to 5000 Da. Fragment ion types included b and y ions, with a maximum fragment charge of 2+. N-Atty. Docket: UCSF-865WOterminal methionine clipping was enabled. The number of variable modifications per peptide was limited to 4, and the number of allowed variable modification combinations was capped at 5000. Spectra were required to have a minimum of 15 peaks and a minimum peak intensity ratio of 0.01. Mass calibration was set to “Mass Calibration and Parameter Optimization,” and deisotoping and de-neutralloss processing were enabled.
[0217] Variable modifications included oxidation of methionine (+15.9949 Da), N-terminal acetylation (+42.0106 Da), and a pair of arginine labels (+695.2942 Da and +701.3014 Da) each with a maximum of 1 occurrence. Cysteine was either unmodified or labeled with a variable modification of +57.02146 Da (up to 3 occurrences). Fixed modifications were not applied except in specific runs evaluating cysteine, where carbamidomethylation (+57.02146 Da) was fixed on C.
[0218] Crystal-C, PTMProphet, and TMT-Integrator were disabled. PeptideProphet was also disabled. ProteinProphet was enabled and ran with the following setting: -maxppmdiff 2000000. Percolator was used for peptide- spectrum match (PSM) validation with the following commandline options: -only-psms -no-terminate -post-processing-tdc, and a minimum probability threshold of 0.5.
[0219] MS 1-based quantification was enabled using lonQuant with the following parameters: m / z tolerance of 10 ppm, retention time (RT) tolerance of 0.4 min, ion mobility tolerance 0.05 (although ion mobility was not used), labeling mode enabled with two arginine labels (+695.2942 Da and +701.3014 Da), re-quantification enabled, minimum frequency of 0.5, a minimum of 2 isotopes and 1 scan per feature, top 3 ions used for quantification, and no normalization applied. Match-between-runs was disabled. Label-free quantification was not used.
[0220] MSBooster, PTMShepherd, SAINTexpress, Skyline, and Spectral Library generation were disabled. Report generation was enabled with the following options: -sequential -razor - mapmods -prot 0.01, with decoys and known contaminants removed. Summary tables were generated at both peptide and protein levels.Atty. Docket: UCSF-865WOSASA Analysis Pipeline
[0221] Output modified peptides from Fragpipe were processed using a custom Python-based pipeline to identify modified residues based on their sequence position in UniProt. Protein identifiers were initially provided as UniProt names, which were converted to their corresponding UniProt accession numbers using the UniProt API (https: / / rest.uniprot.org / ). Structural models were then downloaded from the European Bioinformatics Institute (EBI) AlphaFold Database (https: / / alphafold.ebi.ac.uk / ), selecting the highest-confidence monomeric model (version 4). Each protein structure was parsed using Bio.PDB (Biopython), and the Shrake-Rupley algorithm was applied to compute solvent-accessible surface area (SASA) at the atomic level. Given potential indexing differences between UniProt annotations and AlphaFold structures, residue positions were adjusted by shifting input values down by one to ensure alignment with AlphaFold numbering. To ensure the reliability of structural predictions, residues with a predicted local distance difference test (pLDDT) score below 70 were excluded from the analysis. The relative solvent accessibility (RSA) was then calculated by normalizing SASA values against a reference maximum exposure of 274 A2. Computational storage was optimized by removing downloaded PDB files immediately following analysis, and results were iteratively written to an output file. This automated workflow enabled high-throughput extraction of solvent accessibility data from AlphaFold models, facilitating the structural interpretation of specific residues within predicted protein conformations.
[0222] For experimental structures, PDB accessions were retrieved from UniProt cross-references and corresponding mmCIF files were downloaded using Biopython’s MMCIFParser. Structures with ribosomal content or more than 10 chains were excluded. Each PDB structure was aligned to the UniProt sequence to correctly map the modified residue to the structural coordinates. SASA values were computed using the same Shrake-Rupley algorithm, and RSA values were calculated using the same normalization scheme. In cases where multiple structures existed for a given protein, all qualifying entries were processed independently, and the mean and standard deviation of RSA values for each residue were calculated and used in downstream analysis.Atty. Docket: UCSF-865WOPDB Proximal Ligand Search Pipeline
[0223] To assess the structural context of modified arginine residues in proteins, an in-house computational pipeline was developed to map modified arginines from the Fragpipe output onto full-length UniProt sequences, align them to resolved PDB structures, and evaluate their spatial proximity to bound ligands. The input dataset consisted of a CSV file containing UniProt accession numbers and peptide sequences with modified arginine residues denoted asR[701.3014], Other modifications annotated in brackets, such as oxidation or carbamidomethylation, were ignored. The modified residue’s position within the peptide was extracted, and all bracketed modifications were removed to generate a cleaned peptide sequence for further processing. To determine the position of the modified arginine in the full-length protein sequence, the cleaned peptide was aligned to the UniProt sequence using Biopython’s PairwiseAligner with the BLOSUM62 substitution matrix, and affine gap penalties (-10 for opening, -0.5 for extension). The modified residue’s position in the peptide was mapped to its corresponding position in the aligned UniProt sequence, yielding the residue number in the full-length protein. The script then queried the UniProt cross-references to retrieve all associated PDB accession numbers for each protein. Structures were downloaded in the PDBx / mmCIF format using MMCIFParser from Biopython. To improve computational efficiency and ensure relevance to the analysis, structures containing ribosomal components were excluded, and entries with more than 10 chains were filtered out. For structural mapping, the UniProt sequence was aligned to the sequences of all polypeptide chains in each PDB structure, allowing for potential discrepancies in residue numbering between sequence databases and structural models. The modified residue’s position in the PDB structure was identified based on this alignment, ensuring that subsequent analyses were performed on the correct structural residue. To assess the molecular environment of modified arginines, the spatial relationship between the CZ atom of the modified residue and bound ligands was analyzed. NeighborSearch from Biopython was used to determine whether any ligand atoms were within 6.0 A of the arginine CZ atom. Water molecules (residue name ‘HOFF) were excluded to focus the analysis on small molecules and cofactors. If a ligand met the distance threshold, its residue name and chain position were recorded. To ensure robustness in handling large datasets, results were written incrementally to a CSV file containing the UniProt ID, PDB ID, chain identifier, mapped residue position in the PDB, and the nearest ligand within the specified distance.Atty. Docket: UCSF-865WOAlphaMissense Pathogenicity Annotation
[0224] To evaluate the potential functional relevance of modified arginines, AlphaMissense pathogenicity scores were incorporated into the dataset. UniProt accession numbers and residue positions for all modified arginines were used to generate a queryable index. A custom Python script was developed to query the AlphaMissense API for each protein and residue pair using concurrent requests with up to 10 worker threads to accelerate the process. Isoform-specific identifiers (e.g., P12345-2) were excluded to ensure consistency with canonical AlphaMissense annotations. For each query, the API returned a predicted mean pathogenicity score and a classification (e.g.. 'likely pathogenic', 'likely benign'). If a classification was not explicitly returned, a custom threshold-based scheme was applied: scores < 0.34 were labeled 'likely benign', scores > 0.564 were labeled 'likely pathogenic', and intermediate values were designated 'ambiguous'. All results were written to a CSV file including the protein ID, residue position, score, and classification.Secondary Structure Assignment via DSSP
[0225] To assess local structural context of modified residues, DSSP-based secondary structure annotations were computed for both AlphaFold-predicted and experimentally resolved PDB structures. For AlphaFold models, UniProt IDs were first mapped to primary accessions using the UniProt API. Structural models were downloaded from the AlphaFold Protein Structure Database. For each protein, the corresponding PDB file was parsed using Biopython’s PDBParser, and DSSP secondary structure assignments were computed using Bio.PDB.DSSP. Modified residue positions were matched to AlphaFold residue indices (adjusted for 0-based numbering), and residues with a predicted local distance difference test (pLDDT) score below 75 were excluded. For residues passing the pLDDT filter, a ±4 residue sliding window was used to collect local secondary structure context. Both the DSSP code for the modified residue and its surrounding window were recorded.
[0226] For experimental structures, UniProt cross-references were used to identify relevant PDB accessions. Structures were downloaded in mmCIF format and parsed using Biopython’s MMCIFParser. Resolution and structural complexity filters were applied: only structures with <10 chains and <3.0 A resolution were retained. The best-aligned chain was selected by global pairwise sequence alignment using Biopython’s PairwiseAligner, and modified residues wereAtty. Docket: UCSF-865WOmapped from UniProt positions to structure-specific indices. DSSP was run on each selected PDB model to obtain secondary structure codes. For each modified residue, the DSSP code and ±3 residue sequence context were recorded. Results were aggregated across all qualifying structures for each protein, and the most frequently observed DSSP code (majority vote) was assigned as the representative secondary structure for that residue. Confidence scores were calculated as the frequency of the major assignment relative to total observations.Motif Enrichment Analysis via pLogo
[0227] To investigate sequence preferences surrounding modified arginines, motif enrichment analysis was performed using pLogo (https: / / plogo.uconn.edu). Foreground sequences were generated from a list of experimentally identified arginine modification sites, with UniProt accession numbers and residue positions extracted from FragPipe output. A custom Python script queried the UniProt API to retrieve the full protein sequences and extracted a 21 -residue window (±10 amino acids) centered on each modified arginine. If a full window could not be extracted due to proximity to sequence termini, the missing positions were padded with 'X' characters to maintain fixed length. Background sequences were generated separately from the full UniProt human proteome FASTA (including isoforms). For each protein sequence, a sliding window approach was used to extract 21 -residue sequences centered on every amino acid position, again padding with 'X' as needed. This resulted in an unfiltered and comprehensive background set. Due to file size limitations on the pLogo web interface, the background set was downsampled to 1,000,000 sequences using Python’s random.sample() method. Foreground and background sets were submitted to the pLogo web interface, specifying arginine (R) as the central anchor residue. A binomial statistical model with Bonferroni correction was used to evaluate motif significance. Enrichment scores and residue logos were downloaded and replotted using custom scripts to standardize styling across figures.Salt Bridge Identification in AlphaFold Structures
[0228] To identify potential salt bridge interactions between modified arginines and acidic residues, AlphaFold-predicted structural models were analyzed using a custom Python pipeline. UniProt IDs and modified residue positions were extracted from a CSV input file. Isoform identifiers were excluded. For each canonical UniProt ID, the AlphaFold API was queried to retrieve the corresponding structure metadata and download the associated PDB file. ModelsAtty. Docket: UCSF-865WOwere stored locally and parsed using Biopython's PDBParser. Each arginine residue was scanned for proximity to nearby aspartate (ASP) or glutamate (GLU) side chain atoms (0D1 / 0D2 or 0E1 / 0E2, respectively). Only models with pLDDT scores > 70 for both arginine and acidic residues were considered. Salt bridges were defined by the presence of at least one NE, NH1, or NH2 atom from the modified arginine within 4.0 A of the target carboxylate atom. When multiple atoms satisfied the distance criterion, the closest atom pair was selected. Electrostatic interaction energies were estimated using a simplified Coulombic potential: , assuming unit charges, a dielectric constant of 20.0, and interatomic distance in A. For each interaction, the chain ID, residue names and numbers, interacting atoms, interatomic distance, pLDDT scores, and computed Coulombic energy were recorded. Following analysis, downloaded PDB files were removed to conserve disk space. Final results were compiled into a CSV file for downstream analysis and visualization.Cation-n Interaction Identification in AlphaFold Structures
[0229] To identify potential cation-7t interactions involving modified arginines, AlphaFold-predicted structures were evaluated using a geometry-based Python workflow. For each UniProt ID and arginine position, the corresponding AlphaFold model was downloaded from the EBI database. PDB files were parsed with Biopython’s PDBParser, and residues were scanned for interactions between arginine CZ atoms and aromatic residues (PHE, TYR, TRP, HIS). For each aromatic residue, centroid coordinates were computed from annotated ring atoms (e.g., CZ, CE1, CE2, etc.), and the Euclidean distance to the arginine CZ atom was calculated. Interactions within 6.0 A were considered candidates. If angular filtering was enabled, the orientation between the CZ-centroid vector and the ring plane normal was used to classify geometry as parallel (face-to-face) or T-shaped (edge-to-face), with thresholds of <30° and 90°+20°, respectively. Models with pLDDT scores below 70 for either the arginine or aromatic residue were excluded. For each identified interaction, the protein ID, chain IDs, residue positions and types, distance, angle, geometry classification, and pLDDT scores were recorded. Only one interaction was logged per arginine to avoid multiple counting. Results were written to two CSV outputs: a detailed interaction table and a simplified summary indicating whether each modified arginine participates in a cation-jr interaction.Atty. Docket: UCSF-865WOSalt Bridge and Cation-n Interaction Analysis in Experimental PDB Structures
[0230] To validate and expand upon AlphaFold-based interaction predictions, salt bridges and cation-7t interactions were also identified using experimentally resolved PDB structures. For each UniProt ID and modified residue position, all associated PDB structures were retrieved via UniProt cross-references. Structures were excluded if they represented ribosomal proteins, large macromolecular complexes (taxonomy count > 15), NMR structures, or models containing more than 10 chains. PDBx / mmCIF files were downloaded and parsed using Biopython’s MMCIFParser. For each candidate arginine, neighboring residues were identified using NeighborSearch, and interactions were classified as salt bridges or cation-7t based on geometric and electrostatic criteria. Salt bridges were defined as any acidic residue (ASP, GLU) within 4.0 A of NE, NH1, or NH2 atoms of the arginine. Cation-K interactions were defined as cases where the arginine CZ atom was within 8.0 A of the centroid of an aromatic side chain (PHE, TYR, TRP, HIS). When enabled, angular filtering was used to categorize interaction geometries as face-to-face (<30°) or T-shaped (90° ± 20°). All interactions were annotated with residue identity, chain location, distance, geometry (if applicable), whether the partner residue was part of the same polypeptide chain, and optionally, an estimated electrostatic energy. Final results were compiled into a single CSV table containing all identified intermolecular interactions from experimental PDB models.pKa Estimation for Arginines Using Experimental PDB Structures
[0231] To estimate side-chain protonation states of modified arginines, PROPKA 3.0 was used to calculate predicted pKa values from experimental structures in the Protein Data Bank (PDB). Modified arginines were identified by UniProt accession and residue position. Isoforms and ribosomal proteins were excluded. Associated PDB structures were retrieved via UniProt cross-references and filtered to exclude NMR structures, entries with >10 polymer entities, or structures with >10 chains. Structures were downloaded in mmCIF format and parsed using Biopython’s MMCIFParser. Each structure was aligned to the canonical UniProt sequence using PairwiseAligner to identify the best-matching chain and map UniProt residue indices to structural coordinates. Only arginines with a resolved CZ atom and a B-factor > 30.0 were considered. Structures were converted to PDB format using Biopython’s PDBIO to ensure compatibility with PROPKA. The PROPKA executable was run via subprocess call, and predicted pKa values were parsed from the resulting .pka files. For each modified arginine, theAtty. Docket: UCSF-865WOoutput included UniProt TD, PDB ID, chain ID, UniProt and author residue numbers, and pKa value. Structures and temporary output files were deleted after use to conserve disk space.pKa Estimation for Arginine, Lysine or Cysteine from AlphaFold Models
[0232] To complement structure-based pKa estimation, PROPKA was also applied to AlphaFold-predicted structures. AlphaFold models were downloaded in PDB format for each UniProt accession. Only residues with pLDDT confidence scores >70 were considered for analysis. For each AlphaFold structure, PROPKA 3.0 was executed using the PDB file as input. Predicted pKa values were extracted using a custom parser that accommodated multiple possible output formats. Residue-level pKa values were reported for each modified arginine, cysteine or lysine, alongside the corresponding pLDDT score. Proteins for which PROPKA failed to produce valid output were excluded. Temporary AlphaFold PDB and PROPKA output files were deleted after processing.Intrinsic Disorder Prediction via IUPred2A
[0233] To evaluate the intrinsic disorder propensity of modified arginine residues, disorder scores were retrieved using the IUPred2A web API. A CSV file containing UniProt IDs and modified residue positions served as input. For each protein, the full-length disorder profile was fetched in JSON format and parsed to extract per-residue IUPred2 scores. Modified arginine positions were indexed using 1 -based UniProt residue numbering. Prior to extraction, each index was converted to 0-based to match Python list conventions. If a position fell within the length of the retrieved disorder profile, the corresponding IUPred2 score was recorded. Otherwise, 'NA' was assigned. Each entry in the output file included the UniProt ID, residue position, and corresponding disorder score. Results were written iteratively to a CSV file for downstream integration with structural and reactivity data.Molecular Docking and Covalent Adduct Modeling
[0234] Molecular docking was performed using the Molecular Operating Environment (MOE2024). The crystal structures of aconitase (PDB: 1B0J) and cyclophilin A (PDB: 6GJI) were obtained, with all crystallographic ions removed. Each structure was solvated in a 10 A water sphere to mimic physiological conditions before preparation using MOE’s QuickPrep feature, which optimized hydrogen placement, assigned protonation states, and performed an initial energy minimization. For aconitase, the docking site was defined based on the knownAtty. Docket: UCSF-865WObinding site of isocitrate, and the dehydrate of ninhydrin was docked using MOE’s docking protocol. For cyclophilin A, the ligand 4'CA-Alk was docked using MOE’s template-based docking to ensure proper positioning. In both cases, the top-scoring pose was selected and loaded into the MOE Window, where a covalent adduct was modeled. Energy minimization was performed using the AmberlO:EHT force field with gradient-based optimization until the root mean square gradient reached 0.05 kcal / (mol- A). Visualization and figure preparation were performed using ChimeraX 1.6.1.Cyclophilin A Inhibition Assay
[0235] To measure CA-Nin-Alk (4-CA-Alk) inhibition of cyclophilin A (CypA) activity, a peptidyl-proline isomerase (PPIase) assay was utilized following a previously established protocol54. Recombinant human CypA (R&D Systems) was diluted to a concentration of 1 pM in a reaction buffer of 20 mM HEPES, 100 mM NaCl, and 0.5 mM TCEP (pH 7.5) and kept on ice. Chymotrypsin was diluted to 60 mg / mL in 2 mM CaCh, 1 mM HC1 in H2O. To test inhibition, the inhibitor, CA-Nin-Alk (4-CA-Alk), was diluted in DMSO to generate 100X stock solutions (25 pM, 10 pM). These were added to CypA solutions (final DMSO = 1% v / v) and incubated for 1 hour at room temperature. The substrate, suc-AAPF-pNA (Thomas Scientific), was freshly prepared at 10 mM in 470 mM LiCl / TFE. The spectrophotometer (Cary UV-Vis Spectrophotometer) was cooled to 4°C before use. For uncatalyzed reactions, 980 pL of reaction buffer and 10 pL of chymotrypsin were mixed in a cuvette and placed in the spectrophotometer. The reaction was initiated with 10 pL substrate and A390 was monitored every 0.1s for 5 minutes. For CypA catalyzed reactions, 970 pL of the reaction buffer, 10 pL of chymotrypsin, and 10 pL of CypA solution (pre-incubated ± inhibitor) were combined prior to substrate addition and identical data collection. All assays were performed in triplicate. To ensure calculations depended solely on the first-order reaction of isomerization and not residual transsubstrate in the solution, the absorbance traces, excluding the first 10s, were fit to a singleexponential function to obtain the rate constants for uncatalyzed (kuncat) or CypA-catalyzed (kObs isomerization. kcat / Kmvalues were then calculated via the following equation:‘ k'■catnk'obs — ‘ k'-uncatKM[CypA]using a final enzyme concentration of l*10‘8M.Atty. Docket: UCSF-865WOICso calculations
[0236] Apparent IC50 values were estimated using a two-point linear interpolation. Percent inhibition of CypA catalytic efficiency kcat / Km) was measured at two compound concentrations (25 pM, 10 pM). A linear regression of inhibition (y) vs compound concentration ( ) was used to approximate the relationship (y = mx + b), and the concentration corresponding to 50% inhibition was calculated as:- <50“ VmEXAMPLE 2Synthesis of probes
[0237] To a round-bottom flask containing K2CO3 (1.87 g, 13.50 mmol) and 5-hydroxy-2,3-dihydro-lH-inden-l-one (1.000 g, 6.75 mmol) was added anhydrous DMF (13.5 mL). Propargyl bromide (80wt% in toluene, 829.5 pL, 1.10 g, 7.42 mmol) was added dropwise.The reaction mixture was stirred at 60 °C for 16 h. After cooling to room temperature, the mixture was diluted with CH2CI2 (30 mL) and washed with water (3 x 30 mL) followed by brine (2 x 30 mL). The organic layer was dried over Na2SO4, filtered, and concentrated under reduced pressure to afford 5-(prop-2-yn-l-yloxy)-2,3-dihydro-lH-inden-l-one as a brown solid, which was used in subsequent steps without further purification (1.215 g, 96.7% yield).
[0238] ‘H NMR (400 MHz, CDCh) 5: 7.71 (d, J = 8.4 Hz, 1H), 7.00 (s. 1H), 6.97 (d, J = 8.7 Hz, 1H), 4.77 (t, 7= 2.1 Hz, 2H), 3.11 (t, J = 5.4 Hz, 2H), 2.69 (t, 7= 5.9 Hz, 2H), 2.57 (s, 1H).13C NMR (101 MHz, CDCh) 8: 205.41, 163.06, 158.06, 131.22, 125.50, 115.88, 111.00, 76.35, 56.05, 36.53. 25.99. HRMS (ESP): m / z calcd for Ci2HuO2+([M+H]+): 187.0759; found:187.0756.
[0239] To a solution of 5-(prop-2-yn-l-yloxy)-2,3-dihydro- IH-inden-l-one (500 mg, 2.69 mmol, 1.0 equiv) in 1,4-dioxane (3.25 mL) and water (325 pL) was added selenium dioxide (745 mg, 6.71 mmol, 2.5 equiv). The reaction mixture was heated toreflux (90 °C) and stirred overnight. After cooling to roomAtty. Docket: UCSF-865WOtemperature, the mixture was diluted with ethyl acetate and washed successively with water (2 x 10 mL), 1 M HC1 (2 x 10 mL), and saturated aqueous NaHCCL (2 x 10 mL). The organic layer was washed with brine (1 x 10 mL), dried over Na2SC>4, filtered, and concentrated under reduced pressure. The crude product was purified by flash chromatography (30-50% acetone in hexanes) to afford 2,2-dihydroxy-5-(prop-2-yn-l-yloxy)-lH-indene-l,3(2H)-dione as a red solid. Final purification was performed by reverse-phase HPLC (2-25% MeCN with 0.1% TFA in FLO with 0.1% TFA), and lyophilization of the collected fractions yielded the product as a white powder (561 mg, 2.42 mmol, 90.1%).
[0240] ‘H NMR (400 MHz, CD3CN / D2O) 8: 7.95 (d, J = 8.4 Hz, 1H). 7.51 (d, J = 8.6 Hz. 1H), 7.45 (s, 1H), 4.91 (s, 2H), 2.11 (s, 1H). HRMS (ESI+): m / z calcd for C12H9OC ([M+H]+):233.0450; found: 233.0448.G
[0241] To a round-bottom flask containing K2CO3 (1.87 g, 13.50 mmol) and 4-hydroxy-2,3-dihydro-lH-inden-l-one (1.000 g, 6.75 mmol) was L added anhydrous DMF (13.5 mL). Propargyl bromide (80 wt% in toluene,829.5 pL, 1.10 g, 7.42 mmol) was added dropwise. The reaction mixturewas stirred at 60 °C for 16 h. After cooling to room temperature, the mixture was diluted with CH2CI2 (30 mL) and washed with water (3 x 30 mL) followed by brine (2 x 30 mL). The organic layer was dried over Na2SO4, filtered, and concentrated under reduced pressure to afford 4-(prop-2-yn-l-yloxy)-2,3-dihydro-lH-inden-l-one as a tan solid, which was used in subsequent steps without further purification.
[0242] rH NMR (400 MHz, CDCh) 8: 7.40 (d, J = 7.6 Hz, 1H), 7.36 (t, J= 7.6 Hz, 1H), 7.17 (d, J = 7.7 Hz, 1H), 4.80 (s. 2H). 3.07 (t, J= 6.0 Hz, 2H), 2.69 (t, J= 6.0 Hz. 2H). 2.55 (s, 1H).13C NMR (101 MHz, CDCh) 8: 207.23, 155.10, 144.57, 139.04, 128.86, 116.50, 116.45, 78.18, 76.22, 56.04, 36.26, 22.70. HRMS (ESP): m / z calcd for Ci2HnO2+([M+H]+): 187.0759; found: 187.0757.
[0243] To a solution of 4-(prop-2-yn-l-yloxy)-2,3-dihydro-lH- inden-l-one (500 mg, 2.69 mmol, 1.0 equiv) in 1,4-dioxane (3.25 mL) .OHOH and water (325 pL) was added selenium dioxide (745 mg, 6.71 mmol. 2.5equiv). The reaction mixture was heated to reflux (90 °C) and stirred overnight. After cooling to room temperature, the mixture was dilutedAtty. Docket: UCSF-865WOwith ethyl acetate and washed successively with water (2 x 10 mL), 1 M HO (2 x 10 mL), and saturated aqueous NaHCOs (2 x 10 mL). The organic layer was washed with brine (1 x 10 mL), dried over Na2SO4. filtered, and concentrated under reduced pressure. The crude product was purified by flash chromatography (30-50% acetone in hexanes) to yield 2,2-dihydroxy-4-(prop-2-yn-l-yloxy)-lH-indene-L3(2H)-dione as a red solid. Final purification was performed by reverse-phase HPLC (2-25% MeCN with 0.1% TFA in FLO with 0.1% TFA) to remove selenic acid byproducts yielding a white powder after lyophilization (532 mg, 2.29 mmol, 85.3%).
[0244] ‘H NMR (400 MHz, DMSO-de) 8: 7.95 (q, J= 8.3 Hz, 1H), 7.62 (d, J= 8.4 Hz, 1H), 7.52 (d, J = 7.5 Hz, 1H), 7.40 (s, 2H), 5.05 (d, 7= 2.2 Hz, 2H), 3.68 (d. J= 2.9 Hz, 1H). HRMS (ESP): m / zcalcd for CI2H9O5+([M+H]+): 233.0450; found: 233.0446.NH
[0245] A solution of 4-(trifluoromethyl)benzaldehyde (1.741 g, 10.0 mmol), propargylamine (826.2 mg, 15.0 mmol), sodium triacetoxyborohydride (5.299 g, 25.0 mmol), and glacial acetic acid (600.5 mg, 10.0 mmol) in CH2CI2 (total volume 93.8 mL) was stirred at room temperature for 2.5 h. The reaction was diluted with CH2CI2 (40 mL) and saturated aqueous NaHCCL (40 mL), and the layers were separated. The aqueous layer was extracted with CELCL (2 x 20 mL), and the combined organic layers were concentrated under reduced pressure. The residue was dissolved in Et2O (50 mL) and extracted with 1.0 M HC1 (3 x 30 mL). The combined aqueous extracts were basified to pH ~ 11-12 with 1.0 M NaOH and extracted with Et2O (4 x 20 mL). The combined organic layers were dried over Na2SO4 and concentrated. Purification by flash chromatography (EtOAc / hexanes) afforded N-(4-(trifluoromethyl)benzyl)prop-2-yn-l -amine (2.05 g, 9.62 mmol, 96.2%) as an orange oil.
[0246] JH NMR (CHLOROFORM-D) 8: 7.59 (d, J = 7.9 Hz, 2H), 7.48 (d, J = 7.9 Hz, 2H), 3.95 (s, 2H), 3.43 (s, 2H), 2.27 (s, 1H).13C NMR (CHLOROFORM- / ?) 8: 143.58, 129.73, 129.43, 128.67, 125.42 , 81.78, 71.94, 51.68, 37.41.19F (CHLOROFORM- / ?) 8: -62.31. Calculated [M+H] - 214.0844, found [M+H] - 214.0842.Atty. Docket: UCSF-865WOr ok Jk. -NH,
[0247] To a solution of N-(4-(trifhioromethyl)benzyl)prop-2-yn-l-amine (365 mg. 1.71 mmol) in DMF (8.56 mL) was added DIPEA (664 mg, 895 pL, 5.14 mmol), followed by (tert-butoxycarbonyl)glycine (330 mg, 1.88 mmol) and HATU (976 mg, 2.57 mmol). The mixture was briefly sonicated and stirred at ambient temperature overnight. The reaction was diluted with EtOAc (50 mL) and sequentially washed with 1.0 M HC1 (2x), FLO (2x), saturated NaHCOa (2x), and brine (2x). The organic layer was dried over Na2SO4, filtered, and concentrated under reduced pressure to afford the Boc-protected intermediate. This material was dissolved in 4 M TFA in dioxane and stirred for 3 h at room temperature to yield the TFA salt of the product as a light purple foam. For analytical characterization, the TFA salt was basified with saturated aqueous NaHCOa and extracted with EtOAc. The combined organic extracts were dried, filtered, and concentrated to afford 2-amino-N-(prop-2-yn-l-yl)-N-(4-(trifluoromethyl)benzyl)acetamide (392.7 mg, 1.453 mmol, 84.9%) as a tan oil.
[0248] XH NMR (CHLOROFORM-D) 8: 7.59 (dd, J = 19.2, 7.9 Hz, 2H), 7.34 (dd. J = 21.9, 8.0 Hz, 3H), 4.69 (d, J = 37.0 Hz, 2H), 4.24 (s, 1H), 3.88 (s, 1H), 3.63 (s, 1H), 3.48 (s, 1H), 2.26 (d, J = 26.6 Hz, 1H), 1.72 (bs, 2H).19F (CHLOROFORM -D) 8: -62.47.
[0249] To a solution of 2-amino-N-(prop-2-yn-l-yl)-N-(4-(trifluoromethyl)benzyl)acetamide (100 mg, 370 pmol) in DMF was added DIPEA (191 mg, 266 pL, 1.48 mmol), followed by 1-oxo-2, 3-dihydro-lH-indene-4-carboxylic acid (71.7 mg, 407 pmol) and HATU (211 mg, 555 pmol). The reaction was stirred overnight at room temperature. The reaction was diluted with EtOAc (50 mL) and sequentially washed with 1.0 M HC1 (2x), H2O (2x), saturated NaHCOa (2x), and brine (2x). The organic phase was dried over Na2SO4, filtered, and concentrated under reduced pressure. Purification by flash chromatography (silica gel, gradient 40-60% EtOAc inAtty. Docket: UCSF-865WOhexanes) afforded the product as an off-white solid (137.7 mg, 86.9%). R_f = 0.31 (60% EtOAc / hexanes, TLC).
[0250] NMR (CHLOROFORM-D) 8: 7.97 (t, J = 8.5 Hz, 1H), 7.91 (d, J = 7.7 Hz, 1H), 7.67 (d, J = 8.0 Hz, 1H), 7.62 (d. J = 7.9 Hz, 1H), 7.49 (t, J = 7.8 Hz, 1H), 7.45 - 7.37 (m, 2H), 7.21 (s, 1H), 4.79 (d, J - 16.1 Hz, 2H), 4.48 (d, J = 4.0 Hz, 1H), 4.36 (d, J - 4.0 Hz, 1H), 4.29 (s, 1H), 4.02 (s, 1H), 3.53 - 3.42 (m, 2H), 2.75 (q, J = 5.4 Hz, 2H), 2.34 (d, J = 34.4 Hz, 1H).13C NMR (CHLOROFORM- / )) 8: 206.36, 168.39, 166.61, 154.03, 139.83, 138.27, 133.05, 132.28, 128.48, 127.71, 127.05, 126.61, 126.15, 125.73, 74.18, 73.19, 49.00, 41.86, 36.03, 35.00, 26.21.19F (CHLOROFORM- ) 8: -62.48
[0251] To a solution of 2-amino-N-(prop-2-yn-l-yl)-N-(4-(trifluoromethyl)benzyl)acetamide (60.0 mg, 222 pmol) in DMF was added DIPEA (57.4 mg, 77.3 pL, 444 pmol), followed by 1-oxo-2, 3-dihydro-lH-indene-5-carboxylic acid (43.0 mg, 244 pmol) and HATH (127 mg, 333 pmol). The reaction mixture was stirred at ambient temperature overnight. The reaction was diluted with EtOAc (50 mL) and sequentially washed with 1.0 M HC1 (2x), H2O (2x), saturated NaHCCb (2x), and brine (2x). The organic layer was dried over Na2SC>4, filtered, and concentrated under reduced pressure. The crude product was purified by flash column chromatography (silica gel, gradient 40-60% EtOAc in hexanes) to afford the product as a purple / tan solid (87.6 mg, 92.1%). R_f = 0.29 (60% EtOAc / hexanes, TLC).
[0252] 1H NMR (CHLOROFORM-D) 8: 7.95 (d. J = 11.4 Hz. 1H). 7.82 (s, 1H). 7.80 (s, 1H), 7.66 (d, J = 7.9 Hz, 1H), 7.61 (d, J - 7.9 Hz, 1H), 7.39 (dd, J - 14.1, 7.9 Hz, 2H), 7.32 (s, 1H), 4.77 (d, J = 16.7 Hz, 2H), 4.45 (d, J = 3.9 Hz, 1H), 4.33 (d. J = 3.9 Hz, 1H), 4.28 (s, 1H), 4.01 (s, 1H), 3.24 - 3.16 (m, 2H), 2.75 (t, J = 6.0 Hz, 2H), 2.33 (d, J = 33.2 Hz, 1H).13C NMR (CHLOROFORM-D) 8: 206.50, 168.53, 166.79, 154.20, 138.46, 133.23, 128.67, 127.90, 127.21, 126.81, 126.38, 125.92, 74.35, 73.36, 49.16, 41.98, 36.29, 35.18, 26.38.19F (CHLOROFORM- ) 8: -62.53Atty. Docket: UCSF-865WO2,2-Dihydroxy-l,3-dioxo-N-(2-oxo-2-(prop-2-yn-l-yl(4-(trifluoromethyl)benzyI)amino)ethyl)-2,3-dihydro-lH-indene-carboxamides
[0253] To a solution of the corresponding starting material (1.0 equiv) in l,4-dioxane / H2O (10:1 v / v, 200 mM final concentration) was added selenium dioxide (2.1-2.2 equiv). The reaction mixture was heated to reflux (~96 °C) and stirred overnight. After cooling to room temperature, the mixture was diluted with ethyl acetate and washed successively with water (2 x), 1 M aqueous HC1 (2 x), and saturated aqueous NaHCCT (2 x). The organic layer was washed with brine (1 x), dried over NaiSCL filtered, and concentrated under reduced pressure. The crude product was purified by flash chromatography (silica gel, gradient 40-70% acetone in hexanes) to afford the product as a solid. Final purification was performed by reverse-phase HPLC (gradient: 10-37% MeCN with 0.1% TFA in FEO with 0.1% TFA), followed by lyophilization of the collected fractions, yielding the desired product as an oil.Amide ABPP ProbesStep 1: Amide Coupling with Propargylamine
[0254] To a solution of the corresponding indanone carboxylic acid (5.00 mmol, 1.0 equiv) in anhydrous DMF (final concentration: 0.5 M) at room temperature, were added propargylamine (1.5 equiv, 7.50 mmol), HATU (1.5 equiv, 7.50 mmol), and DIPEA (3.0 equiv, 15.0 mmol). The reaction mixture was stirred overnight at room temperature. After completion (monitored by TLC), the mixture was diluted with ethyl acetate (50 mL) and washed successively with 1 M aqueous HC1 (2 x 30 mL), water (2 x 30 mL), saturated aqueous NaHCCL (2 x 30 mL), and brineAtty. Docket: UCSF-865WO(1 x 30 mL). The organic layer was dried over Na2SO4, filtered, and concentrated under reduced pressure. The crude amide was purified by flash chromatography (gradient 0-50% EtOAc in hexanes) to yield the pure propargyl amide intermediate.Step 2: Oxidation to 2,2-Dihydroxy-l,3-indanedione derivatives
[0255] To a solution of the propargyl amide intermediate (obtained above, 1.0 equiv) in 1,4-dioxane / H2O (10:1 v / v, -200 mM)was added selenium dioxide (2.1 equiv). The reaction mixture was heated to reflux (96 °C) and stirred overnight. After cooling, the reaction was diluted with ethyl acetate and washed sequentially with water (2 x 20 mL), 1 M aqueous HC1 (2 x 20 mL), saturated aqueous NaHCCL (2 x 20 mL), and brine (20 mL). The organic layer was dried over Na2SO4, filtered, and concentrated under reduced pressure. Purification by flash chromatography (gradient 30-50% acetone in hexanes), followed by reverse-phase HPLC (gradient 2-25% MeCN with 0.1% TFA in FLO with 0.1% TFA) and lyophilization yielded the final 2,2-dihydroxy-l,3-indanedione derivatives as a pink solid.Enantiopair ProbesStep 1: Formation of tert-butyl 3-(prop-2-yn-l-yloxy)piperidine-l-carboxylate derivatives
[0256] To a round-bottom flask containing tert-butyl 3-hydroxypiperidine-l -carboxylate (1.00 g.4.97 mmol, 1.0 equiv) and I COa (2.0 equiv, 10.0 mmol) was added anhydrous DMF (10 mL), followed by propargyl bromide (80% wt solution in toluene, 1.5 equiv, 7.45 mmol). The reaction mixture was heated to 60 °C for 16 hours. The reaction was then cooled to room temperature, diluted with CFLCL (30 mL), and washed sequentially with water (3 x 30 mL) and brine (2 x 30Atty. Docket: UCSF-865WOmL). The organic layer was dried over MgSCL, filtered, and concentrated under reduced pressure. Purification by flash chromatography (SiC>2, gradient: 100% hexanes to 10% EtOAc in hexanes) afforded tert-butyl 3-(prop-2-yn-l-yloxy)piperidine-l -carboxylate as a colorless oil or white solid. The Boc-protected intermediate (tert-butyl 3-(prop-2-yn-l-yloxy)piperidine-l-carboxylate) was dissolved in 20% TFA / DCM and stirred at room temperature for 1 hours. The reaction mixture was diluted in toluene and concentrated under reduced pressure. The residue was diluted with saturated aqueous NaHCCh (20 mL) to neutralize the residual acid to a basic pH (-8-9), and extracted with ethyl acetate (3 x 20 mL). The organic layers were combined, dried over Na2SC>4. filtered, and concentrated under reduced pressure to yield the corresponding free amine (3-(prop-2-yn-l-yloxy)piperidine derivatives) directly as a pure oil suitable for use in the next step without further purification.Step 2: Amide Coupling with Indanone Carboxylic Acids
[0257] To a solution of 4carboxy-indanone(1.0 equiv) in anhydrous DMF (0.5 M) were added the corresponding free amine intermediate (from step 1, 1.5 equiv), HATU (1.5 equiv), and DIPEA (3.0 equiv). The reaction mixture was stirred overnight at room temperature. After completion (monitored by TLC), the reaction was diluted with ethyl acetate (50 mL) and washed sequentially with 1 M aqueous HC1 (2 x 20 mL), water (2 x 20 mL), saturated aqueous NallCCh (2 x 20 mL), and brine (20 mL). The organic layer was dried over Na2SC>4, filtered, and concentrated under reduced pressure. Purification by flash chromatography (gradient: 0-50% EtOAc in hexanes) afforded the pure amide intermediate as an oil.Step 3: Oxidation to 2, 2-Dihydroxy- 1,3 -Indanedione Derivatives
[0258] To a solution of the amide intermediate from step 3 (1.0 equiv) in l,4-dioxane / H20 (10:1 v / v, -200 mM) was added SeO2 (2.1 equiv). The reaction mixture was heated to reflux (-96 °C) and stirred overnight. After cooling to room temperature, the reaction mixture was diluted with ethyl acetate and washed sequentially with water (2 x 20 mL), 1 M aqueous HO (2 x 20 mL), saturated aqueous NaHCOs (2 x 20 mL), and brine (20 mL). The organic layer was dried over Na2SO4. filtered, and concentrated. Purification by flash chromatography (30-50% acetone in hexanes) yielded 2,2-dihydroxy-l,3-indanedione derivatives a red oil.Atty. Docket: UCSF-865WOUnnatural amino acid - Phenylalanine mimicStep 1: Organozinc Coupling Reaction
[0259] To a rapidly stirred suspension of zinc dust (929.3 mg, 14.21 mmol, 3.0 equiv) in anhydrous DMF (23.7 mL) in a 25 mL one-necked round-bottom flask fitted with a three-way tap and a magnetic stirrer bar was added a catalytic amount (5-10 mg) of iodine. The mixture was stirred rapidly at room temperature until the iodine color faded, indicating activated zinc formation. To this activated zinc suspension was added methyl (R)-2-((tert- butoxycarbonyl)amino)-3-iodopropanoate (1.949 g, 5.922 mmol, 1.25 equiv) in one portion. A slight exothermic reaction was observed. The mixture was stirred vigorously for an additional 5 min until it returned to room temperature. Stirring was stopped and zinc dust was allowed to settle. The clear supernatant solution containing the zinc intermediate was carefully withdrawn using a syringe. The obtained activated zinc solution was then transferred via syringe to a separate side-arm round-bottom flask containing 6-bromo-2,3-dihydro-lH-inden-l-one (1.000 g, 4.738 mmol, 1.0 equiv) and trans-bis(acetato)bis[2-[bis(2- methylphenyl)phosphino]benzyl]dipalladium(II) (222.1 mg, 236.9 pmol, 0.05 equiv) under a nitrogen atmosphere. The reaction mixture was stirred at 19 °C for 24 hours under nitrogen. Upon completion, the concentrated crude mixture was dry loaded onto silica and purification by flash chromatography (SiCL, gradient: 100% hexanes to 10% EtOAc in hexanes)) to obtain pure methyl (S)-2-((tert-butoxycarbonyl)amino)-3-(2,3-dihydro-lH-inden-l-one)propanoate as an oil (83% yield).Step 2: Oxidation with Selenium Dioxide
[0260] To a solution of the intermediate from Step 1 (1.0 equiv) in a mixture of l,4-dioxane / H20 (10:1 v / v, -200 mM) was added SeCL (2.1 equiv). The reaction mixture was heated at reflux (96 °C) overnight. After cooling to room temperature, the reaction mixture was diluted with ethyl acetate and sequentially washed with water (2 x 20 mL), 1 M aqueous HC1 (2 x 20mL), saturated aqueous NaHCOs (2 x 20 mL), and brine (20 mL). The organic phase was driedAtty. Docket: UCSF-865WO(Na2SC>4), filtered, and concentrated. The crude material was purified by flash chromatography (30-50% acetone in hexanes), to yield methyl (S)-2-((tert-butoxycarbonyl)amino)-3-(2,2-dihydroxy-l,3-dioxo-indan-l-yl)propanoate as a brown solid.Step 3: Saponification (Ester Hydrolysis) and Boc-Deprotection
[0261] To the purified ester intermediate obtained from Step 2 dissolved in a 3:1 mixture of THF / water was added LiOH (3 equiv) at room temperature. The reaction was stirred at room temperature for 4 hours, after which completion was verified by TLC. The reaction mixture was acidified to approximately pH = 2 with 1 M aqueous HC1 and extracted with ethyl acetate (3 x 20 mL). The combined organic extracts were dried (JM^SCb), filtered, and concentrated under reduced pressure to yield the free carboxylic acid intermediate, which was used without further purification. The carboxylic acid intermediate was dissolved in 4 M HC1 in dioxane was stirred at 0 °C for 2 hours. After removal of solvents under reduced pressure, the residue was purified by evaporation of excess acid in toluene to provide the final product: (S)-2-amino-3-(2,2-dihydroxy-l,3-dioxo-indan-l-yl)propanoic acid hydrochloride as a pure white solid.Unnatural amino acid - Lysine AmideStep 1: Amide Coupling of Indanone Carboxylic Acid with tert-butyl (tert-butoxycarbonyl)-L-lysinate
[0173] To a solution of l-oxo-2,3-dihydro-lH-indene-5-carboxylic acid (500 mg, 2.84 mmol, 1.0 equiv) in anhydrous DMF (11.4 mL, 0.25 M) at room temperature were added sequentially tert-butyl (tert-butoxycarbonyl)-L-lysinate (944 mg, 3.12 mmol, 1.1 equiv), HATH (1.30 g. 3.41 mmol, 1.2 equiv), and DIPEA (1.10 g, 1.48 mL, 8.51 mmol, 3.0 equiv). The reaction mixture was stirred overnight at room temperature. After reaction completion (monitored by TLC), the mixture was diluted with ethyl acetate (50 mL) and washed successively with 1 M aqueous HC1 (2 x 20 mL), water (2 x 20 mL), saturated aqueous NaHCCL (2 x 20 mL), and brine (20 mL). The organic layer was dried over NaiSCh, filtered, andAtty. Docket: UCSF-865WOconcentrated under reduced pressure. Purification by flash chromatography (gradient: 0-50% EtOAc in hexanes) afforded tert-butyl N-2-(tert-butoxycarbonyl)-N-6-(l-oxo-2,3-dihydro-lH-indene-5-carbonyl)-L-lysinate as a yellow oil (1.21 g, 2.63 mmol, 92.6% yield).Step 2: Oxidation to the 2,2-Dihydroxy-l,3-Indanedione Derivative
[0173] To a solution of the above amide intermediate (1.0 equiv) in l,4-dioxane / H2O (10:1 v / v, -200 mM) was added selenium dioxide (2.1 equiv). The reaction mixture was heated at reflux (96 °C) overnight. After cooling to room temperature, the mixture was diluted with ethyl acetate and washed sequentially with water (2 x 20 mL), 1 M aqueous HC1 (2 x 20 mL), saturated aqueous NaHCCb (2 x 20 mL), and brine (20 mL). The organic layer was dried over Na2SO4, filtered, and concentrated. Purification by flash chromatography (gradient 30-50% acetone in hexanes), yielded tert-butyl N-2-(tert-butoxycarbonyl)-N-6-(2,2-dihydroxy-l,3-dioxo-indan-5-carbonyl)-L-lysinate as a red oil.Step 3: Global Deprotection under Acidic Conditions
[0173] The intermediate obtained from step 2 was dissolved in 4 M HC1 in dioxane at 0 °C and stirred for 2 hours. The reaction mixture was then concentrated under reduced pressure. The residue was co-evaporated with toluene (3 x 10 mL) to remove residual HC1 and dioxane completely, affording the final deprotected product, N6-(2,2-dihydroxy-l,3-dioxo-indan-5-carbonyl)-L-lysine hydrochloride, as an orange-red oil.Synthesis of PHGDH covalent inhibitorStep 1: Formation of the Sulfonate Ester
[0173] To a stirred suspension of 4-(chlorosulfonyl)benzoic acid (882 mg, 4.00 mmol, 1.0 equiv) and 4-hydroxy-2,3-dihydro-lH-inden-l-one (622 mg, 4.20 mmol, 1.05 equiv) in anhydrous DCM (20.0 mL) at 0 °C was added triethylamine (TEA, 1.01 g, 1.39 mL, 10.0 mmol, 2.5 equiv) dropwise. The reaction was stirred for an additional 2 hours at room temperature, afterAtty. Docket: UCSF-865WOwhich the reaction was quenched by adding water. The reaction mixture was diluted with additional DCM (30 mL) and acidified carefully with 1 M aqueous HC1. The aqueous layer was extracted with DCM (3 x 30 mL). The combined organic layers were washed with brine, dried over Na2SO4, filtered, and concentrated under reduced pressure to give the crude product 4-(((l-oxo-2, 3-dihydro-lH-inden-4-yl)oxy)sulfonyl)benzoic acid as a solid (1.29 g, 3.88 mmol, 97% yield), used directly in the subsequent amidation reaction without further purification.Step 2: Amide Coupling with 2-Amino-4-(2,4-difluorophenyl)thiazole
[0173] To a solution of 4-(((l-oxo-2,3-dihydro-lH-inden-4-yl)oxy)sulfonyl)benzoic acid (1.0 equiv) from step 1 in anhydrous DMF (0.25 M) were added 2-amino-4-(2,4-difluorophenyl)thiazole (1.2 equiv), HATU (1.2 equiv), and DIPEA (3.0 equiv) at room temperature. The reaction was stirred overnight at room temperature. Upon completion (monitored by TLC), the reaction mixture was diluted with ethyl acetate (50 mL) and washed successively with 1 M aqueous HC1 (2 x 20 mL), water (2 x 20 mL), saturated aqueous NaHCCL (2 x 20 mL), and brine (20 mL). The organic phase was dried (Na2SO4), filtered, and concentrated under reduced pressure. The residue was purified by flash chromatography (gradient 20-60% EtOAc in hexanes) to yield the amide intermediate as a brown solid.Step 3: Oxidation with Selenium Dioxide
[0173] To a solution of the amide intermediate from step 2 (1.0 equiv) in l,4-dioxane / H2O (10:1 v / v, -200 mM) was added selenium dioxide (2.1 equiv). The reaction mixture was stirred at reflux (96 °C) overnight. After cooling, the reaction mixture was diluted with ethyl acetate, washed successively with water (2 x 20 mL), 1 M aqueous HC1 (2 x 20 mL), saturated aqueous NaHCCh (2 x 20 mL), and brine (20 mL). The organic layer was dried (Na2SO4), filtered, and concentrated under reduced pressure. The crude material was purified by flash chromatography (gradient 30-50% acetone in hexanes) without further HPLC purification to afford the final oxidized product 2,2-dihydroxy-l ,3-indanedione sulfonate-amide derivative as a red oil.NOTE- All of the following ninhydrin analogs have very poor water solubility even when resolubilized in DMSO and then added to water up to 10% DMSO in water solution. We hypothesized these ligands would be the simplest starting points but found in each case after diluting into water they would crash out in at concentrations that are not biologically compatible.Atty. Docket: UCSF-865WOStep 1: Azide Substitution of 6 -Bromoindanone
[0173] To a stirred solution of 6-bromo-2,3-dihydro-lH-inden-l-one (1.06 g, 5.00 mmol, 1.0 equiv) in a mixture of ethanol (11.7 mL) and water (5.0 mL) was added sodium azide (650 mg, 10.0 mmol, 2.0 equiv), sodium ascorbate (49.5 mg, 250 pmol, 0.05 equiv), Cui (95.2 mg, 500 pmol. 0.1 equiv), and N,N'-dimethylethylenediamine (DMEDA, 88.2 mg, 107 pL, 1.00 mmol, 0.2 equiv). The reaction mixture was stirred at room temperature overnight. Upon completion (monitored by TLC), the reaction was diluted with ethyl acetate (50 mL) and washed sequentially with water (3 x 20 mL), brine (20 mL), and dried over Na2SC>4. The organic phase was filtered and concentrated under reduced pressure. The crude product was purified by flash chromatography (0-20% EtOAc in hexanes) to yield 6-azido-2,3-dihydro-lH-inden-l-one as a solid (327 mg, 1.89 mmol, 37.8% yield).Step 2: Selenium Dioxide Oxidation
[0173] To a solution of the azide intermediate (1.0 equiv) obtained above in a mixture of 1,4- dioxane / IUO (10:1 v / v, -200 mM)was added selenium dioxide (2.1 equiv). The reaction mixture was stirred at reflux (96 °C) overnight. After cooling to room temperature, the mixture was diluted with ethyl acetate and washed sequentially with water (2 x 20 mL), 1 M aqueous HC1 (2 x 20 mL), saturated aqueous NaHCCh (2 x 20 mL), and brine (20 mL). The organic phase was dried over Na2SC>4, filtered, and concentrated under reduced pressure. The crude product was purified by flash chromatography (30-50% acetone in hexanes) to yield the final product, 6- azido-2,2-dihydroxy-lH-indene-l,3(2H)-dione as a red oil.5 Alkynyl Ninhydrin- -Step 1: Sonogashira Coupling Reaction
[0173] In a flame-dried round-bottom flask under nitrogen atmosphere, to a mixture of 5-bromo- 2,2-dihydroxy-lH-indene-l,3(2H)-dione (25 mg, 97 pmol, 1.0equiv), Bis(triphenylphosphine)palladium(II) chloride (0.68 mg, 0.97 pmol, 0.01 equiv), and CuiAtty. Docket: UCSF-865WO(0.37 mg, 1.9 pmol, 0.02 equiv) in triethylamine (120 pL) and DMF (41 pL), wasadded ethynyltrimethylsilane (16 mg, 23 pL, 0.17 mmol, 1.7 equiv) dropwise. The reaction mixture was heated at 80 °C for 2 hours. After cooling to room temperature, the mixture was transferred to a separatory funnel and diluted with CH2CI2 (10 mL). The organic phase was washed sequentially with 10% HC1 (10 mL), 10% aqueous Na2COs (10 mL), and water (10 mL). The organic layer was dried over N 3286)4. filtered, and concentrated under reduced pressure. The crude residue was purified by flash chromatography (gradient 30-60% acetone in hexanes) to yield 2,2-dihydroxy-5-((trimethylsilyl)ethynyl)-lH-indene-l,3(2H)-dione as a solid (16 mg, 58 pmol, 60% yield).Step 2: Deprotection of the Trimethylsilyl Group
[0173] To a stirred solution of 2,2-dihydroxy-5-((trimethylsilyl)ethynyl)-lH-indene-l,3(2H)-dione (145.0 mg, 528.5 pmol, 1.0 equiv) in dry DCM (2.64 mL) at room temperature, was added tetrabutylammonium fluoride (345.5 mg, 379 pL, 1.321 mmol, 2.5 equiv) dropwise. The reaction mixture instantly turned dark black. Reaction completion was confirmed by LCMS within 5 minutes. The stir bar was removed, and the mixture was concentrated under reduced pressure to yield the final product, 5-ethynyl-2,2-dihydroxy-lH-indene-l,3(2H)-dione, suitable for use directly without further purification.&Step 1: Preparation of 3-oxo-2,3-dihydro-lH-indene-5-carbonitrile
[0173] A dry 100 mL round-bottom flask containing a magnetic stir bar was chargedwith copper(I) cyanide (CuCN, 5.1 g, 57 mmol, 1.2 equiv), 6-bromo-2,3-dihydro-lH-inden-l-one (10.0 g, 47 mmol, 1.0 equiv), and anhydrous DMF (40 mL). The flask was fitted with a condenser, placed under nitrogen, and heated to 140 °C for 16 hours. The reaction mixture was then cooled to room temperature and diluted with dichloromethane (500 mL). The resulting solids were removed by vacuum filtration, and the filtrate was washed sequentiallywith saturated aqueous NFLOAc (2 x 150 mL) and brine (150 mL). The organic phase was dried over MgSC>4, filtered, concentrated under reduced pressure, and dry-loaded onto silica gel. The product was purified via column chromatography (30% ethyl acetate / hexanes — 40% ethylAtty. Docket: UCSF-865WOacetate / hexanes) to yield 3-oxo-2,3-dihydro-lH-indene-5-carbonitrile as a tan solid (6.3 g, 40 mmol, 85% yield).Step 2: Oxidation to 2,2-Dihydroxy-5-cyano-lH-indene-l,3(2H)-dione
[0173] To a solution of 3-oxo-2,3-dihydro-lH-indene-5-carbonitrile (1.0 equiv) obtained in Step 1 dissolved in l,4-dioxane / H20 (10:1 v / v, -200 mM), was added selenium dioxide (2.1 equiv). The reaction mixture was heated at reflux (96 °C) overnight. After cooling to room temperature, the mixture was diluted with ethyl acetate and washed sequentially with water (2 x 20 mL), 1 M aqueous HC1 (2 x 20 mL), saturated aqueous NaHCCh (2 x 20 mL), and brine (20 mL). The organic phase was dried over Na2SO4, filtered, and concentrated under reduced pressure. The crude residue was purified by flash chromatography (gradient 30-50% acetone in hexanes) to afford the final product, 2,2-dihydroxy-5-cyano-lH-indene-l,3(2H)-dione, as a dark brown solid.Rhodamine conjugated ninhydrin triazole
[0262] To a solution of 3',6'-bis(diethylamino)-2-(prop-2-yn-l-yl)spiro[isoindoline-l,9'-xanthen] -3-one (48.0 mg, 0.100 mmol, 1.0 equiv) and 5-azido-2,2-dihydroxy-lH-indene-l,3(2H)-dione (21.9 mg, 0.100 mmol, 1.0 equiv) in CILCL (5 mL) and H2O (5 mL) were added sequentially CuSCb SfLO (4.99 mg, 20 pmol, 0.20 equiv) and sodium ascorbate (19.8 mg, 0.100 mmol, 1.0 equiv). The reaction mixture was vigorously stirred at room temperature for 12-20 hours. After reaction completion (monitored by TLC), additional CfLCL (10 mL) was added. The organic layer was separated, and the aqueous layer was extracted with CfLCL (3 x 10 mL). The combined organic extracts were washed with brine (2 x 20 mL), dried over MgSCL. filtered, and concentrated under reduced pressure. The residue was purified by flash chromatography (gradient 20-80% acetone in hexanes over 30 column volumes) to yield the final product, 5-(4-((3',6'-bis(diethylamino)-3-oxospiro[isoindoline-l,9'-xanthen]-2-yl)methyl)-lH-l,2,3-triazol-l-yl)-2,2-dihydroxy-lH-indene-1,3(2H)-dione, as a bright pink oil (32 mg, 46 pmol, 46% yield).Atty. Docket: UCSF-865WOEXAMPLE 3In-vitro characterization of arginine engagement of ninhydrin.
[0263] Vicinal dicarbonyl compounds in aqueous environments predominantly exist in their hydrate state, rendering them relatively inert to nucleophilic residues. However, the equilibrium between this hydrate state and the reactive dehydrate state governs their ability to engage in covalent modification. To investigate this equilibrium for ninhydrin versus the known argininereactive but promiscuous warhead phenyl glyoxal (PGO), we employed density functional theory (DFT) calculations (FIG 1A). These values predict the relative occupancy of each state, allowing us to assess their potential for selective arginine modification. These calculations revealed that both probes predominantly exist in their hydrate state; however, ninhydrin has a calculated Keqof 3.55 x IO3, indicating a higher fraction of the molecule resides in the reactive dehydrate state. Conversely, PGO has a significantly lower calculated Keqof 1.33 x IO , suggesting a much lower occupancy in the reactive form (FIG 1A). This favorable equilibrium positioning suggested that ninhydrin would have a strong potential for arginine modification.
[0264] To validate these predictions, we characterized ninhydrin reactivity toward arginine using an HPLC-based kinetics assay monitoring the consumption of Fmoc-Arginine (FIG IB). This approach provided direct evidence that the reaction proceeds at the terminal guanidine, while also enabling precise quantification via UV absorption. The reaction kinetics were evaluated across a pH gradient ranging from 6.0 to 10.0, with time points recorded at 2, 7, 15, 30, and 60 minutes. At physiologic pH (7.5), complete consumption of free arginine was observed within 60 minutes, whereas in more basic conditions reaction completion was achieved within 30 minutes (FIG IB). These fast 2nd order reaction kinetics are comparable to that of iodoacetamide to cysteine [Nelson et al. Anal. Biochem. 375, 187-195 (2008)].
[0265] The 2nd order reaction kinetics were evaluated across a pH gradient ranging from 6.0 to 10.0, with time points recorded at 2, 7, 15. 30. and 60 minutes. At pH 7.5, complete consumption of free arginine was observed within 60 minutes (kArg,PH7.5 = 828.0 ± 86.8 M'1min1), whereas in more basic conditions reaction completion was achieved within 30 minutes (kArg,PH9.5 = 1761.6 ± 1.2 M1min1) (FIG IB). The relatively fast 2nd order reaction kinetics are comparable to that of iodoacetamide with cysteine (kcys,PH9.o = 210.1 ± 30.2 M1min1) and further supported our hypothesis that ninhydrin may be a warhead amenable to engaging arginines in biologicalAtty. Docket: UCSF-865WOsystems on a reasonable timescale, Nelson et al. Anal. Biochem. 2008, 375 (2), 187-195. When we compared these rates to the reaction of PGO towards arginine we found that on average a 6-fold decrease in reactivity(kArg.PH 7.5 = 170.2 ± 0.2 M1min-1) with PGO never fully consuming the Fmoc-Arg at an hour even at more basic pHs (FIG IE, 1G, and 2N). A similar study was performed with ninhydrin towards Fmoc-lysine and glutathione, Bohme et al. Chem. Res.Toxicol. 2009, 22 (4), 742-750, which did not result in appreciable engagement. Conversely, PGO showed engagement of glutathione, consistent with prior reports of PGOs reactivity towards other nucleophilic sidechains, Zanon et al. Nat. Chem. 2025, 1-10.
[0266] To probe the selectivity of ninhydrin labeling and its potential to be pan-reactive towards arginines, we examined its modification stoichiometry on a model protein, bovine serum albumin (BSA), which contains 23 arginine residues (FIG 1C). Following 90 minutes of ninhydrin treatment, intact protein mass spectrometry revealed selective engagement of up to 9 residues under physiologic conditions (FIG ID), with increased reactivity in a pH-dependent manner, as we observed in our kinetics assay.
[0267] Following 30 minutes of ninhydrin treatment, intact protein mass spectrometry revealed up to 6 labeling events under physiologic conditions (FIG IF), with increased reactivity in a pH-dependent manner, as we observed in our free-arginine kinetics assay. The mass shift for each labeling event indicated addition of ninhydrin, consistent with engagement of ninhydrin with a nucleophilic sidechain forming dehydration products. Similar studies were also conducted for lysozyme, revealing a comparable pH-dependent reactivity profile modifying up to 11 residues on this protein with 11 arginines, Canfield et al. J. Biol. Chem. 1963, 238 (8), 2698-2707. While it was possible that engagement of other sidechains could result in a similar mass shift, it was unlikely that all labeling events could be attributed to non-arginine sidechains due to the reversibility of those reactions compared to arginine, especially considering that free lysine and cysteine are not detectably engaged by ninhydrin. Therefore, though possible that some of this labeling could be attributed to lysine or cysteine engagement, we postulated that it was unlikely that all of the labeling could be explained through this mechanism.Atty. Docket: UCSF-865WOEXAMPLE 4Ninhydrin probe development
[0268] Encouraged by robust arginine modification on a purified protein, we pursued development of a ninhydrin chemical probe and for testing in complex proteomes. We synthesized two regioisomeric ninhydrin probes and the reported arginine reactive probe PGO (FIG 2A). We visualized their relative reactivity across the proteome using gel-based chemical proteomic techniques (FIG 2B). Mino cell lysates were treated with the indicated probe for Ih. Following incubation, lysates were subjected to treatment with an azide-conjugated TAMRA fluorochrome under Cu(I)-catalyzed Azide- Alkyne Cycloaddition (CuAAC) conditions. The reaction was quenched with methanol and precipitated proteins were resuspended and separated by SDS-PAGE. In-gel fluorescence scanning revealed dose-dependent increases in labeling across the proteome with 5-Nin-Alk (Nin-Alk) exhibiting higher levels of labeling at all concentrations tested (FIG 2G). As a benchmark, we compared proteome- wide labeling by our ninhydrin-containing probes to the pan-cysteine reactive probe iodoacetamide-alkyne (lA-Alk). Ninhydrin probes at 100 pM exhibited proteomic coverage comparable to lA-Alk at 10 pM, whereas PGO displayed markedly reduced labeling efficiencies, despite a known lack of specificity for arginine residues. These results suggested that ninhydrin-based probes achieve improved proteomic labeling efficiencies compared to the best-in-class PGO probe. Both 5-Nin-Alk (Nin-Alk) and PGO exhibited reduced labeling compared to lA-Alk (FIG 2G). Curious as to if a regioisomer of 5-Nin-Alk (Nin-Alk) would have similar labeling, we also synthesized 4-Nin-Alk (FIG 2A), which exhibited similar labeling of cell lysates compared to 5-Nin-Alk (Nin-Alk).
[0269] We next asked if our ninhydrin-containing probes were capable of labeling live cells in a time-dependent manner. A549 cells were treated for 0, 30, or 60 min with the indicated probe (10 pM). Labeled cells were pelleted, lysed, and subjected to conjugation with azide-TAMRA as described previously. In-gel fluorescence of scanning of SDS-PAGE separated proteins revealed time-dependent incorporation of both ninhydrin probes compared to PGO (FIG 2C). Having demonstrated the ability of both ninhydrin probes to label proteins, we opted to focus our next studies on the 5-Nin-Alk, given its superior labeling in cells.
[0270] To assess 5-Nin-Alk’ s potential as a probe for pan-arginine chemical proteomics, we modified the established isotopically distinct azide-bearing desthiobiotin (isoDTB) workflowAtty. Docket: UCSF-865WOdeveloped by Hacker and co-workers to be amenable to characterizing reactive arginines, a workflow we call Arginine-specific isotopic Labeling of Functional Guanidines (RisoLFG [RAP], FIG 2D). Cell lysates were treated with 5-Nin-Alk or DMSO vehicle at a desired concentration. Alkyne-labeled arginines were further modified with an azide-bearing desthiobiotin tag that is isotopically heavy or light via copper(I) -catalyzed azide-alkyne cycloaddition (CuAAC). Isotopically distinct sample proteome pairs (in this case, probe-modified and no-probe) were subjected to single-pot, solid-phase sample preparation (SP3). Enriched and desalted peptides were analyzed by mass spectrometry. Modified arginine-containing peptides were identified and quantified by FragPipe22.
[0271] We performed an open search to directly assess the selectivity of our probe for arginines across the proteome. Open searches characterize the frequency of mass modifications to peptides, which allowing for the identification of large mass differences between unmodified peptide sequences and experimentally observed precursors [Yu, F. et al. Nat. Commun. 11, 4065 (2020)]. For our 5-Nin-Alk treated lysates, we observed just one major mass modification, 701.3014 da. which corresponds to the cyclic adduct to arginine of 5-Nin-Alk clicked to the heavy isoDTB tag (FIG 2E). In further support of arginine selectivity, our isoDTB samples were prepared in the absence of a reducing agent, which is often required for the trapping of lysine modifications (Yang, T. et al. Nat. Chem. Biol. 18, 934-941 (2022)). To evaluate potential free thiol engagement through orthogonal methods, we performed an in vitro glutathione (GSH) reactivity assay. While the PGO probe displayed some GSH engagement, our probe did not appear to appreciably react (FIG 2F).
[0272] We next investigated whether 5-Nin-Alk (Nin-Alk) was capable of labeling live cells in a time-dependent manner. Mino cells were treated for 0, 30, or 60 min with the indicated probe (100 pM). Labeled cells were pelleted, lysed, and subjected to conjugation with TAMRA-azide as described above. In-gel fluorescence scanning of SDS-PAGE separated proteins revealed improved time-dependent labeling with 5-Nin-Alk (Nin-Alk) compared to PGO (FIG 2H). Notably, 4-Nin-Alk also exhibited markedly reduced labeling in live cells compared to 5-Nin-Alk (Nin-Alk), affirming our focus on 5-Nin-Alk (Nin-Alk) as a potential probe for characterizing arginine reactivity.Atty. Docket: UCSF-865WO
[0273] To evaluate 5-Nin-Alk (Nin-Alk)’ s potential as a pan-arginine chemical probe, we developed a workflow utilizing the isotopically distinct azide-bearing desthiobiotin (isoDTB) tags developed by Hacker and co-workers [Zanon et al. Angew. Chem. Int. Ed Engl. 2020, 59 (7), 2829-2836]. (FIG 21). Cell lysates were treated with 5-Nin-Alk (Nin-Alk) at either 100 pM or 1 mM for Ih, and the resulting alkyne-labeled arginines were conjugated to light or heavy IsoDTB tags via copper(I)-catalyzed azide-alkyne cycloaddition (CuAAC). Labeled proteome pairs were combined and processed by single-pot, solid-phase sample preparation (SP3), yielding enriched, desalted peptides for subsequent mass spectrometry analysishttps: / / aperpile.com / c / vOklQw / Ty3Xd [Hughes et al. Nat. Protoc. 2019, 14 (1), 68-85.] We call this pipeline Reactive Arginine Profiling (RAP).
[0274] We first performed an open search to validate detection of intact probe modification. Open searches characterize the frequency of mass modifications to peptides, which allows for the identification of large mass differences between unmodified peptide sequences and experimentally observed precursors [Yu et al. Nat. Commun. 2020, 11 (1), 4065.]. For our 5-Nin-Alk (Nin-Alk)-treated lysates, we observed one major mass modification, +701.308600 da, which corresponds to the cyclic mono-dehydration adduct of arginine and 5-Nin-Alk (Nin-Alk) clicked to the heavy isoDTB tag. A less abundant mass corresponding to the cyclic monodehydration adduct of arginine and 5-Nin-Alk (Nin-Alk) clicked to the light isoDTB tag was also detected (+695.310118 da, FIG 2J), which we expected given the 10-fold lower concentration of 5-Nin-Alk (Nin-Alk) used in the light channel. pLOGO analysis of peptides with a modified arginine found no significant 2D sequence determinants of labeling with some minor enrichment for acidic residues 3-4 residues away (FIG 2K) [O’Shea et al. Nat. Methods 2013, 10 (12), 1211-1212.]. Finally, in further support of arginine selectivity, our RAP samples were prepared in the absence of a reducing agent, which is often required for the trapping of lysine modifications.
[0275] We next compared the performance of 5-Nin-Alk (Nin-Alk) to the PGO probe in a comparative RAP experiment in which paired lysates were each treated with either probe (ImM; Ih) and conjugated to either a heavy or light isoDTB tag. UpSet plot analysis revealed that the majority of PGO-engaged peptides were also engaged by 5-Nin-Alk (Nin-Alk) (n = 2862 of 2871 arginines are shared) while remarkably few arginines are uniquely engaged with PGO (n = 9) (FIG 2L). This high overlap suggests that most PGO-reactive arginines are also accessible by 5-Atty. Docket: UCSF-865WONin-Alk (Nin-Alk). Conversely, an additional 713 arginines are uniquely engaged by 5-Nin-Alk (Nin-Alk), suggestive that the probe engages a broader subset of arginines in the proteome. While less intrinsically reactive, PGO remains an important electrophile within the arginine-targeting toolkit, offering a more tempered reactivity profile that can facilitate more discriminating labeling and may prove advantageous in experimental contexts demanding precise or attenuated engagement of reactive residues.
[0276] Most arginines engaged by both probes were preferentially engaged by 5-Nin-Alk (Nin-Alk) (2704 arginines), whereas a significantly smaller fraction of sites were favored or exclusively engaged by PGO (n = 44) (FIG 2M). This pattern suggests that PGO frequently does not reach equivalent occupancy of shared arginines under identical conditions. Solution-phase Fmoc-Arginine reactivity mirrored these proteomic trends, with ninhydrin showing faster and more extensive arginine consumption than PGO across all tested pH values and time points (FIG 1G and 2N). Together, these findings suggest that 5-Nin-Alk (Nin-Alk) achieves a broader and more uniform engagement of reactive arginines and is therefore well suited for proteome-wide screening applications analogous to iodoacetamide-alkyne in cysteine profiling. It is also worth noting that the PGO adduct has been reported to oxidize during sample preparation and data acquisition [Zanon et al. Nat. Chem. 2025. 1-10.]. We also observed this limitation, which further complicates quantitative analysis. For our comparative studies we were careful to perform searches with both the non-oxidized modification and oxidized modification (FIG 2M and 2N). These data revealed comparable reactivity profiles and combining both searches into our quantitative pipeline did not alter our results (FIG 1H, II, 1J, IK, 2M). In contrast, 5-Nin-Alk (Nin-Alk) yields a single dominant mass adduct and clean isotopic labeling profiles, affirming its suitability as a general arginine-targeting chemical probe.EXAMPLE 5Chemical proteomic profiling of arginine labeling
[0277] Inspired by established workflows to characterize cysteine and lysine reactivity, we employed RisoLFG [RAP] to identify reactive arginines embedded in proteins and characterize their relative reactivities (FIG 3A). Cell lysates were treated with our alkyne-bearing panspecific arginine probe, 5-Nin-Alk, at a high concentration, which saturates reactive arginine labeling (1 mM, Ih) or a low concentration in which a subset of reactive arginines are labeled,Atty. Docket: UCSF-865WOthe most reactive of which are saturated (100 pM, Ih). Labeled lysates are further tagged with a ‘heavy’ or ‘light’ azide-bearing desthiobiotin reagent under CuAAC conditions. Lysates labeled under each condition, now also isotopically distinct, can be combined, processed by SP3, and purified peptides subjected to quantitative mass spectrometry analysis using a modified isoDTB closed-search workflow in FragPipe v22.0 as developed by Hacker and co-workers [Zanon et al. Angew. Chem. Int. Ed. 59, 2829-2836 (2020)].
[0278] We defined hyperreactive arginines as those which display similar enrichment at both concentrations (comparative ratio of intensity between conditions of less than 2), whereas less reactive arginines display concentration-dependent labeling (comparative ratio of at least 2) (FIG 3A). We performed these studies in Mino cells, revealing 10,642 modified peptides across 2,094 distinct proteins (FIG 3B). The median ratio of modification between 1 mM and 100 pM 5-Nin-Alk across all arginines was 4.38. Among these reactive arginines, 750 displayed hyperreactivity, while the majority (9,982) demonstrated concentration-dependent engagement.
[0279] To contextualize the determinants of 5-Nin-Alk modification, solvent-accessible surface area (SASA) analysis revealed that ninhydrin-engaging arginines exhibited an average relative solvent accessibility of approximately 0.21, indicating that they are most often only partially solvent-exposed, rather than fully surface-accessible (FIG 3C). This trend was consistent for hyper-reactive arginines. Unlike what is currently expected for cysteine and lysine reactivity, we found that arginine reactivity is independent of pKa (FIG 3D). Taken together, these data suggest that arginine reactivity is driven by local environment variables exclusive of protonation state or solvent accessibility.
[0280] Having validated the arginine-targeted pan-reactivity of 5-Nin-Alk (Nin-Alk), we next leveraged the data to identify reactive arginines embedded in proteins. Raw data collected from the RAP workflow comparing 1 mM and 100 pM 5-Nin-Alk (Nin-Alk) labeling events (which has been standard for arginine-targeted datasets, given their increased frequency in the proteome compared to cysteine) were subjected to a modified isoDTB closed-search workflow (FIG 2I)27.From Mino cell lysates, we captured 10,642 modified peptide ions mapping to 6,888 unique arginines across 2,076 distinct proteins (FIG 3E). Despite a recent observation from the Hacker group that modified arginines are very rarely (0.2 %) recognized by trypsin as a proteolytic cleavage site, we found that a greater subset (6.74 %) of our modifications were still beingAtty. Docket: UCSF-865WOrecognized by trypsin as a cleavage site. We performed similar RAP studies using 5-Nin-Alk (Nin-Alk) in lysates and live cells at a lower concentration range, 100 pM and 10 pM, as we felt these concentrations were more appropriate for live cell studies (FIGs 3S, 3T, 3U, 3V).Strikingly, our lysate and in cell studies at these lower concentrations resulted in less than 100 quantified arginines for each condition. While live cell data contained intracellular proteins of interest such as the Akt modulator TCL1A, mitochondrial GTPase EFTU, and the endoplasmic reticulum disulfide isomerase PDIA3, suggestive of cell uptake of the 5-Nin-Alk (Nin-Alk) probe, the low coverage nature of our lower concentration pairing deprioritized these experimental conditions (FIG 3S). Therefore, we moved forward analyzing our cell lysate dataset comparing 1 mM and 100 pM 5-Nin-Alk (Nin-Alk) engagement of nearly 7,000 arginines.
[0281] Statistical overrepresentation analysis of proteins containing reactive arginines revealed enrichment for protein classes associated with protein translation, mRNA processing, and nucleic acid binding, among others (FIG 3F) [Thomas et al. Protein Sci. 2022, 31 (1), 8-22.; Mi et al. Nat. Protoc. 2019. 14 (3). 703-721.]. We also cross-referenced proteins harboring reactive arginines with Pharos, a database developed by the NIH to categorize a protein’s status as a drug target based upon available chemical and biological evidence. This analysis revealed that the majority of proteins harboring reactive arginines are potential therapeutic targets supported by strong biological evidence but currently lack chemical tools for modulation (FIG 3G).Complementary annotations from DrugBank further suggest an opportunity to expand covalent small-molecule engagement to a broader set of potentially druggable proteins through reactive arginine engagement (FIG 3H).
[0282] Beyond these structural contexts, we turned to cation-7t contacts with aromatic side chains, quantifying interactions across all available PDB models for each protein (FIG 3K). We applied two distance thresholds, 8 A, which is the commonly used cutoff, and a more stringent 5.5 A cutoff for true energetic relevance, measured from the arginine CZ atom to the aromatic ring-centroid. To ensure only geometrically plausible face-on and edge-on interactions were included, we added angular filters: for parallel (face-on) interactions, the CZ-to-centroid vector was required to lie within 30° of the ring normal, while for T-shaped (edge-on) interactions, deviation from 90° had to be < 20°. Using these criteria, we identified 905 total cation -it contacts at 8.0 A, while only 332 satisfied the stricter 5.5 A threshold (FIG 3K). Phenylalanine andAtty. Docket: UCSF-865WOtyrosine accounted for the majority of these contacts, compared to histidine and tryptophan. At the more stringent distance threshold, the distribution of aromatic partners mirrors the natural abundance of aromatic residues in the human proteome. Unlike recent efforts to selectively label tryptophans that showed enrichment for tryptophans engaged in cation-7t interactions [Xie et al. Nature 2024, 627 (8004), 680-687.], we do not see an enrichment of arginines engaged in cation-n interactions. This indicates that arginine reactivity is not biased by a specific type of tt-donor or the stabilizing effect of a cation-7r interaction.
[0283] We next measured salt bridges as the minimum distance between any arginine guanidinium nitrogen and Asp / Glu carboxylate oxygens for both PDB entries and AlphaFold models. In AlphaFold structures, these distances cluster tightly around ~2.6 to 2.8 A, whereas PDB entries span a broader 2.5 to 4.0 A range with a median closer to 3.2 A (FIG 3L). Likewise, relative solvent accessibility values derived from solvent accessible surface area (SASA) [Shrake et al. J. Mol. Biol. 1973, 79 (2), 351-371.] analysis on PDB models are systematically higher and more variable than those from AlphaFold (median RS A for AlphaFold predicted models = 0.19, median RSA for PDB experimental models = 0.33: FIG 3M). Although AlphaFold models offer near proteome-wide coverage, their monomeric, in-silico nature often yields overly “idealized” geometries that may not reflect the full range of biologically relevant conformationsHe, X. et al. Acta Pharmacol. Sin. 2025, 46 (4), 1111-1122.: Shen, S.-Y. et al. Acta Pharmacol. Sin. 2025 DOI: 10.1038 / s41401 -025-01617-4.; He, X.-H. et al. Acta Pharmacol. Sin. 2023, 44 (1), 1-7. This was exemplified in our secondary structure analysis on AlphaFold models, which showed a greater enrichment of alpha helices than determined in the PDB models (FIG 31). which is a common feature of AlphaFold structures [Stevens, A. O. et al. Biomolecules 2022, 12 (7), 985.]. To address this, we prioritized experimentally determined PDB structures, trading breadth for depth by focusing on a smaller set of proteins with multiple crystallographic or complexed snapshots that better approximate physiological contexts. In both cases, the experimental data capture greater conformational diversity and exposure compared to the predicted models, which we believe are critical parameters for accurately mapping covalent-probe hotspots.
[0284] Transitioning from these broader structural contexts, we sought to dissect the biophysical and functional determinants that govern differential arginine reactivity using datasets of reactiveAtty. Docket: UCSF-865WOand hyper-reactive arginines (FIG 3M). We defined hyper-reactive arginines as those which display similar enrichment at high and low concentrations (comparative ratio of intensity between conditions of less than 2), whereas less reactive arginines display concentrationdependent labeling (comparative ratio of at least 2) (FIG 3M and FIG 21). The median log2 ratio of enrichment between 1 mM and 100 iiM 5-Nin-Alk (Nin-Alk) across all arginines was 2.50, reflecting dose-dependent labeling of most reactive arginines. Notably. 410 sites (~ 6.0% of detected arginines) were hyper-reactive, fully engaged at 100 pM, while the remaining 6478 sites exhibited clear, concentration-dependent labeling. This percentage of arginines annotated as hyper-reactive is consistent with reports for cysteine and arginine [Weerapana, E. et al. Nature 122010, 468 (7325), 790-795.; Backus, K. M. et al. Nature 062016, 534 (7608), 570-574.; Hacker, S. M. et al. Nat. Chem. 2017, 9 (12), 1181-1190.]. Moreover, proteins typically harbor a single hyper-reactive arginine but can contain multiple concentration-dependent reactive arginines (FIG 3N).
[0285] With our lists of either all reactive (any arginine engaging ninhydrin) or hyper-reactive (enrichment ratio less than 2) arginines, we set out to identify distinguishing features between the two sets. Even though arginine remains protonated under physiological conditions, cysteine and lysine reactivity have been linked to residue pKa values [Liu, R. et al. J. Am. Chem. Soc. 2019, 141 (16), 6553-6560.; Awoonor- Williams, E. et al. J. Chem.. Inf. Model. 2023, 63 (7), 2170-2180.], so we first asked whether arginine labeling follows a similar trend. Predicted pKa values for all reactive arginines within the PDB were, as expected, tightly clustered around 12.5.Distributions for hyper-reactive versus all reactive arginine sites showed no significant differences (FIG 30), indicating that, unlike lysine, intrinsic protonation state does not drive labeling efficiency. This finding was consistent for predicted AlphaFold structures. Given the established importance of solvent accessibility for cysteine modification by iodoacetamide [Whitehurst, C. B. et al. J. Virol. 2007, 81 (12), 6231-6240.], we next assessed the relative solvent exposure of reactive arginines. SASA analysis revealed comparable relative solvent exposure distributions for hyper-reactive and all reactive sites, with both populations exhibiting partial rather than full surface exposure (median RS A for hyper-reactive = 0.31, median RS A for all reactive = 0.33, mean RSA for both = 0.33; FIG 3P. The pKa and solvent accessibility independence of arginine labeling implies that: arginine-directed probes can be designed to engage arginines across diverse structural and environmental contexts, and selectivity would beAtty. Docket: UCSF-865WObased on interactions between the probe’s scaffold and the binding site, not on the intrinsic reactivity of the engaged arginine. This highlights a notable difference compared to cysteinereactive probe design, wherein pKa and solvent accessibility of targeted cysteines play major roles in the reactivity and selectivity of covalent probes.
[0286] To examine whether salt-bridge geometry contributes to hyper-reactivity, we compared the salt bridge distances for hyper-reactive and all reactive arginines in PDB models (FIG 3Q).Both populations exhibit broad, overlapping distributions (median distances = 3.2 A for both hyper-reactive and all reactive sites), indicating that salt-bridge proximity alone does not distinguish the hyper-reactive subset. Lastly, we explored functional constraints by mapping mean pathogenicity scores (AlphaMissense) onto each reactive arginine (FIG 3R). A high pathogenicity score indicates that mutations of a given residue are associated with significant functional impairment or disease-related phenotypes. Both hyper-reactive and all reactive residues are enriched at the high end of the pathogenicity scale, with 51% of enriched arginines falling into the likely pathogenic range, indicating a labeling bias towards arginines that are mutationally constrained, likely reflecting their critical structural or functional roles in maintaining cellular integrity. Since these arginines play critical roles in cellular function, they are less likely to undergo genetic mutations that could evade covalent engagement. Thus, covalent engagement at these essential positions likely mimics genetic perturbation, directly impacting residue functionality and consequently cellular processes, making them attractive and robust targets for drug and tool compound development.EXAMPLE 6Rationale for ninhydrin engagement at R479 on aconitase
[0287] In researching historical uses of ninhydrin in proteins we found a 1992 report (Lenzen et al. Naunyn. Schmiedeb ergs Arch. Pharmacol. 346, (1992)) that reported ninhydrin-mediated inhibition of mitochondrial aconitase catalytic activity at doses comparable to those used in our assays. Mitochondrial aconitase plays a central role in the tricarboxylic acid (TCA) cycle, linking cellular energy production to redox homeostasis and iron-sulfur cluster maintenance. Dysregulation of its activity has been implicated in metabolic disorders, neurodegeneration and tumor progression. While not widely targeted clinically, mitochondrial aconitase has garnered interest as a metabolic vulnerability in cancer and as a redox-sensitive node in neurodegenerativeAtty. Docket: UCSF-865WOdiseases, with small molecules and metabolic poisons, such as fluoroacetate, demonstrating proof-of-concept for pharmacological modulation. Our 5-Nin-Alk hyper-reactivity RisoLFG [RAP] dataset revealed a single modified reactive arginine on aconitase, R479. despite the presence of multiple arginine residues in the active site. Structural analysis of an aconitase crystal structure (PDB 1B0J) identified R479 as a key residue in the positioning of aconitase’s native substrate (FIG 4A). Covalent docking simulations to this crystal structure further demonstrated that ninhydrin modification at this site would sterically occlude the active site, likely preventing substrate binding (FIG 4B). These findings suggest that ninhydrin-mediated modification of R479 contributes to the observed inhibitory effect and provides mechanistic insight into how covalent modification of arginine influences aconitase function.
[0288] Ninhydrin has been reported to inhibit the catalytic activity of mitochondrial aconitase at doses comparable to those used in our assays [Lenzen, S. et al. Naunyn. Schmiedehergs. Arch. Pharmacol. 1992, 346 (5), 532-536.], but the mechanism of this inhibition has not been elucidated. Mitochondrial aconitase plays a central role in the tricarboxylic acid cycle, linking cellular energy production to redox homeostasis and iron-sulfur cluster maintenance [Castro, L. et al.. Acc. Chem. Res. 2019, 52 (9), 2609-2619.]. Dysregulation of its activity has been implicated in metabolic disorders, neurodegeneration, and tumor progression [Bak, D. W. et al. Biochim. Biophys. Acta Mol. Cell Res. 2024, 1871 (7), 119791.]. While not widely targeted clinically, mitochondrial aconitase has garnered interest as a metabolic vulnerability in cancer and as a redox-sensitive node in neurodegenerative diseases, with small molecules and metabolic poisons such as fluoroacetate demonstrating proof-of-concept for pharmacological modulation [Kim, E. et al. Nat. Commun. 2023, 14 (1), 3716.]. Our 5-Nin-Alk (Nin-Alk) hyper-reactivity RAP dataset revealed four modified arginines on mitochondrial aconitase (R479, R474, R564 and R607) which are all located within or proximal to the enzymes active site playing roles in the potioning of the native citrate / isocitrate. Among these, R479 was consistently detected across all samples and occupies a central position in substrate coordination. Structural analysis of an aconitase crystal structure (PDB 1B0J) [Lloyd, S. J. et al. Protein Sci. 1999, 8 (12), 2655-2662.] identified R479 as a key residue in the positioning of aconitase’s native substrate (FIG 4A). Covalent docking simulations to this crystal structure further demonstrated that ninhydrin modification at this site would sterically occlude the active site, likely preventing substrate binding (FIG 6B). The presence of multiple reactive arginines clustered around the catalyticAtty. Docket: UCSF-865WOpocket supports competitive engagement of the ninhydrin scaffold within the same binding region, providing mechanistic insight into how covalent modification of these residues may contribute to the observed inhibitory effect.EXAMPLE 7Development of a pipeline for ligand-directed arginine modification
[0289] To systematically identify which reactive arginines in our RisoLFG [RAP] dataset would be amenable to targeting with ninhy drin-endowed ligands, we developed an in-house computational structural proteomics pipeline (FIG 4C). The workflow compiles a list of UniProt IDs alongside their respective modified arginine residues and cross-references them with available high-resolution structures in the Protein Data Bank (PDB). For each structurally resolved arginine, the pipeline identifies nearby ligands within 3 A of the guanidinium group. To prioritize functionally relevant targets, we exclude primary human metabolites, crystallographic ions, and water molecules, focusing instead on ligand-binding arginines with the potential for covalent modification.
[0290] The analysis parsed through 9,348 unique PDB files, identifying 33,402 occurrences of arginines from our dataset. For each structurally resolved arginine, the pipeline identifies nearby ligands within 3 A of the guanidinium group. From these occurrences. 292 arginines were found to engage in 884 unique contacts with 461 distinct ligands. To prioritize functionally relevant ligands, we exclude primary human metabolites, crystallographic ions, and buffer components.
[0291] Applying this strategy, we identified R55 in cyclophilin A as a reactive arginine proximal to recently reported tri- vector ligands that provided synthetically tractable starting points (FIG 4D and 4E). Structure-guided ligand optimization led to the development of 4-CA-Alk (FIG 4D), where the aniline NFL was replaced with an isosteric trifluoromethyl group, the urea was substituted with an amide, and the ethyl ester was replaced with a ninhydrin warhead. This substitution leveraged the trifluoromethyl group’s ability to electronically and spatially mimic an aniline while avoiding potential liabilities of a free amine noted in the original tri-vector ligand report, thereby maintaining the binding geometry observed in the tri-vector ligands. This design positioned the ninhydrin warhead appropriately for covalent engagement with R55 in our docking models (FIG 4F). An in vitro washout-like activity assay revealed that CA-Nin-Alk (4-CA-Alk) attenuated CypA catalytic activity [Kofron, J. L. et al Biochemistry 1991, 30 (25),Atty. Docket: UCSF-865WO6127-6134.]. A regioisomer featuring an amide substitution at the 5 position of ninhydrin, designed to spatially disfavor R55 modification, 5-CA-Alk (iso-CA-Nin-Alk).
[0292] Using a modified coupled peptidyl-prolyl isomerase (PPIase) assay with the chromogenic substrate suc-AAPF-pNA, recombinant CypA (1 pM) was pre-incubated with CA-Nin-Alk (4-CA-Alk) (10 pM or 25 pM; 1 h) or vehicle and subsequently diluted 100-fold into the reaction mixture for a final CypA concentration of 100 nM immediately before initiating catalysis.Importantly, this dilution step would be expected to dissociate noncovalent complexes [Copeland, R. A. et al. Anal. Biochem. 2011, 416 (2), 206-210.] thereby mimicking a washout step, which would help indicate whether the interaction was covalent or non-covalent. Under these conditions, CA-Nin-Alk (4-CA-Alk) treatment resulted in a dose-dependent decrease in catalytic efficiency (FIG 41). Specifically, CA-Nin-Alk (4-CA-Alk) treatment at 10 pM and 25 pM reduced (kcat / Km) by 20.3 ± 8.1% and 56.4 ± 4.8%. respectively, relative to DMSO. These findings suggest that CA-Nin-Alk (4-CA-Alk) is either capable of covalently inhibiting CypA enzymatic activity or can function as a non-covalent inhibitor at 100 nM.
[0293] To quantitatively assess the proteome-wide selectivity of 4-CA-Alk, we performed probe- versus-no-probe RisoLFG [RAP] experiment (FIG 2A). At 100 pM, 4-CA-Alk enriched 132 modified peptides, whereas, at 10 pM, only 6 modified peptides were detected, demonstrating a marked reduction in off-target engagement at lower concentrations (FIG 4G).Notably, R55 of cyclophilin A was among these six peptides, indicating that 4-CA-Alk selectively modifies a limited subset of reactive arginines under physiologic conditions.Conversely, 5-CA-Alk (iso-CA-Nin-Alk) exhibited significantly reduced reactivity, confirming the necessity of precise warhead placement for selective engagement.
[0294] The ligand scaffold likely plays a key role in directing this selectivity by engaging R55 within its binding pocket, thereby restricting the modification of other reactive arginines throughout the proteome. Comparative analysis of 4-CA-Alk and its regioisomer 5-CA-Alk (iso-CA-Nin-Alk), which mispositions the ninhydrin warhead relative to R55, revealed a log2 selectivity ratio of 3.6, further supporting the importance of rational warhead placement in achieving site-specific labeling (FIG 4H).
[0295] To quantitatively assess the proteome-wide selectivity of CA-Nin-Alk (4-CA-Alk) and validate covalency for CypA at R55, we performed a probe-versus-no-probe RAP experiment atAtty. Docket: UCSF-865WOtwo concentrations. At 100 pM, CA-Nin-Alk (4-CA-Alk) enriched 287 modified peptides, indicating that CA-Nin-Alk (4-CA-Alk) selectively modifies a limited subset of reactive arginines under physiologic conditions. At 10 pM, only 7 modified peptides were detected, demonstrating a marked reduction in off-target engagement at lower concentrations (FIG 4J).Notably, R55 of cyclophilin A was among these 7 peptides. The ligand scaffold likely plays a key role in directing this selectivity by engaging the binding pocket containing R55 and restricting the modification of other reactive arginines throughout the proteome. Together with the reduction in CypA catalytic efficiency observed in the PPIase assay (FIG 41), these results suggest that CA-Nin-Alk (4-CA-Alk) engages CypA through its active-site arginine, consistent with a covalent mode of inhibition.
[0296] A comparative RAP analysis at 1 mM between CA-Nin-Alk (4-CA-Alk) and its regioisomer 5-CA-Alk (iso-CA-Nin-Alk), which repositions the ninhydrin warhead away relative to R55, revealed both a significantly reduced overall reactivity for 5-CA-Alk (iso-CA-Nin-Alk) at CypA R55 and a median log2 selectivity between the probes ratio of 2.49 in favor of CA-Nin-Alk (4-CA-Alk), likely driven by a combination of improved interactions with the binding pocket and better electrophile placement for improved engagement.EXAMPLE 8Discussion
[0297] Covalent probes have transformed chemical biology and drug discovery, but expanding covalent strategies to amino acids beyond cysteine and lysine has been challenging. Our study establishes ninhydrin-based probes as a selective platform for covalent arginine labeling. Our work expands the chemical proteomics toolkit, enabling new approaches to study arginine-mediated processes and develop covalent molecules that selectively and irreversibly engage targets of interest. Our ninhydrin-derived probes leverage the equilibrium between the hydrate and reactive dehydrate states of vicinal dicarbonyls. This strategy optimizes probe reactivity while minimizing non-specific interactions. Both in vitro and cell-based studies revealed that modified arginines are often partially buried rather than fully solvent-exposed, highlighting the influence of the local protein environment on reactivity. Additionally, modified arginines were frequently enriched near active sites, suggesting a potential regulatory role for selective arginine modification in enzyme function. Our 5-Nin-Alk probe exhibits distinct modification patternsAtty. Docket: UCSF-865WOcompared to pan-specific cysteine targeting covalent reagents. Reactive and hyper-reactive arginines have a consistent pKa and are only partially solvent-exposed, suggesting that other environmental factors drive arginine reactivity, which remains a major focus of our groups.
[0298] Analysis of the distribution of modified arginine distances from active sites revealed that the majority of modifications occurred within 30 A, with a peak at 20 A (FIG 5).Beyond 30 A, the frequency of modified arginines declined sharply, with few detected beyond 50 A. This distribution suggests that while active- site proximity may influence reactivity, a significant subset of modified arginines resides at intermediate distances, where they may contribute to allosteric regulation or protein-protein interactions. Furthermore, among well-resolved structures, very few modified arginines were fully solvated, suggesting that local protein environments contribute to reactivity. Alternatively, this observation may reflect a structural bias, as fully solvated surface-exposed arginines are often not well resolved in crystallographic datasets, limiting their inclusion in the analysis.
[0299] Covalent probes have transformed chemical biology and drug discovery, but broadly expanding covalent strategies to amino acids beyond cysteine and lysine has been challenging. Our study establishes ninhydrin-based probes as selective and robust tools for covalent arginine labeling, significantly expanding the chemical proteomics toolkit and enabling new approaches to study arginine-mediated processes. Our ninhydrin-derived probes leverage the equilibrium between the hydrate and reactive dehydrate states of vicinal dicarbonyls. This strategy optimizes probe reactivity while minimizing non-specific interactions and positions arginine as a promising target for broadly applicable covalent strategies. Although 5-Nin-Alk (Nin-Alk) demonstrates more complete proteome engagement than PGO (Figure 3A, 3B), the latter will be valuable in contexts where selective or attenuated reactivity is advantageous. PGO’s greater rotational freedom and lower intrinsic electrophilicity suggest that it could function analogously to chloroacetamides in cysteine-directed ABPP, providing a less globally-reactive warhead that enables competition-based ligand screening or focused mapping of highly nucleophilic arginines. Future efforts might leverage this tunability in designing arginine-targeting covalent inhibitors with site-specific, tailored electrophilicity.
[0300] Unlike cysteine and lysine, whose reactivities heavily depend on intrinsic residue pKa and high solvent accessibility, arginine residues demonstrate consistent reactivity independent ofAtty. Docket: UCSF-865WOpKa variations or extensive solvent exposure. This independence could be advantageous or disadvantageous for probe development: it allows for covalent engagement in structurally diverse and otherwise challenging environments, but privileged intrinsic reactivity is minimal for arginines compared to cysteines. Cell lysate labeling studies revealed that modified arginines are often partially buried rather than fully solvent-exposed, emphasizing that determinants of arginine reactivity remain poorly defined, unlike other amino acids, where factors such as pKa or solvent exposure more clearly predict reactivity (Figures 1C, 41 and 5D).
[0301] Arginines are frequently localized at protein-protein interaction (PPI) sites, where their covalent modification is poised to directly modulate or disrupt critical interaction surfaces. Both cation-7t interactions and salt bridges are more prevalent in proteins from thermophilic organisms, enhancing protein stability. We observe that reactive arginines frequently engage in these stabilizing interactions, although these features were not enriched in hyper- reactive arginines compared to generally reactive ones. This trend is consistent with the absence of significant differences in predicted pKa or solvent accessibility between the two groups. In our RAP dataset, we identify 290 unique arginines stabilizing PPIs within the PDB, forming 519 unique PPI pairs. Among these, 145 arginines participated in 193 cation-7t interactions, and 212 arginines were involved in 326 unique salt bridge pairs. Despite the prevalence of these specific interactions, our analysis found that fully solvated arginines are underrepresented based on their physicochemical properties. This rarity may partly reflect the inherent dynamic nature of fully solvated arginines, as their positions might only be reliably resolved when stabilized at interaction interfaces, resulting in relatively lower RS A values being captured in our bioinformatic analyses. Importantly, these interactions provide valuable insights into structural contexts that could guide rational optimization of arginine- targeted covalent probes.
[0302] Integration of experimental PDB structures significantly enhanced our structural analyses compared to predictions based on computational AlphaFold models. Although AlphaFold offers extensive near-proteome-wide coverage, these models frequently depict idealized conformations, lacking critical contextual interactions present in biological environments. By contrast, the use of PDB structures, though computationally more intensive, owing to the need for extensive sequence alignments, the handling of large cryo-EM complexes, and the presence of multiple structural entries for individual proteins, provided more realistic and biologically representative data. Specifically, experimental structures captured greater diversity in salt bridge geometries,Atty. Docket: UCSF-865WOsolvent accessibility, and interaction distances, all crucial for accurately predicting arginine reactivity (FIG 3L, 3M). This approach underscores the value of incorporating experimentally determined structures when studying residue-level reactivity, particularly for interpreting complex biophysical features that may be underrepresented in predicted models.
[0303] Recent work studying iodoacetamide-cysteine reactivity found that simple features such as 2D sequence, pKa or even data derived from AlphaFold models were not sufficient to predict a cysteine’s relative reactivity. More accurate prediction required experimentally determined structures to start to model residue-level reactivity [Boatner, L. M. et al. ACS Chem. Biol. 2025, 20 (7), 1669-1682.]. While AlphaFold offers transformative accessibility and continues to enable exciting applications such as predictive covalent docking, the current generation of models often struggles with accurate representation of sidechain packing and interaction geometries. As we expand our platform to profile reactive arginines across more cell types, the growing bulk of arginine reactivity data may make AlphaFold-derived structures increasingly useful at scale, particularly given the computational demands of mining the PDB. Future iterations of Alphafold, which are showing improvements in sidechain accuracy, may further enhance its utility for predicting covalent probe engagement.
[0304] Selective modification of functionally relevant arginines is exemplified by ninhydrin-mediated labeling of aconitase at R479. Previous studies suggested that ninhydrin inhibits aconitase catalytic activity, and this hypothesis is further supported by our RisoLFG [RAP] data. Structural analysis confirmed that R479 is critical for substrate positioning, and covalent modeling indicated that modification at this site sterically occludes substrate binding, likely contributing to inhibition. These findings suggest that selective arginine modification could serve as a powerful approach for probing enzyme function and developing covalent inhibitors targeting active-site arginines.
[0305] Beyond aconitase, the selective modification of cyclophilin A at R55 underscores the broader utility of ninhydrin-based probes in ligand-directed covalent targeting. 4-CA-Alk exhibited dose-dependent selectivity, with lower concentrations engaging a restricted subset of modified peptides, including R55 of cyclophilin A. The ligand scaffold played a crucial role in directing this selectivity while minimizing off-target interactions. The rapid development of a ~22 pM inhibitor of cyclophilin A was achieved with the design of a single lead molecule,Atty. Docket: UCSF-865WOhighlighting the efficacy of synergizing covalent chemical probes and structural biology to facilitate rapid tool compound development.
[0306] Structural comparisons with its regioisomer 5-CA-Alk (iso-CA-Nin-Alk), which mispositioned the ninhydrin warhead, further reinforced the importance of rational warhead placement in achieving site- specific engagement. The selectivity demonstrated between the regioisomeric pairs likely is assisted by subtle differences in binding affinity and probe conformation as well as proper electrophile placement. These results demonstrate how integrated computational modeling, proteomics, and structure-guided ligand design can systematically guide the development of selective covalent ligands. These findings highlight the potential for covalent arginine probes to guide future ligand design and drug discovery efforts.
[0307] This study establishes a framework for selective covalent arginine modification using ninhydrin-derived probes, integrating computational modeling, proteomics, and structure-guided ligand design. We demonstrate that ninhydrin functions as a selective arginine- modifying electrophile with distinct proteomic coverage. We also introduce a chemical proteomic approach, RisoLFG [RAP], for classifying arginine reactivity, facilitating rational probe development, identifying arginines of interest, and a future competitive screening platform for argininereactive small molecules. The successful targeting of cyclophilin A at R55 further underscores the potential of covalent arginine probes in drug discovery and functional proteomics. Future studies will focus on refining probe specificity, elucidating additional structural determinants of arginine reactivity, and expanding the application of ninhydrin-based probes to other therapeutically relevant targets. By advancing strategies for arginine-selective covalent modification, this work paves the way for novel chemical biology tools and next-generation covalent therapeutics.
Claims
Atty. Docket: UCSF-865WOCLAIMS:
1. A modified arginine-containing protein comprising: a small molecule fragment moiety, covalently bonded to an arginine residue of an arginine-containing protein, wherein a covalent bond is formed by reaction with a non-naturally occurring small molecule probe having a structure of Formula (I): F1— Q (I), wherein F1is a small molecule fragment moiety comprising an azide moiety, an alkyne moiety, a fluorophore moiety, a labeling group, or a combination thereof; and Q comprises a cyclic trione moiety.
2. The modified arginine-containing protein of claim 1, wherein Q is ninhydrin.
3. The modified arginine-containing protein of claim 1, wherein the arginine residue is attached to the small molecule fragment through an imidazolidinimine moiety.
4. The modified arginine-containing protein of claim 1, wherein the small-molecule probe has a structure according to:oo5. The modified arginine-containing protein of claim 1, wherein F1is an alkyne moiety.
6. The modified arginine-containing protein of claim 1, wherein F1is a labeling group which comprises desthiobiotin.
7. The modified arginine-containing protein of claim 1, wherein the arginine-containing protein is described herein.
8. A modified arginine-containing protein comprising: a small molecule fragment moiety, covalently bonded to an arginine residue of an arginine-containing protein, wherein a covalent bond is formed by reaction with a non-naturally occurring ligand electrophile having a structure of Formula (X): F2— Q (X), wherein F2is a small molecule fragment moiety; and Q comprises a cyclic trione moiety.
9. The modified arginine-containing protein of claim 8, wherein Q is ninhydrin.Atty. Docket: UCSF-865WO10. The modified arginine-containing protein of claim 8, wherein the arginine residue is attached to the small molecule fragment through an imidazolidinimine moiety.
11. The modified arginine-containing protein of claim 8, wherein the ligand electrophile has a structure according to:oo12. The modified arginine-containing protein of claim 8, wherein F2comprises Ci-Ce alkyl, Ci-Ce fluoroalkyl, Ci-Ce heteroalkyl, a substituted or unsubstituted C3-C6 cycloalkyl, a substituted or unsubstituted C2-C6 heterocycloalkyl, a substituted or unsubstituted aryl, or a substituted or unsubstituted heteroaryl.
13. The modified arginine-containing protein of claim 8, wherein the arginine-containing protein is described herein.
14. A compound having a structure according to Formula (la):wherein F1is a small molecule fragment moiety comprising an azide moiety, an alkyne moiety, a fluorophore moiety, a labeling group, or a combination thereof.
15. The compound of claim 14, having a structure according to:whereinZ is an alkyne, a fluorophore. or a labeling group:Y is substituted or unsubstituted Ci-Ce alkylene;Atty. Docket: UCSF-865WOX is -CH2-, O, -C(O)NH, -NH-, -OC(O)O-, triazole, sulfonate, and phosphonate; and Raand Rbare independently selected from unsubstituted methyl, ethyl, propyl, isopropyl, butyl, isobutyl, and sec-butyl.
16. A method of identifying a reactive arginine of a protein, comprising:(a) providing a protein sample comprising isolated proteins, living cells (such as primary immune cells), cell lysate, or tissue (such as blood);(b) contacting the protein sample with a probe compound of Formula (I) at a first concentration for a time sufficient for the probe compound to react with the reactive arginine of the protein sample; and(c) analyzing the proteins of the protein sample to identify the reactive arginine that bound with the probe compound at the first concentration;wherein the probe compound has a structure represented by Formula (I): F1— Q (I), wherein F1is a small molecule fragment moiety comprising an alkyne moiety, a fluorophore moiety, a labeling group, or a combination thereof; and Q comprises a cyclic trione moiety.
17. The method of claim 16, wherein F1comprises an alkyne moiety or a fluorophore moiety.
18. The method of claim 16 or 17, wherein Q is ninhydrin.
19. The method of claims 16-18, wherein the analyzing of step (c) further comprises tagging at least one arginine-containing protein-ligand complex of step (b) to generate a tagged arginine-containing protein ligand complex.
20. The method of claims 16-19, wherein the analyzing of step (c) further comprises isolating the tagged arginine-containing protein-ligand complex.
21. The method of claim 19 or 20, wherein the tagging comprises a biotin moiety.
22. The method of claim 21, wherein the biotin moiety comprises biotin or a biotin derivative, such as desthiobiotin, biotin alkyne or biotin azide.Atty. Docket: UCSF-865WO23. A method of identifying a reactive arginine of a protein, comprising:(a) providing a protein sample comprising isolated proteins, living cells, or a cell lysate and separating the protein sample into a first protein sample and a second protein sample;(b) contacting the first protein sample with a probe compound of Formula I at a first concentration for a time sufficient for the probe compound to react with a reactive arginine of the first protein sample, and contacting the second protein sample with the probe compound of Formula (I) at a second concentration for a sufficient time for the probe compound to react with a reactive arginine of the second protein sample;(c) analyzing the proteins of the first protein sample and the second protein sample of step (b) to identify the reactive arginines that bound with the probe compound;(d) comparing the identity of the reactive arginines of step (c) from the first protein sample at the first concentration of probe compound to the reactive arginines from the second protein sample at the second concentration of probe compound; and(e) based on step (d), determining a reactive arginine of a protein;wherein the probe compound has a structure represented by Formula (I): F1— Q (I), wherein F1is a small molecule fragment moiety comprising an alkyne moiety, a fluorophore moiety, a labeling group, or a combination thereof; and Q comprises a cyclic trione moiety.
24. The method of claim 23, wherein F1comprises an alkyne moiety or a fluorophore moiety.
25. The method of claim 23 or 24, wherein Q is ninhydrin.
26. The method of claims 23-25, wherein the analyzing of step (c) further comprises tagging at least one arginine-containing protein-ligand complex of step (b) to generate a tagged arginine-containing protein ligand complex.
27. The method of claims 23-26, wherein the analyzing of step (c) further comprises isolating the tagged arginine-containing protein-ligand complex.
28. The method of claim 26 or 27, wherein the tagging comprises attaching a biotin moiety.
29. The method of claim 28, wherein the biotin moiety comprises biotin or a biotin derivative.Atty. Docket: UCSF-865WO30. The method of claim 29, wherein the biotin derivative comprises desthiobiotin, biotin alkyne or biotin azide.
31. A method of identifying a protein that interacts with a ligand of interest, comprising:(a) providing a protein sample comprising isolated proteins, living cells, or a cell lysate and separating the protein sample into a first protein sample and a second protein sample;(b) contacting the first protein sample with a ligand for a sufficient time for the ligand to react with a reactive arginine of the first protein sample;(c) contacting the first protein sample and the second protein sample with a probe compound of Formula (I) for a sufficient time for the probe compound to react with the reactive arginines of the first and second protein samples;(d) analyzing the proteins of the first and second protein samples to identify the reactive arginines that bound with the probe compound;(e) comparing the reactivity of the reactive arginine from the first protein sample to the reactivity of the reactive arginine from the second protein sample, wherein a decrease in the reactivity of the reactive arginine of the first protein sample relative to the reactive arginine of the second protein sample indicates interaction of the ligand with the reactive arginine of the first protein sample; and(f) determining the protein comprising the reactive arginine of the first protein sample that interacts with the ligand; wherein the probe compound has a structure represented by Formula (I): F1— Q (I), wherein F1is a small molecule fragment moiety comprising an alkyne moiety, a fluorophore moiety, a labeling group, or a combination thereof; and Q comprises a cyclic trione moiety.
32. The method of claim 31, wherein the ligand in step (b) comprises a small molecule compound, a polynucleotide, a polypeptide or its fragments thereof, or a peptidomimetic.
33. The method of claim 31, wherein the small molecule compound comprises a ligandelectrophile compound that has a structure represented by Formula (X): F2— Q (X), wherein F2is a small molecule fragment moiety; and Q comprises a cyclic trione moiety.
34. The method of claim 31, wherein the ligand in step (b) comprises a polypeptide or its fragments thereof.Atty. Docket: UCSF-865WO35. The method of claims 31-34, wherein the analyzing of step (d) further comprises tagging at least one arginine-containing protein-ligand complex of step (c) to generate a tagged arginine-containing protein ligand complex.
36. The method of claims 31-34, wherein the analyzing of step (d) further comprises isolating the tagged arginine-containing protein-ligand complex.
37. The method of claim 35 or 36, wherein the tagging comprises attaching a biotin moiety.
38. The method of claim 37, wherein the biotin moiety comprises biotin or a biotin derivative.
39. The method of claim 38, wherein the biotin derivative comprises desthiobiotin, biotin alkyne or biotin azide.
40. The method of claim 38, wherein the biotin derivative comprises desthiobiotin.