Unnatural amino acids, bioreactive proteins, and uses thereof

The introduction of unnatural amino acids like PFY and PFK into proteins addresses the need for new tools in protein identification and therapeutics, enhancing bioreactive functions and drug development capabilities.

WO2025128629A1PCT designated stage expired Publication Date: 2025-06-19RGT UNIV OF CALIFORNIA
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
PCT/US2024/059463
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-12-28
Filing Date
2024-12-11
Publication Date
2025-06-19

AI Technical Summary

Technical Problem

There is a need for new unnatural amino acids that can be used for protein identification, drug target discovery, or biotherapeutics, as existing technologies are limited in this regard.

Method used

The development of compounds containing unnatural amino acids with specific side chains, such as PFY and PFK, which can be genetically incorporated into proteins, enabling novel protein-based therapeutics and bioreactive functions.

Benefits of technology

These unnatural amino acids facilitate protein cross-linking, drug development, and bioreactive protein synthesis, expanding the capabilities in protein-based therapeutics and biotechnology.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2024059463_19062025_PF_FP_ABST
    Figure US2024059463_19062025_PF_FP_ABST
Patent Text Reader

Abstract

Provided herein are, inter alia, unnatural amino acids, proteins comprising unnatural amino acids, protein conjugates, and methods of making the unnatural amino acids, proteins, and protein conjugates. In embodiments, the unnatural amino acids are compounds having the following formula or a stereoisomer thereof: wherein Z is sulfur or phosphorous, and the remaining substituents are as described herein.
Need to check novelty before this filing date? Find Prior Art

Description

PATENT Docket No.: 048536-780001WO SF-2024-080 UNNATURAL AMINO ACIDS, BIOREACTIVE PROTEINS, AND USES THEREOF CROSS-REFERENCE TO RELATED APPLICATIONS

[0001] This application claims the benefit of priority to US Application No.63 / 615,691 filed December 28, 2023, and US Application No.63 / 608,411 filed December 11, 2023, the disclosures of which are incorporated by reference herein in their entirety. STATEMENT AS TO RIGHTS TO INVENTIONS MADE UNDER FEDERALLY SPONSORED RESEARCH AND DEVELOPMENT

[0002] This invention was made with government support under grant R01 GM118384 awarded by The National Institutes of Health. The government has certain rights in the invention. REFERENCE TO A "SEQUENCE LISTING," A TABLE, OR A COMPUTER PROGRAM LISTING APPENDIX SUBMITTED AS AN ASCII FILE

[0003] The Sequence Listing written in XML file entitled “048536-780001WO-SL.xml” created on November 28, 2024, and having 60,638 bytes, is incorporated by reference herein. BACKGROUND

[0004] Introducing new chemical bonds into proteins provides innovative avenues for manipulating protein structure and function. Unnatural amino acids (Uaas) containing diverse latent bioreactive functional groups have recently been introduced into proteins via genetic code expansion. This offers an exquisite tool not only to study cellular protein interactions but also create novel protein-based therapeutics. SuFEx click chemistry via the latent aryl fluorosulfate group has demonstrated value in aiding modular organic synthesis, chemical biology, and drug development. As set forth in US Publication No.2021 / 0002325, the inventors incorporated fluorosulfate-L-tyrosine (FSY) into proteins for protein crosslinking and generating covalent protein drugs. There is a need in the art, inter alia, for new and other unnatural amino acids that can be used for protein identification, drug target discovery, or biotherapeutics. Provided herein are solutions to these and other needs in the art. SUMMARY

[0005] Provided herein are compounds having the following structures or stereoisomers thereof:nd wherein Z is sulfur o -, or –O-; x is an integer from 1 to 8; x1 is an integer from 0 to 5; L1is a bond, substituted or unsubstituted alkylene, or substituted or unsubstituted heteroalkylene; R1is hydrogen or an electron withdrawing group; and R2and R3are each independently substituted or unsubstituted C1-5alkyl. In embodiments, Z is sulfur. In embodiments, Z is phosphorous.

[0006] Provided herein are proteins comprising an unnatural amino acid, wherein the unnatural amino comprises a side chain having the following structures: ,integer from 1 to 8; x1 is an integer from 0 to 5; L1is a bond, substituted or unsubstituted alkylene, or substituted or unsubstituted heteroalkylene; R1is hydrogen or an electron withdrawing group; and R2and R3are each independently substituted or unsubstituted C1-5 alkyl. In embodiments, Z is sulfur. In embodiments, Z is phosphorous.

[0007] Provided herein are protein conjugates having the following structures: ;is a bond, -(CH2)1-5-, -O-(CH2)1-5-, or –O-; x is an integer from 1 to 8; x1 is an integer from 0 to 5; L1is a bond, substituted or unsubstituted alkylene, or substituted or unsubstituted heteroalkylene; and R1is hydrogen or an electron withdrawing group; R2and R3are each independently substituted or unsubstituted C1-5alkyl; and L2and L3are as defined herein. In embodiments, Z is sulfur. In embodiments, Z is phosphorous. In embodiments, L2is a bond andL3-R5.

[0008] BRIEF DESCRIPTION OF THE DRAWINGS

[0009] FIGS.1A-1H show the design, synthesis, and genetic incorporation of PFY into proteins in E. coli and mammalian cells. FIG.1A: Structure of PFY. FIG.1B: PFY to react with nucleophilic residues in close proximity via proximity-enabled PFEx reactivity. FIG.1C: Chemical synthesis of PFY. FIG.1D: SDS-PAGE and Western blot analysis of PFY incorporation into Afb(36TAG) by the tRNAPyl / PFYRS in E. coli. FIG.1E: Bright field and fluorescence microscopic images of HEK-293T cells. The cells were transfected with the tRNAPyl / PFYRS and the EGFP(182TAG) gene, and grown in the absence and presence of PFY. Scale bar, 100 µm. FIG.1F: Western blot analysis of HEK-293T cells in (J). FIG.1G: Flow cytometric analysis of HEK-293T cells, which were transfected with the tRNAPyl / PFYRS and the EGFP(182TAG) gene and grown in the presence of various concentrations of PFY for 24 h and 48 h. FIG.1H: Flow cytometric quantification of fluorescence intensity of the HEK-239T cells in (I) incubated with PFY for 48 h.

[0010] FIGS.2A-2E show that PFY reacts with proximal His, Tyr, Lys, and Cys in proteins through PFEx. In embodiments, the reaction occurs at an acidic pH. In embodiments, the acidic pH is from 5 to 6.9. In embodiments, the acidic pH is from 5.5 to 6.5. FIG.2A: Structure of the Afb-Z complex (PDB code: 1LP1) showing D36 in Afb for PFY incorporation and N6 in Z protein for mutation to different residues. FIG.2B: Western blot analysis of Afb(36PFY) protein incubation with Z(6X) protein. X represents mutated residues. FIG.2C: SDS-PAGE analysis of Afb(36PFY) protein incubation with Z(6X) protein. FIG.2D: Structure of ecGST (PDB code: 1A0F) showing residue T103 in one monomer for PFY incorporation and the proximal residue H106 in the other monomer for mutations. FIG.2E: Western blot analysis of ecGST dimeric cross-linking in HEK-293T cells. ecGST(103PFY / 106X) was expressed in HEK-293T cells and the cell lysate was probed with anti-Hisx6 antibody to detect the Hisx6 tag appended at the C-terminus of ecGST.

[0011] FIGS.3A-3D show that basic pH increases PFY-Tyr cross-linking but decreases PFY- His cross-linking. FIG.3A: SDS-PAGE analysis of Afb(36PFY) cross-linking with MBP-Z(6X) at pH 7.4 and pH 8.8. FIG.3B: Cross-linking efficiency between Afb(36PFY) and MBP- Z(6His) or MBP-Z(6Tyr) at pH 7.4 and pH 8.8. The line and error bar represent mean ^ SEM; n = 6 independent experiments. **p < 0.01; ****p < 0.001; unpaired t-test. FIG.3C: SDS-PAGEanalysis Afb(36PFY) or Afb(36FSY) cross-linking with MBP-Z(6His) at various pH. FIG.3D: changes in cross-linking efficiency with varying pH for reaction of Afb(36PFY) or Afb(36FSY) with MBP-Z(6His). The line and error bar represents mean + / - SEM; n=5 independent experiments.

[0012] FIGS.4A-4F show that temperature affects PFY reaction with Tyr and His differently. FIG.4A: Scheme showing how Afb(36PFY) was incubated with MBP-Z(6X) and then treated at different temperature. FIG.4B: SDS-PAGE analysis of Afb(36PFY) cross-linking with MBP- Z(6X) under different incubation and treatment temperatures. FIG.4C: Cross-linking efficiency of Afb(36PFY) with MBP-Z(6Tyr) at different temperatures measured from (B) using densitometry. FIG.4D: Cross-linking efficiency of Afb(36PFY) with MBP-Z(6His) at different temperatures measured from (B) using densitometry. For FIGS.4C-4D, the line and error bar represent mean ^ SEM; n = 6 independent experiments. ns, not significant; ****p < 0.0001; unpaired t-test. FIG.4E: thermal sensitivity of the P(V)-N linkage resulting from PFY reaction with His. Top: SDS-PAGE analysis of the Afb(36PFY) / MBP-Z(6His) cross-linking product after it was subjected to various temperatures for durations of 5 and 10 minutes. Bottom: the cross-linking efficiency of Afb(36PFY) with MBP-Z(6His) was evaluated after their cross- linked product was exposed to various temperatures for 5 and 10 min. The line and error bar represents mean + / - SEM; n=3 independent experiments. FIG.4F: PFY exhibited enhanced durability compared with FSY in proteins. The figure shows the remaining activity of AFb(36PFY) and Afb(36FSY) after incubation at 37°C for the indicated number of days. The cross-linking efficiency of Afb(36PFY) and Afb(36FSY) with MBP-Z(6His) was measured to quantify their remaining reactivity. The line and error bar represents mean + / - SEM; n=8 independent experiments.

[0013] FIGS.5A-5G show that Na2SiO3 increases PFY reaction with Tyr and Cys but decreases PFY reaction with His.FIG.5A: Structure of the Afb-Z complex (PDB code: 1LP1) showing E24 in Z protein for PFY incorporation and K7 in Afb for mutation to Cys, Tyr, or His. FIGS.5B-5D: SDS-PAGE analysis (top panel) and quantification of MBP-Z(24PFY) cross- linking with Afb(7X) (bottom panel) in the presence of different concentration of Na2SiO3. (B) Afb(7Cys); (C) Afb(7Tyr); (D) Afb(7His). FIG.5E: Structure of the Afb-Z complex (PDB code: 1LP1) showing D36 in Afb for PFY incorporation and N6 in Z protein for mutation to Tyr, or His. FIG.5F-5G: SDS-PAGE analysis (top panel) and quantification of Afb(36PFY) cross-linking with MBP-Z(6X) (bottom panel) in the presence of different concentration of Na2SiO3. (F) MPB-Z(6Tyr); (G) MPB-Z(6His). The line and error bar represent mean ± SEM; n = 6 independent experiments. ns, not significant; *p < 0.05; **p < 0.01; ***p < 0.001; ****p <0.0001; unpaired t-test.

[0014] FIGS.6A-6I show that genetic incorporation of PFK expands protein cross-linking unreachable by PFY in vitro and in cells. FIG.6A: Structure of PFK. FIG.6B: Western blot analysis of PFK incorporation in mNb6(54TAG) in E. coli cells. FIG.6C: Western blot analysis of PFK incorporation in EGFP(182TAG) in HEK-293T cells. FIG.6D: Structure of nanobody mNb6 binding with the Spike protein of SARS-CoV-2 (PDB code 7KKL). Sites 50-59 for PFK incorporation are colored in purple. R54 in mNb6 and the target residue Y351 in the Spike are shown in stick. FIG.6E: Western blot analysis of cross-linking between the Spike protein’s receptor binding domain (RBD) and mNb6 mutants with PFK incorporated at the indicated sites. FIG.6F: Western blot comparison between mNb6(54PFK) and mNb6(54PFY) for cross-linking with the Spike protein’s RBD. FIG.6G: Structure of ecGST (PDB code: 1A0F) showing residue T103 in one monomer for PFK / PFY incorporation and the proximal residue C10 in the other monomer for mutations. FIG.6H: Western blot analysis of HEK-293T cell lysate of cells expressing ecGST mutants with PFK / PFY incorporated at site 103 and different mutations at site 10. FIG.6I: Quantification of the cross-linked dimer to monomer ratio of ecGST in (I), showing the difference between PFK and PFY in cross-linking His, Tyr, and Cys. The line and error bar represent mean ^ SEM; n = 4 independent experiments. **p < 0.01; ***p < 0.001; unpaired t- test.

[0015] FIG.7 shows the chemical synthesis of PFK.

[0016] FIGS.8A-8B show that PFY was nontoxic to E. coli and HEK-293T cells at 4 mM and 2 mM concentration, respectively. FIG.8A: Growth curve of E. coli DH10B cells at 37 ^C in the absence or presence of different concentrations of PFY in the growth media. n=3 independent experiments; error bars represent SEM. FIG.8B: Cell viability assay for HEK- 293T cells incubated with various concentrations of PFY. n = 12, error bars represent SEM.

[0017] FIG.9 shows the suppression of TAG in the EGFP-182TAG gene in E. coli by tRNAPyl / mFSYRS or tRNAPyl / NpYRS in the presence of different concentrations of PFY. OD600-normalized fluorescence intensity of EGFP measured from cells showed that mFSYRS was able to incorporate PFY more efficiently than NpYRS.

[0018] FIGS.10A-10C show that PFY reactivity in proteins was proximity driven. FIG.10A: Structure of nanobody SR4 in complex with the Spike protein of SARS-CoV-2 (PDB code: 7C8V). Sites S57 and H54 in SR4 for PFY incorporation and the target Y505 in the Spike protein are shown in stick.FIGS.10B-10C: SDS-PAGE analysis of the cross-linking of the Spike protein’s receptor binding domain (RBD) with either SR4(54PFY) or SR4(57PFY).

[0019] FIGS.11A-11B show the P(V)-S linkage resultant from PFY-Cys reaction was stableat 95 °C. FIG.11A: SDS-PAGE analysis of MBP-Z(24PFY) cross-linking with Afb(7Cys).95 °C sample, 37 °C sample, and 4 °C sample were prepared using the same incubation and treatment temperatures as described in FIG 4A. Na2SiO3(2 mM) was added to boost PFY reaction with Cys as described in FIG 5. FIG.11B: Cross-linking efficiency of MBP-Z(24PFY) with Afb(7Cys) at different temperatures measured from (FIG.11A) using densitometry. The line and error bar represent mean ^ SEM; n = 4 independent experiments. ns, not significant; ***p < 0.001; ****p < 0.0001; unpaired t-test.

[0020] FIG.12 shows flow cytometric analysis of PFK incorporation into GFP in HeLa cells. HeLa-GFP(182TAG) reporter cells were transfected with the tRNAPyl / PFKRS genes and grown in the absence or presence of 1 mM PFK for 24 h.

[0021] FIG.13 is an SDS-PAGE analysis of affibody(36PFY) at left and affibody(36PFY / 32R) at right crosslinking with MBP-Z (N6Y) protein.

[0022] FIGS.14A-14B show primers for cloning as described in the examples. DETAILED DESCRIPTION

[0023] Definitions

[0024] Unless defined otherwise, technical and scientific terms used herein have the same meaning as commonly understood by a person of ordinary skill in the art. See, e.g., Singleton et al., Dictionary of Microbiology and Molecular Biology, 2nd ed., J. Wiley & Sons (New York, NY 1994); Sambrook et al., Molecular Cloning, A Laboratory Manual, Cold Springs Harbor Press (Cold Springs Harbor, NY 1989). Any methods, devices and materials similar or equivalent to those described herein can be used in the practice of this disclosure. The following definitions are provided to facilitate understanding of certain terms used frequently herein and are not meant to limit the scope of the present disclosure.

[0025] The term “PFY” refers to a compound having the following structure: .

[0026] The term “PFK”structure:H2N 2 .

[0027] In can be writtenwithout identification of the stereoisomer, however, the stereoisomer should be considered as encompassed by the compound in each instance, whether or not it is specifically stated. Thus the compounds described herein (e.g., PFY and PFK) have the following group: which can alternatively be described by the .“antibody” is used according to its commonlyart. Antibodies exist, e.g., as intact immunoglobulins or as a number of well-characterized fragments produced by digestion with various peptidases. Thus, for example, pepsin digests an antibody below the disulfide linkages in the hinge region to produce F(ab)'2, a dimer of Fab which itself is a light chain joined to VH-CH1by a disulfide bond. The term “F(ab)'2” is used interchangeably with “Fab dimer.” The F(ab)'2 may be reduced under mild conditions to break the disulfide linkage in the hinge region, thereby converting the F(ab)'2dimer into an Fab' monomer. The Fab' monomer is essentially Fab with part of the hinge region (see Fundamental Immunology (Paul ed., 3d ed.1993)). The term “Fab’ monomer” is used interchangeably with “Fab” and “or an antigen-binding fragment.” While various antibody fragments are defined in terms of the digestion of an intact antibody, one of skill will appreciate that such fragments may be synthesized de novo either chemically or by using recombinant DNA methodology. Thus, the term antibody, as used herein, also includes antibody fragments either produced by the modification of whole antibodies, or those synthesized de novo using recombinant DNA methodologies (e.g., single chain Fv) or those identified using phage display libraries (e.g., McCafferty et al., Nature 348:552-554 (1990)).

[0029] Antibodies are large, complex proteins with an intricate internal structure. A natural antibody molecule contains two identical pairs of polypeptide chains, each pair having one lightchain and one heavy chain. Each light chain and heavy chain in turn consists of two regions: avariable (“V”) region involved in binding the target antigen, and a constant (“C”) region that interacts with other components of the immune system. The light and heavy chain variable regions come together in 3-dimensional space to form a variable region that binds the antigen(for example, a receptor on the surface of a cell). Within each light or heavy chain variable region, there are three short segments (averaging 10 amino acids in length) called the complementarity determining regions (“CDRs”). The six CDRs in an antibody variable domain (three from the light chain and three from the heavy chain) fold up together in 3-dimensional space to form the actual antibody binding site which docks onto the target antigen. The position and length of the CDRs have been precisely defined by Kabat et al, Sequences of Proteins of Immunological Interest, U.S. Department of Health and Human Services, 1987. The part of a variable region not contained in the CDRs is called the framework (“FR”), which forms the environment for the CDRs.

[0030] An “antibody variant” as provided herein refers to a polypeptide capable of binding to a receptor protein or an antigen and including one or more structural domains of an antibody or fragment thereof. Non-limiting examples of antibody variants include single-domain antibodies (nanobodies), affibodies (polypeptides smaller than monoclonal antibodies and capable of binding receptor proteins or antigens with high affinity and imitating monoclonal antibodies), antigen-binding fragments (Fab), Fab dimers (monospecific Fab2, bispecific Fab2), trispecific Fab3, monovalent IgGs, single-chain variable fragments (scFv), bispecific diabodies, trispecific triabodies, scFv-Fc, minibodies, IgNAR, V-NAR, hcIgG, VhH, and peptibodies. A “peptibody” as provided herein refers to a peptide moiety attached (through a covalent or non-covalent linker) to the Fc domain of an antibody.

[0031] A “single-domain antibody” or “nanobody” refers to an antibody fragment having a single monomeric variable antibody domain. Like a whole antibody, it is able to bind selectively to a specific antigen. In embodiments, the single domain antibody is a human or humanized single-domain antibody.

[0032] A single-chain variable fragment (scFv) is typically a fusion protein of the variable regions of the heavy (VH) and light chains (VL) of immunoglobulins, connected with a short linker peptide of 10 to about 25 amino acids. The linker is usually rich in glycine for flexibility, as well as serine or threonine for solubility. The linker can either connect the N-terminus of the VH with the C-terminus of the VL, or vice versa.

[0033] The term “target protein” refers to a targeting molecule having, e.g., a regulatory role in a cell. In embodiments, a “target protein” is a receptor protein, a cytosolic protein, a transcriptional factor, or an enzyme. In embodiments, a receptor protein is an extracellular domain receptor protein, a transmembrane domain receptor protein, or an intracellular domain receptor protein. In embodiments, the “target protein” is the binding target of the antibody or antibody variant described herein. In embodiments, the protein (e.g., antibody or antibodyvariant) described herein is capable of inhibiting or activating the biological activity of the target protein upon binding. In embodiments, the activity of the target protein is increased or decreased.

[0034] “Receptor protein” or “membrane receptor” refers to a receptor (protein) that is embedded in the plasma membrane of a cell. In embodiments, the receptor protein is located in the extracellular domain of a cell, the transmembrane domain of a cell, or the intracellular domain of a cell. In embodiments, the receptor protein is a cell-surface receptor. In embodiments, the receptor protein is in the extracellular domain. In embodiments, the receptor protein is in the transmembrane domain. In embodiments, the receptor protein is an ion channel- linked receptor, an enzyme-linked receptor, or a G protein-coupled receptor. In embodiments, the receptor protein is a hormone receptor.

[0035] The term “peptidyl moiety” as used herein refers to a protein, protein fragment, or peptide that may form part of a biomolecule or a biomolecule conjugate. In aspects, the peptidyl moiety forms part of a biomolecule (e.g., protein). In aspects, the peptidyl moiety forms part of a biomolecule (e.g., protein) conjugate. The peptidyl moiety may also be substituted with additional chemical moieties (e.g., additional R substituents). In aspects, the peptidyl moiety forms part of an antibody or an antibody variant. In aspects, the peptidyl moiety forms part of a receptor protein. In aspects, a peptidyl moiety is a protein, protein fragment, or peptide that contains a monovalent radical of an amino acid.

[0036] “Nucleic acid” refers to nucleotides (e.g., deoxyribonucleotides or ribonucleotides) and polymers thereof in either single-, double- or multiple-stranded form, or complements thereof. The terms “polynucleotide,” “oligonucleotide,” “oligo” or the like refer, in the usual and customary sense, to a linear sequence of nucleotides. The term “nucleotide” refers, in the usual and customary sense, to a single unit of a polynucleotide, i.e., a monomer. Nucleotides can be ribonucleotides, deoxyribonucleotides, or modified versions thereof. Examples of polynucleotides contemplated herein include single and double stranded DNA, single and double stranded RNA, and hybrid molecules having mixtures of single and double stranded DNA and RNA. Examples of nucleic acid, e.g. polynucleotides contemplated herein include any types of RNA, e.g. mRNA, siRNA, miRNA, and guide RNA and any types of DNA, genomic DNA, plasmid DNA, and minicircle DNA, and any fragments thereof. The term “duplex” in the context of polynucleotides refers, in the usual and customary sense, to double strandedness. Nucleic acids can be linear or branched. For example, nucleic acids can be a linear chain of nucleotides or the nucleic acids can be branched, e.g., such that the nucleic acids comprise one or more arms or branches of nucleotides. Optionally, the branched nucleic acids are repetitivelybranched to form higher ordered structures such as dendrimers and the like.

[0037] A polynucleotide is typically composed of a specific sequence of four nucleotide bases: adenine (A); cytosine (C); guanine (G); and thymine (T) (uracil (U) for thymine (T) when the polynucleotide is RNA). Thus, the term “polynucleotide sequence” is the alphabetical representation of a polynucleotide molecule; alternatively, the term may be applied to the polynucleotide molecule itself. This alphabetical representation can be input into databases in a computer having a central processing unit and used for bioinformatics applications such as functional genomics and homology searching. Polynucleotides may optionally include one or more non-standard nucleotide(s), nucleotide analog(s) and / or modified nucleotides.

[0038] The term “amino acid” refers to naturally occurring and synthetic amino acids, as well as amino acid analogs and amino acid mimetics that function in a manner similar to the naturally occurring amino acids. Naturally occurring amino acids are those encoded by the genetic code, as well as those amino acids that are later modified, e.g., hydroxyproline, γ-carboxyglutamate, and O-phosphoserine. Amino acid analogs refers to compounds that have the same basic chemical structure as a naturally occurring amino acid, i.e., an α carbon that is bound to a hydrogen, a carboxyl group, an amino group, and an R group, e.g., homoserine, norleucine, methionine sulfoxide, methionine methyl sulfonium. Such analogs have modified R groups (e.g., norleucine) or modified peptide backbones, but retain the same basic chemical structure as a naturally occurring amino acid. Amino acid mimetics refers to chemical compounds that have a structure that is different from the general chemical structure of an amino acid, but that functions in a manner similar to a naturally occurring amino acid. The terms “non-naturally occurring amino acid” and “unnatural amino acid” refer to amino acid analogs, synthetic amino acids, and amino acid mimetics which are not found in nature. In embodiments, the unnatural amino acid is an amino acid described herein, including embodiments thereof. Amino acids may be referred to herein by either their commonly known three letter symbols or by the one-letter symbols recommended by the IUPAC-IUB Biochemical Nomenclature Commission. Nucleotides, likewise, may be referred to by their commonly accepted single-letter codes.

[0039] The term “amino acid side chain” refers to the functional substituent contained on amino acids. For example, an amino acid side chain may be the side chain of a naturally occurring amino acid. Naturally occurring amino acids are those encoded by the genetic code (e.g., alanine, arginine, asparagine, aspartic acid, cysteine, glutamine, glutamic acid, glycine, histidine, isoleucine, leucine, lysine, methionine, phenylalanine, proline, serine, threonine,tryptophan, tyrosine, or valine), as well as those amino acids that are later modified, e.g.,hydroxyproline, γ-carboxyglutamate, and O-phosphoserine. In aspects, the amino acid side chainmay be a non-natural amino acid side chain. In aspects, the amino acid side chain is H, ,refers to the functional substituent of compounds that have the same basic chemical structure as a naturally occurring amino acid, i.e., an α carbon that is bound to a hydrogen, a carboxyl group,an amino group, and an R group, e.g., homoserine, norleucine, methionine sulfoxide, methioninemethyl sulfonium, allylalanine, 2-aminoisobutryric acid. Non-natural amino acids are non- proteinogenic amino acids that either occur naturally or are chemically synthesized. Such analogs have modified R groups (e.g., norleucine) or modified peptide backbones, but retain the same basic chemical structure as a naturally occurring amino acid. Non-limiting examples include exo-cis-3-aminobicyclo[2.2.1]hept-5-ene-2-carboxylic acid hydrochloride, cis-2- aminocycloheptane-carboxylic acid hydrochloride, cis-6-amino-3-cyclohexene-1-carboxylic acid hydrochloride, cis-2-amino-2-methylcyclohexanecarboxylic acid hydrochloride, cis-2- amino-2-methylcyclopentane-carboxylic acid hydrochloride, 2-(Boc-aminomethyl)benzoic acid, 2-(Boc-amino)octanedioic acid, Boc-4,5-dehydro-Leu-OH (dicyclohexylammonium), Boc-4- (Fmoc-amino)-L-phenylalanine, Boc-β-Homopyr-OH, Boc-(2-indanyl)-Gly-OH, 4-Boc-3- morpholineacetic acid, 4-Boc-3-morpholine acetic acid, Boc-pentafluoro-D-phenylalanine, Boc- pentafluoro-L-phenylalanine, Boc-Phe(2-Br)-OH, Boc-Phe(4-Br)-OH, Boc-D-Phe(4-Br)-OH, Boc-D-Phe(3-Cl)-OH , Boc-Phe(4-NH2)-OH, Boc-Phe(3-NO2)-OH, Boc-Phe(3,5-F2)-OH, 2- (4-Boc-piperazino)-2-(3,4-dimethoxy-phenyl)acetic acid purum, 2-(4-Boc-piperazino)-2-(2- fluorophenyl)acetic acid purum, 2-(4-Boc-piperazino)-2-(3-fluorophenyl)acetic acid purum, 2- (4-Boc-piperazino)-2-(4-fluorophenyl)acetic acid purum, 2-(4-Boc-piperazino)-2-(4-methoxy- phenyl)acetic acid purum, 2-(4-Boc-piperazino)-2-phenylacetic acid purum, 2-(4-Boc-piperazino)-2-(3-pyridyl)acetic acid purum, 2-(4-Boc-piperazino)-2-[4-(trifluoromethyl)phenyl]- acetic acid purum, Boc-β-(2-quinolyl)-Ala-OH, N-Boc-1,2,3,6-tetrahydro-2-pyridinecarboxylic acid, Boc-β-(4-thiazolyl)-Ala-OH, Boc-β-(2-thienyl)-D-Ala-OH, Fmoc-N-(4-Boc-aminobutyl)- Gly-OH, Fmoc-N-(2-Boc-aminoethyl)-Gly-OH , Fmoc-N-(2,4-dimethoxybenzyl)-Gly-OH, Fmoc-(2-indanyl)-Gly-OH, Fmoc-pentafluoro-L-phenylalanine, Fmoc-Pen(Trt)-OH, Fmoc- Phe(2-Br)-OH, Fmoc-Phe(4-Br)-OH, Fmoc-Phe(3,5-F2)-OH, Fmoc-β-(4-thiazolyl)-Ala-OH, Fmoc-β-(2-thienyl)-Ala-OH, 4-(Hydroxymethyl)-D-phenylalanine.

[0041] The following eight groups each contain amino acids that are conservative substitutions for one another: (i) Alanine (A), Glycine (G); (ii) Aspartic acid (D), Glutamic acid (E); (iii) Asparagine (N), Glutamine (Q); (iv) Arginine (R), Lysine (K); (v) Isoleucine (I), Leucine (L), Methionine (M), Valine (V); (vi) Phenylalanine (F), Tyrosine (Y), Tryptophan (W); (vii) Serine (S), Threonine (T); and (viii) Cysteine (C), Methionine (M).

[0042] The terms “polypeptide,” “peptide” and “protein” are used interchangeably herein to refer to a polymer of amino acid residues. The polymer of amino acids may, in embodiments, be conjugated to a moiety that does not consist of amino acids. The terms apply to amino acid polymers in which one or more amino acid residue is an artificial chemical mimetic of a corresponding naturally occurring amino acid, as well as to naturally occurring amino acid polymers and non-naturally occurring amino acid polymers. A “fusion protein” refers to a chimeric protein encoding two or more separate protein sequences that are recombinantly expressed as a single moiety.

[0043] An amino acid or nucleotide base “position” is denoted by a number that sequentially identifies each amino acid (or nucleotide base) in the reference sequence based on its position relative to the N-terminus (or 5'-end). Due to deletions, insertions, truncations, fusions, and the like that must be taken into account when determining an optimal alignment, in general the amino acid residue number in a test sequence determined by simply counting from the N- terminus will not necessarily be the same as the number of its corresponding position in the reference sequence. For example, in a case where a variant has a deletion relative to an aligned reference sequence, there will be no amino acid in the variant that corresponds to a position in the reference sequence at the site of deletion. Where there is an insertion in an aligned reference sequence, that insertion will not correspond to a numbered amino acid position in the reference sequence. In the case of truncations or fusions there can be stretches of amino acids in either the reference or aligned sequence that do not correspond to any amino acid in the corresponding sequence.

[0044] The terms “numbered with reference to” or “corresponding to,” when used in thecontext of the numbering of a given amino acid or polynucleotide sequence, refers to the numbering of the residues of a specified reference sequence when the given amino acid or polynucleotide sequence is compared to the reference sequence.

[0045] “Percentage of sequence identity” is determined by comparing two optimally aligned sequences over a comparison window, wherein the portion of the polynucleotide or polypeptide sequence in the comparison window may comprise additions or deletions (i.e., gaps) as compared to the reference sequence (which does not comprise additions or deletions) for optimal alignment of the two sequences. The percentage is calculated by determining the number of positions at which the identical nucleic acid base or amino acid residue occurs in both sequences to yield the number of matched positions, dividing the number of matched positions by the total number of positions in the window of comparison and multiplying the result by 100 to yield the percentage of sequence identity.

[0046] The terms “identical” or percent “identity,” in the context of two or more nucleic acids or polypeptide sequences, refer to two or more sequences or subsequences that are the same or have a specified percentage of amino acid residues or nucleotides that are the same (i.e., about 60% identity, or at least 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% identity over a specified region, when compared and aligned for maximum correspondence over a comparison window or designated region) as measured using a BLAST or BLAST 2.0 sequence comparison algorithms with default parameters described below, or by manual alignment and visual inspection (e.g., NCBI web site ncbi.nlm.nih.gov / BLAST / or the like). Such sequences are then said to be “substantially identical.” This definition also refers to, or may be applied to, the compliment of a test sequence. The definition also includes sequences that have deletions and / or additions, as well as those that have substitutions. As described below, the preferred algorithms can account for gaps and the like. Preferably, identity exists over a region that is at least about 25 amino acids or nucleotides in length, or more preferably over a region that is 50-100 amino acids or nucleotides in length.

[0047] The term “pyrrolysyl-tRNA synthetase” refers to an enzyme (including homologs, isoforms, and functional fragments thereof) with pyrrolysyl-tRNA synthetase activity. Pyrrolysyl-tRNA synthetase is an aminoacyl-tRNA synthetase that catalyzes the reaction necessary to attach α-amino acid pyrrolysine to the cognate tRNA (tRNApyl), thereby allowing incorporation of pyrrolysine during proteinogenesis at amber stop codons (i.e., UAG). The term includes any recombinant or naturally-occurring form of pyrrolysyl-tRNA synthetase or variants, homologs, or isoforms thereof that maintain pyrrolysyl-tRNA synthetase activity (e.g. within at least 30%, 40%, 50%, 60%, 70%, 80%, 90%, 95%, or 100% activity compared to wild-type pyrrolysyl-tRNA synthetase). In embodiments, the variants, homologs, or isoforms have at least 90%, 95%, 96%, 97%, 98%, 99% or 100% amino acid sequence identity across the whole sequence or a portion of the sequence (e.g., a 50, 100, 150 or 200 continuous amino acid portion) compared to a naturally occurring pyrrolysyl-tRNA synthetase. In embodiments, the mutant pyrrolysyl-tRNA synthetase catalyzes the attachment of the compound of Formula (1) and embodiments thereof to a tRNApyl. In embodiments, the mutant pyrrolysyl-tRNA synthetase catalyzes the attachment of the compound of Formula (IV) and embodiments thereof to a tRNApyl. In embodiments, the mutant pyrrolysyl-tRNA synthetase catalyzes the attachment of the compound of Formula (VII) and embodiments thereof to a tRNApyl. In embodiments, the pyrrolysyl-tRNA synthetase comprises the amino acid sequence set forth as SEQ ID NO:9.

[0048] The term “mutant pyrrolysyl-tRNA synthetase” or “mutant PylRS” refers to any pyrrolysyl-tRNA synthetase that has a different amino acid sequence from wild-type amino acid sequence.

[0049] The terms “tRNAPyl” and “rTNAPylCUA” and “tRNA Pyl CUA” (i.e., tRNA(superscript Pyl)(subscript CUA)) are used interchangeably and all refer to a single-stranded RNA molecule containing about 70 to 90 nucleotides which fold via intrastrand base pairing to form a characteristic cloverleaf structure that carries a specific amino acid (e.g., compound of Formula (1) or embodiments thereof) and matches it to its corresponding codon (i.e., a complementary to the anticodon of the tRNA) on an mRNA during protein synthesis. In tRNAPyl, the anticodon is CUA. Anticodon CUA is complementary to amber stop codon UAG. In embodiments, the tRNAPylcomprises an anticodon. In embodiments, the anticodon is CUA, TTA, or TCA. In embodiments, the tRNAPylcomprises an anticodon, wherein the anticodon comprises at least one non-cannonical base. The abbreviation “Pyl” of tRNAPylstands for pyrrolysine and the “CUA” of tRNAPylrefers to its anticodon CUA. In embodiments, tRNAPylis attached to the compound of Formula (1) or embodiments thereof.

[0050] The term “substrate-binding site” as used herein refers to residues located in the enzyme active site that form temporary bonds or interactions with the substrate. In embodiments, the substrate-binding site of pyrrolysyl-tRNA synthetase refers to residues located in the active site of pyrrolysyl-tRNA synthetase that form temporary bonds or interactions with the amino acid substrate.

[0051] The term “vector” refers to a nucleic acid molecule capable of transporting another nucleic acid to which it has been linked. One type of vector is a “plasmid”, which refers to a linear or circular double stranded DNA loop into which additional DNA segments can be ligated. Another type of vector is a viral vector, wherein additional DNA segments can beligated into the viral genome. Certain vectors are capable of autonomous replication in a host cell into which they are introduced (e.g., bacterial vectors having a bacterial origin of replication and episomal mammalian vectors). Other vectors (e.g., non episomal mammalian vectors) are integrated into the genome of a host cell upon introduction into the host cell, and thereby are replicated along with the host genome. Moreover, certain vectors are capable of directing the expression of genes to which they are operatively linked. Such vectors are referred to herein as “expression vectors.” In general, expression vectors of utility in recombinant DNA techniques are often in the form of plasmids. The terms “plasmid” and “vector” can be used interchangeably as the plasmid is the most commonly used form of vector. However, the disclosure is intended to include such other forms of expression vectors, such as viral vectors (e.g., replication defective retroviruses, adenoviruses and adeno-associated viruses), which serve equivalent functions. Some viral vectors are capable of targeting a particular cells type either specifically or non- specifically. Exemplary vectors that can be used include, but are not limited to, pEvol vector, pMP vector, pET vector, pTak vector, pBad vector.

[0052] The term “complex” refers to a composition that includes two or more components, where the components bind together to make a functional unit. In embodiments, a complex described herein include a mutant pyrrolysyl-tRNA synthetase described herein and an amino acid substrate (e.g., the compound of Formula (1) or embodiments thereof; the compound of Formula (5) or embodiments thereof). In embodiments, a complex described herein includes a mutant pyrrolysyl-tRNA synthetase described herein and a tRNA (e.g., tRNAPy). In embodiments, a complex described herein includes a mutant pyrrolysyl-tRNA synthetase described herein, an amino acid substrate (e.g., PFY, PFK) and a tRNA (e.g., tRNAPy). In embodiments, a complex described herein includes at least two components selected from the group consisting of a mutant pyrrolysyl-tRNA synthetase described herein, an amino acid substrate (e.g., the compound of Formula (1) or embodiments thereof), a polypeptide containing the compound of Formula (1) or embodiments thereof, and a tRNA (e.g., tRNAPy). In embodiments, a complex described herein includes at least two components selected from the group consisting of a mutant pyrrolysyl-tRNA synthetase described herein, an amino acid substrate (e.g., the compound of Formula (5) or embodiments thereof), a polypeptide containing the compound of Formula (5) or embodiments thereof, and a tRNA (e.g., tRNAPy).

[0053] The term “protein / protein complex” refers to a composition that includes one protein- binding protein (e.g., comprising an unnatural amino acid as described herein) and one protein, where the protein-binding protein and protein are proximal to each other but not bound together; the protein-binding protein and protein are covalently bound together; or the protein-bindingprotein and protein are ionically bound together. In embodiments, the protein-binding protein and protein are proximal to each other but not bound together. In embodiments, the protein- binding protein and protein are covalently bonded together. In embodiments, the protein-binding protein and protein are ionically bonded together. In embodiments, the protein-binding protein and protein are covalently and ionically bonded together. In embodiments, the chemical reaction forming the protein / protein complex is a SuFEx reaction.

[0054] The terms “transfection”, “transduction”, “transfecting” or “transducing” can be used interchangeably and are defined as a process of introducing a nucleic acid molecule or a protein to a cell. Nucleic acids are introduced to a cell using non-viral or viral-based methods. The nucleic acid molecules may be gene sequences encoding complete proteins or functional portions thereof. Non-viral methods of transfection include any appropriate transfection method that does not use viral DNA or viral particles as a delivery system to introduce the nucleic acid molecule into the cell. Exemplary non-viral transfection methods include calcium phosphate transfection, liposomal transfection, nucleofection, sonoporation, transfection through heat shock, magnetifection and electroporation. In embodiments, the nucleic acid molecules are introduced into a cell using electroporation following standard procedures well known in the art. For viral-based methods of transfection any useful viral vector may be used in the methods described herein. Examples for viral vectors include, but are not limited to retroviral, adenoviral, lentiviral and adeno-associated viral vectors. In embodiments, the nucleic acid molecules are introduced into a cell using a retroviral vector following standard procedures well known in the art. The terms ″transfection″ or ″transduction″ also refer to introducing proteins into a cell from the external environment. Typically, transduction or transfection of a protein relies on attachment of a peptide or protein capable of crossing the cell membrane to the protein of interest.

[0055] The term “isolated,” when applied to a nucleic acid or protein, denotes that the nucleic acid or protein is essentially free of other cellular components with which it is associated in the natural state. It can be, for example, in a homogeneous state and may be in either a dry or aqueous solution. Purity and homogeneity are typically determined using analytical chemistry techniques such as polyacrylamide gel electrophoresis or high performance liquid chromatography. A protein that is the predominant species present in a preparation is substantially purified.

[0056] The term “therapeutic agent” refers to any agent useful in treating and / or preventing adisease. “Therapeutic agent“ includes, without limitation, small molecule drugs, proteins,nucleic acids (e.g., DNA, RNA), and the like. “Small-molecule drugs” refers tochemical compounds with low molecular weight that are capable of treating and / or preventing diseases. In embodiments, the proteins described herein are bonded to a therapeutic agent. Methods for covalently bonding therapeutic agents to proteins are well-known in the art.

[0057] The term “intermolecular linker” refers to a linking group between two biomolecules. For example, when the compounds of Formula (4) or (8) (or embodiments thereof) are an intermolecular linker, then the peptidyl moiety of R4is a first protein and the peptidyl moiety of R5is a second (different) protein, such that the first protein and the second protein are covalently bonded. In aspects, the first protein and the second protein can have the same sequence, e.g., providing an intermolecular linker between two different proteins having the same amino acid sequence. In aspects, the first protein and the second protein are different proteins, e.g., providing an intermolecular linker between two different proteins, such as a nanobody and a receptor protein.

[0058] The term “intramolecular linker” refers to a linking group within a biomolecule. For example, when the compounds of Formula (4) or (8) (or embodiments thereof) are an intramolecular linker, then the peptidyl moiety of R4and the peptidyl moiety of R5are in the same protein. A compound having an intramolecular linker may also be referred to as an intramolecularly conjugated protein.

[0059] Where substituent groups are specified by their conventional chemical formulae, written from left to right, they equally encompass the chemically identical substituents that would result from writing the structure from right to left, e.g., -CH2O- is equivalent to -OCH2-.

[0060] The term “alkyl,” by itself or as part of another substituent, means, unless otherwise stated, a straight (i.e., unbranched) or branched carbon chain (or carbon), or combination thereof, which may be fully saturated, mono- or polyunsaturated and can include mono-, di- and multivalent radicals. The alkyl may include a designated number of carbons (e.g., C1-C10means one to ten carbons). Alkyl is an uncyclized chain. Examples of saturated hydrocarbon radicals include, but are not limited to, groups such as methyl, ethyl, n-propyl, isopropyl, n-butyl, t-butyl, isobutyl, sec-butyl, methyl, homologs and isomers of, for example, n-pentyl, n-hexyl, n-heptyl, n-octyl, and the like. An unsaturated alkyl group is one having one or more double bonds or triple bonds. Examples of unsaturated alkyl groups include, but are not limited to, vinyl, 2- propenyl, crotyl, 2-isopentenyl, 2-(butadienyl), 2,4-pentadienyl, 3-(1,4-pentadienyl), ethynyl, 1- and 3-propynyl, 3-butynyl, and the higher homologs and isomers. An alkoxy is an alkyl attached to the remainder of the molecule via an oxygen linker (-O-). An alkyl moiety may be an alkenyl moiety. An alkyl moiety may be an alkynyl moiety. An alkyl moiety may be fully saturated. An alkenyl may include more than one double bond and / or one or more triple bonds in addition tothe one or more double bonds. An alkynyl may include more than one triple bond and / or one or more double bonds in addition to the one or more triple bonds.

[0061] The term “alkylene,” by itself or as part of another substituent, means, unless otherwise stated, a divalent radical derived from an alkyl, as exemplified by, e.g., -CH2CH2CH2CH2-. Typically, an alkyl (or alkylene) group will have from 1 to 24 carbon atoms, with those groups having 10 or fewer carbon atoms being preferred herein. A “lower alkyl” or “lower alkylene” is a shorter chain alkyl or alkylene group, generally having eight or fewer carbon atoms. The term “alkenylene,” by itself or as part of another substituent, means, unless otherwise stated, a divalent radical derived from an alkene.

[0062] The term “heteroalkyl,” by itself or in combination with another term, means, unless otherwise stated, a stable straight or branched chain, or combinations thereof, including at least one carbon atom and at least one heteroatom (e.g., O, N, P, Si, and S), and wherein the nitrogen and sulfur atoms may optionally be oxidized, and the nitrogen heteroatom may optionally be quaternized. The heteroatom(s) may be placed at any interior position of the heteroalkyl group or at the position at which the alkyl group is attached to the remainder of the molecule. Heteroalkyl is an uncyclized chain. Examples include, but are not limited to: -S(O)-CH3, -CH2-CH2-O-CH3, -CH2-CH2-NH-CH3, -CH2-CH2-N(CH3)-CH3, -CH2-S-CH2-CH3, -O-CH3, -CH2-CH2, -CH2-NH2, -CH2-NO2, -NHCH3, -CH2-CH2-S(O)2-CH3, -CH=CH-O-CH3, -Si(CH3)3, -CH2-CH=N-OCH3, -CH=CH-N(CH3)-CH3, -O-CH2-CH3, and -CN. Up to two or three heteroatoms may be consecutive, such as, for example, -CH2-NH-OCH3and -CH2-O-Si(CH3)3. A heteroalkyl moiety may include one heteroatom. A heteroalkyl moiety may include two optionally different heteroatoms. A heteroalkyl moiety may include three optionally different heteroatoms. A heteroalkyl moiety may include four optionally different heteroatoms. A heteroalkyl moiety may include five optionally different heteroatoms. A heteroalkyl moiety may include up to 8 optionally different heteroatoms. The term “heteroalkenyl,” by itself or in combination with another term, means, unless otherwise stated, a heteroalkyl including at least one double bond. A heteroalkenyl may optionally include more than one double bond and / or one or more triple bonds in additional to the one or more double bonds. The term “heteroalkynyl,” by itself or in combination with another term, means, unless otherwise stated, a heteroalkyl including at least one triple bond. A heteroalkynyl may optionally include more than one triple bond and / or one or more double bonds in additional to the one or more triple bonds.

[0063] Similarly, the term “heteroalkylene,” by itself or as part of another substituent, means, unless otherwise stated, a divalent radical derived from heteroalkyl, as exemplified, but not limited by, -CH2-CH2-S-CH2-CH2- and -CH2-S-CH2-CH2-NH-CH2-. For heteroalkylene groups,heteroatoms can also occupy either or both of the chain termini (e.g., alkyleneoxy, alkylenedioxy, alkyleneamino, alkylenediamino, and the like). Still further, for alkylene and heteroalkylene linking groups, no orientation of the linking group is implied by the direction in which the formula of the linking group is written. For example, the formula -C(O)2R'- represents both -C(O)2R'- and -R'C(O)2-. As described above, heteroalkyl groups, as used herein, include those groups that are attached to the remainder of the molecule through a heteroatom, such as - C(O)R', -C(O)NR', -NR'R'', -OR', -SR', and / or -SO2R'. Where “heteroalkyl” is recited, followed by recitations of specific heteroalkyl groups, such as -NR'R'' or the like, it will be understood that the terms heteroalkyl and -NR'R'' are not redundant or mutually exclusive. Rather, the specific heteroalkyl groups are recited to add clarity. Thus, the term “heteroalkyl” should not be interpreted herein as excluding specific heteroalkyl groups, such as -NR'R'' or the like.

[0064] The terms “cycloalkyl” and “heterocycloalkyl,” by themselves or in combination with other terms, mean, unless otherwise stated, cyclic versions of “alkyl” and “heteroalkyl,” respectively. Cycloalkyl and heterocycloalkyl are not aromatic. Additionally, for heterocycloalkyl, a heteroatom can occupy the position at which the heterocycle is attached to the remainder of the molecule. Examples of cycloalkyl include, but are not limited to, cyclopropyl, cyclobutyl, cyclopentyl, cyclohexyl, 1-cyclohexenyl, 3-cyclohexenyl, cycloheptyl, and the like. Examples of heterocycloalkyl include, but are not limited to, 1-(1,2,5,6- tetrahydropyridyl), 1-piperidinyl, 2-piperidinyl, 3-piperidinyl, 4-morpholinyl, 3-morpholinyl, tetrahydrofuran-2-yl, tetrahydrofuran-3-yl, tetrahydrothien-2-yl, tetrahydrothien-3-yl, 1- piperazinyl, 2-piperazinyl, and the like. A “cycloalkylene” and a “heterocycloalkylene,” alone or as part of another substituent, means a divalent radical derived from a cycloalkyl and heterocycloalkyl, respectively.

[0065] In embodiments, the term “cycloalkyl” means a monocyclic, bicyclic, or a multicyclic cycloalkyl ring system. In embodiments, monocyclic ring systems are cyclic hydrocarbon groups containing from 3 to 8 carbon atoms, where such groups can be saturated or unsaturated, but not aromatic. In embodiments, cycloalkyl groups are fully saturated. Examples of monocyclic cycloalkyls include cyclopropyl, cyclobutyl, cyclopentyl, cyclopentenyl, cyclohexyl, cyclohexenyl, cycloheptyl, and cyclooctyl. Bicyclic cycloalkyl ring systems are bridged monocyclic rings or fused bicyclic rings. In embodiments, bridged monocyclic rings contain a monocyclic cycloalkyl ring where two non adjacent carbon atoms of the monocyclic ring are linked by an alkylene bridge of between one and three additional carbon atoms (i.e., a bridging group of the form (CH2)w , where w is 1, 2, or 3). Representative examples of bicyclic ring systems include, but are not limited to, bicyclo[3.1.1]heptane, bicyclo[2.2.1]heptane,bicyclo[2.2.2]octane, bicyclo[3.2.2]nonane, bicyclo[3.3.1]nonane, and bicyclo[4.2.1]nonane. In embodiments, fused bicyclic cycloalkyl ring systems contain a monocyclic cycloalkyl ring fused to either a phenyl, a monocyclic cycloalkyl, a monocyclic cycloalkenyl, a monocyclic heterocyclyl, or a monocyclic heteroaryl. In embodiments, the bridged or fused bicyclic cycloalkyl is attached to the parent molecular moiety through any carbon atom contained within the monocyclic cycloalkyl ring. In embodiments, cycloalkyl groups are optionally substituted with one or two groups which are independently oxo or thia. In embodiments, the fused bicyclic cycloalkyl is a 5 or 6 membered monocyclic cycloalkyl ring fused to either a phenyl ring, a 5 or 6 membered monocyclic cycloalkyl, a 5 or 6 membered monocyclic cycloalkenyl, a 5 or 6 membered monocyclic heterocyclyl, or a 5 or 6 membered monocyclic heteroaryl, wherein the fused bicyclic cycloalkyl is optionally substituted by one or two groups which are independently oxo or thia. In embodiments, multicyclic cycloalkyl ring systems are a monocyclic cycloalkyl ring (base ring) fused to either (i) one ring system selected from the group consisting of a bicyclic aryl, a bicyclic heteroaryl, a bicyclic cycloalkyl, a bicyclic cycloalkenyl, and a bicyclic heterocyclyl; or (ii) two other ring systems independently selected from the group consisting of a phenyl, a bicyclic aryl, a monocyclic or bicyclic heteroaryl, a monocyclic or bicyclic cycloalkyl, a monocyclic or bicyclic cycloalkenyl, and a monocyclic or bicyclic heterocyclyl. In embodiments, the multicyclic cycloalkyl is attached to the parent molecular moiety through any carbon atom contained within the base ring. In embodiments, multicyclic cycloalkyl ring systems are a monocyclic cycloalkyl ring (base ring) fused to either (i) one ring system selected from the group consisting of a bicyclic aryl, a bicyclic heteroaryl, a bicyclic cycloalkyl, a bicyclic cycloalkenyl, and a bicyclic heterocyclyl; or (ii) two other ring systems independently selected from the group consisting of a phenyl, a monocyclic heteroaryl, a monocyclic cycloalkyl, a monocyclic cycloalkenyl, and a monocyclic heterocyclyl. Examples of multicyclic cycloalkyl groups include, but are not limited to tetradecahydrophenanthrenyl, perhydrophenothiazin-1-yl, and perhydrophenoxazin-1-yl.

[0066] In embodiments, a cycloalkyl is a cycloalkenyl. The term “cycloalkenyl” is used in accordance with its plain ordinary meaning. In embodiments, a cycloalkenyl is a monocyclic, bicyclic, or a multicyclic cycloalkenyl ring system. In embodiments, monocyclic cycloalkenyl ring systems are cyclic hydrocarbon groups containing from 3 to 8 carbon atoms, where such groups are unsaturated (i.e., containing at least one annular carbon carbon double bond), but not aromatic. Examples of monocyclic cycloalkenyl ring systems include cyclopentenyl and cyclohexenyl. In embodiments, bicyclic cycloalkenyl rings are bridged monocyclic rings or a fused bicyclic rings. In embodiments, bridged monocyclic rings contain a monocycliccycloalkenyl ring where two non adjacent carbon atoms of the monocyclic ring are linked by an alkylene bridge of between one and three additional carbon atoms (i.e., a bridging group of the form (CH2)w, where w is 1, 2, or 3). Representative examples of bicyclic cycloalkenyls include, but are not limited to, norbornenyl and bicyclo[2.2.2]oct 2 enyl. In embodiments, fused bicyclic cycloalkenyl ring systems contain a monocyclic cycloalkenyl ring fused to either a phenyl, a monocyclic cycloalkyl, a monocyclic cycloalkenyl, a monocyclic heterocyclyl, or a monocyclic heteroaryl. In embodiments, the bridged or fused bicyclic cycloalkenyl is attached to the parent molecular moiety through any carbon atom contained within the monocyclic cycloalkenyl ring. In embodiments, cycloalkenyl groups are optionally substituted with one or two groups which are independently oxo or thia. In embodiments, multicyclic cycloalkenyl rings contain a monocyclic cycloalkenyl ring (base ring) fused to either (i) one ring system selected from the group consisting of a bicyclic aryl, a bicyclic heteroaryl, a bicyclic cycloalkyl, a bicyclic cycloalkenyl, and a bicyclic heterocyclyl; or (ii) two ring systems independently selected from the group consisting of a phenyl, a bicyclic aryl, a monocyclic or bicyclic heteroaryl, a monocyclic or bicyclic cycloalkyl, a monocyclic or bicyclic cycloalkenyl, and a monocyclic or bicyclic heterocyclyl. In embodiments, the multicyclic cycloalkenyl is attached to the parent molecular moiety through any carbon atom contained within the base ring. In embodiments, multicyclic cycloalkenyl rings contain a monocyclic cycloalkenyl ring (base ring) fused to either (i) one ring system selected from the group consisting of a bicyclic aryl, a bicyclic heteroaryl, a bicyclic cycloalkyl, a bicyclic cycloalkenyl, and a bicyclic heterocyclyl; or (ii) two ring systems independently selected from the group consisting of a phenyl, a monocyclic heteroaryl, a monocyclic cycloalkyl, a monocyclic cycloalkenyl, and a monocyclic heterocyclyl.

[0067] In embodiments, a heterocycloalkyl is a heterocyclyl. The term “heterocyclyl” as used herein, means a monocyclic, bicyclic, or multicyclic heterocycle. The heterocyclyl monocyclic heterocycle is a 3, 4, 5, 6 or 7 membered ring containing at least one heteroatom independently selected from the group consisting of O, N, and S where the ring is saturated or unsaturated, but not aromatic. The 3 or 4 membered ring contains 1 heteroatom selected from the group consisting of O, N and S. The 5 membered ring can contain zero or one double bond and one, two or three heteroatoms selected from the group consisting of O, N and S. The 6 or 7 membered ring contains zero, one or two double bonds and one, two or three heteroatoms selected from the group consisting of O, N and S. The heterocyclyl monocyclic heterocycle is connected to the parent molecular moiety through any carbon atom or any nitrogen atom contained within the heterocyclyl monocyclic heterocycle. Representative examples of heterocyclyl monocyclic heterocycles include, but are not limited to, azetidinyl, azepanyl,aziridinyl, diazepanyl, 1,3-dioxanyl, 1,3-dioxolanyl, 1,3-dithiolanyl, 1,3-dithianyl, imidazolinyl, imidazolidinyl, isothiazolinyl, isothiazolidinyl, isoxazolinyl, isoxazolidinyl, morpholinyl, oxadiazolinyl, oxadiazolidinyl, oxazolinyl, oxazolidinyl, piperazinyl, piperidinyl, pyranyl, pyrazolinyl, pyrazolidinyl, pyrrolinyl, pyrrolidinyl, tetrahydrofuranyl, tetrahydrothienyl, thiadiazolinyl, thiadiazolidinyl, thiazolinyl, thiazolidinyl, thiomorpholinyl, 1,1- dioxidothiomorpholinyl (thiomorpholine sulfone), thiopyranyl, and trithianyl. The heterocyclyl bicyclic heterocycle is a monocyclic heterocycle fused to either a phenyl, a monocyclic cycloalkyl, a monocyclic cycloalkenyl, a monocyclic heterocycle, or a monocyclic heteroaryl. The heterocyclyl bicyclic heterocycle is connected to the parent molecular moiety through any carbon atom or any nitrogen atom contained within the monocyclic heterocycle portion of the bicyclic ring system. Representative examples of bicyclic heterocyclyls include, but are not limited to, 2,3-dihydrobenzofuran-2-yl, 2,3-dihydrobenzofuran-3-yl, indolin-1-yl, indolin-2-yl, indolin-3-yl, 2,3-dihydrobenzothien-2-yl, decahydroquinolinyl, decahydroisoquinolinyl, octahydro-1H-indolyl, and octahydrobenzofuranyl. In embodiments, heterocyclyl groups are optionally substituted with one or two groups which are independently oxo or thia. In certain embodiments, the bicyclic heterocyclyl is a 5 or 6 membered monocyclic heterocyclyl ring fused to a phenyl ring, a 5 or 6 membered monocyclic cycloalkyl, a 5 or 6 membered monocyclic cycloalkenyl, a 5 or 6 membered monocyclic heterocyclyl, or a 5 or 6 membered monocyclic heteroaryl, wherein the bicyclic heterocyclyl is optionally substituted by one or two groups which are independently oxo or thia. Multicyclic heterocyclyl ring systems are a monocyclic heterocyclyl ring (base ring) fused to either (i) one ring system selected from the group consisting of a bicyclic aryl, a bicyclic heteroaryl, a bicyclic cycloalkyl, a bicyclic cycloalkenyl, and a bicyclic heterocyclyl; or (ii) two other ring systems independently selected from the group consisting of a phenyl, a bicyclic aryl, a monocyclic or bicyclic heteroaryl, a monocyclic or bicyclic cycloalkyl, a monocyclic or bicyclic cycloalkenyl, and a monocyclic or bicyclic heterocyclyl. The multicyclic heterocyclyl is attached to the parent molecular moiety through any carbon atom or nitrogen atom contained within the base ring. In embodiments, multicyclic heterocyclyl ring systems are a monocyclic heterocyclyl ring (base ring) fused to either (i) one ring system selected from the group consisting of a bicyclic aryl, a bicyclic heteroaryl, a bicyclic cycloalkyl, a bicyclic cycloalkenyl, and a bicyclic heterocyclyl; or (ii) two other ring systems independently selected from the group consisting of a phenyl, a monocyclic heteroaryl, a monocyclic cycloalkyl, a monocyclic cycloalkenyl, and a monocyclic heterocyclyl. Examples of multicyclic heterocyclyl groups include, but are not limited to 10H-phenothiazin-10-yl, 9,10- dihydroacridin-9-yl, 9,10-dihydroacridin-10-yl, 10H-phenoxazin-10-yl, 10,11-dihydro-5H-dibenzo[b,f]azepin-5-yl, 1,2,3,4-tetrahydropyrido[4,3-g]isoquinolin-2-yl, 12H- benzo[b]phenoxazin-12-yl, and dodecahydro-1H-carbazol-9-yl.

[0068] The terms “halo” or “halogen,” by themselves or as part of another substituent, mean, unless otherwise stated, a fluorine, chlorine, bromine, or iodine atom. Additionally, terms such as “haloalkyl” are meant to include monohaloalkyl and polyhaloalkyl. For example, the term “halo(C1-C4)alkyl” includes, but is not limited to, fluoromethyl, difluoromethyl, trifluoromethyl, 2,2,2-trifluoroethyl, 4-chlorobutyl, 3-bromopropyl, and the like.

[0069] The term “acyl” means, unless otherwise stated, -C(O)R where R is a substituted or unsubstituted alkyl, substituted or unsubstituted cycloalkyl, substituted or unsubstituted heteroalkyl, substituted or unsubstituted heterocycloalkyl, substituted or unsubstituted aryl, or substituted or unsubstituted heteroaryl.

[0070] The term “aryl” means, unless otherwise stated, a polyunsaturated, aromatic, hydrocarbon substituent, which can be a single ring or multiple rings (preferably from 1 to 3 rings) that are fused together (i.e., a fused ring aryl) or linked covalently. A fused ring aryl refers to multiple rings fused together wherein at least one of the fused rings is an aryl ring. The term “heteroaryl” refers to aryl groups (or rings) that contain at least one heteroatom such as N, O, or S, wherein the nitrogen and sulfur atoms are optionally oxidized, and the nitrogen atom(s) are optionally quaternized. Thus, the term “heteroaryl” includes fused ring heteroaryl groups (i.e., multiple rings fused together wherein at least one of the fused rings is a heteroaromatic ring). A 5,6-fused ring heteroarylene refers to two rings fused together, wherein one ring has 5 members and the other ring has 6 members, and wherein at least one ring is a heteroaryl ring. Likewise, a 6,6-fused ring heteroarylene refers to two rings fused together, wherein one ring has 6 members and the other ring has 6 members, and wherein at least one ring is a heteroaryl ring. And a 6,5- fused ring heteroarylene refers to two rings fused together, wherein one ring has 6 members and the other ring has 5 members, and wherein at least one ring is a heteroaryl ring. A heteroaryl group can be attached to the remainder of the molecule through a carbon or heteroatom. Non- limiting examples of aryl and heteroaryl groups include phenyl, naphthyl, pyrrolyl, pyrazolyl, pyridazinyl, triazinyl, pyrimidinyl, imidazolyl, pyrazinyl, purinyl, oxazolyl, isoxazolyl, thiazolyl, furyl, thienyl, pyridyl, pyrimidyl, benzothiazolyl, benzoxazoyl benzimidazolyl, benzofuran, isobenzofuranyl, indolyl, isoindolyl, benzothiophenyl, isoquinolyl, quinoxalinyl, quinolyl, 1-naphthyl, 2-naphthyl, 4-biphenyl, 1-pyrrolyl, 2-pyrrolyl, 3-pyrrolyl, 3-pyrazolyl, 2- imidazolyl, 4-imidazolyl, pyrazinyl, 2-oxazolyl, 4-oxazolyl, 2-phenyl-4-oxazolyl, 5-oxazolyl, 3- isoxazolyl, 4-isoxazolyl, 5-isoxazolyl, 2-thiazolyl, 4-thiazolyl, 5-thiazolyl, 2-furyl, 3-furyl, 2- thienyl, 3-thienyl, 2-pyridyl, 3-pyridyl, 4-pyridyl, 2-pyrimidyl, 4-pyrimidyl, 5-benzothiazolyl,purinyl, 2-benzimidazolyl, 5-indolyl, 1-isoquinolyl, 5-isoquinolyl, 2-quinoxalinyl, 5- quinoxalinyl, 3-quinolyl, and 6-quinolyl. Substituents for each of the above noted aryl and heteroaryl ring systems are selected from the group of acceptable substituents described below. An “arylene” and a “heteroarylene,” alone or as part of another substituent, mean a divalent radical derived from an aryl and heteroaryl, respectively. A heteroaryl group substituent may be -O- bonded to a ring heteroatom nitrogen.

[0071] A fused ring heterocyloalkyl-aryl is an aryl fused to a heterocycloalkyl. A fused ring heterocycloalkyl-heteroaryl is a heteroaryl fused to a heterocycloalkyl. A fused ring heterocycloalkyl-cycloalkyl is a heterocycloalkyl fused to a cycloalkyl. A fused ring heterocycloalkyl-heterocycloalkyl is a heterocycloalkyl fused to another heterocycloalkyl. Fused ring heterocycloalkyl-aryl, fused ring heterocycloalkyl-heteroaryl, fused ring heterocycloalkyl- cycloalkyl, or fused ring heterocycloalkyl-heterocycloalkyl may each independently be unsubstituted or substituted with one or more of the substituents described herein.

[0072] Spirocyclic rings are two or more rings wherein adjacent rings are attached through a single atom. The individual rings within spirocyclic rings may be identical or different. Individual rings in spirocyclic rings may be substituted or unsubstituted and may have different substituents from other individual rings within a set of spirocyclic rings. Possible substituents for individual rings within spirocyclic rings are the possible substituents for the same ring when not part of spirocyclic rings (e.g. substituents for cycloalkyl or heterocycloalkyl rings). Spirocyclic rings may be substituted or unsubstituted cycloalkyl, substituted or unsubstituted cycloalkylene, substituted or unsubstituted heterocycloalkyl or substituted or unsubstituted heterocycloalkylene and individual rings within a spirocyclic ring group may be any of the immediately previous list, including having all rings of one type (e.g. all rings being substituted heterocycloalkylene wherein each ring may be the same or different substituted heterocycloalkylene). When referring to a spirocyclic ring system, heterocyclic spirocyclic rings means a spirocyclic rings wherein at least one ring is a heterocyclic ring and wherein each ring may be a different ring. When referring to a spirocyclic ring system, substituted spirocyclic rings means that at least one ring is substituted and each substituent may optionally be different.

[0073] The symbol “ ” or “-” denotes the point of attachment of a chemical moiety to theremainder of a molecule or chemical formula.

[0074] The term “oxo,” as used herein, means an oxygen that is double bonded to a carbon atom.

[0075] The term “alkylsulfonyl,” as used herein, means a moiety having the formula -S(O2)-R', where R' is a substituted or unsubstituted alkyl group as defined above. R'may have a specified number of carbons (e.g., “C1-C4alkylsulfonyl”).

[0076] The term “alkylarylene” as an arylene moiety covalently bonded to an alkylene moiety (also referred to herein as an alkylene linker).

[0077] An alkylarylene moiety may be substituted (e.g. with a substituent group) on the alkylene moiety or the arylene linker (e.g. at carbons 2, 3, 4, or 6) with halogen, oxo, -N3, -CF3, -CCl3, -CBr3, -CI3, -CN, -CHO, -OH, -NH2, -COOH, -CONH2, -NO2, -SH, -SO2CH3, -SO3H, -OSO3H, -SO2NH2, ^NHNH2, ^ONH2, ^NHC(O)NHNH2, substituted or unsubstituted C1-C5 alkyl or substituted or unsubstituted 2 to 5 membered heteroalkyl). In embodiments, the alkylarylene is unsubstituted.

[0078] Each of the above terms (e.g., “alkyl,” “heteroalkyl,” “cycloalkyl,” “heterocycloalkyl,” “aryl,” and “heteroaryl”) includes both substituted and unsubstituted forms of the indicated radical. Preferred substituents for each type of radical are provided below.

[0079] Substituents for the alkyl and heteroalkyl radicals (including those groups often referred to as alkylene, alkenyl, heteroalkylene, heteroalkenyl, alkynyl, cycloalkyl, heterocycloalkyl, cycloalkenyl, and heterocycloalkenyl) can be one or more of a variety of groups selected from, but not limited to, -OR', =O, =NR', =N-OR', -NR'R'', -SR', -halogen, -SiR'R''R''', -OC(O)R', -C(O)R', -CO2R', -CONR'R'', -OC(O)NR'R'', -NR''C(O)R', -NR'-C(O)NR''R''', -NR''C(O)2R', -NR-C(NR'R''R''')=NR'''', -NR-C(NR'R'')=NR''', -S(O)R', -S(O)2R', -S(O)2NR'R'', -NRSO2R', ^NR'NR''R''', ^ONR'R'', ^NR'C(O)NR''NR'''R'''', -CN, -NO2, -NR'SO2R'', -NR'C(O)R'', -NR'C(O)-OR'', -NR'OR'', in a number ranging from zero to (2m'+1), where m' is the total number of carbon atoms in such radical. R, R', R'', R''', and R'''' each preferably independently refer to hydrogen, substituted or unsubstituted heteroalkyl, substituted or unsubstituted cycloalkyl, substituted or unsubstituted heterocycloalkyl, substituted or unsubstituted aryl (e.g., aryl substituted with 1-3 halogens), substituted or unsubstituted heteroaryl, substituted or unsubstituted alkyl, alkoxy, or thioalkoxy groups, or arylalkyl groups. When a compound described herein includes more than one R group, for example, each of the R groups is independently selected as are each R', R'', R''', and R'''' group when more than one of these groups is present. When R' and R'' are attached to the same nitrogen atom, they can be combined with the nitrogen atom to form a 4-, 5-, 6-, or 7-membered ring. For example, -NR'R'' includes, but is not limited to, 1-pyrrolidinyl and 4-morpholinyl. From the above discussion of substituents, one of skill in the art will understand that the term “alkyl” is meant to include groups including carbon atoms bound to groups other than hydrogen groups, such as haloalkyl (e.g., -CF3 and -CH2CF3) and acyl (e.g., -C(O)CH3, -C(O)CF3, -C(O)CH2OCH3, and the like).

[0080] Similar to the substituents described for the alkyl radical, substituents for the aryl andheteroaryl groups are varied and are selected from, for example: -OR', -NR'R'', -SR', -halogen, -SiR'R''R''', -OC(O)R', -C(O)R', -CO2R', -CONR'R'', -OC(O)NR'R'', -NR''C(O)R', -NR'-C(O)NR''R''', -NR''C(O)2R', -NR-C(NR'R''R''')=NR'''', -NR-C(NR'R'')=NR''', -S(O)R', -S(O)2R', -S(O)2NR'R'', -NRSO2R', ^NR'NR''R''', ^ONR'R'', ^NR'C(O)NR''NR'''R'''', -CN, -NO2, -R', -N3, -CH(Ph)2, fluoro(C1-C4)alkoxy, and fluoro(C1-C4)alkyl, -NR'SO2R'', -NR'C(O)R'', -NR'C(O)-OR'', -NR'OR'', in a number ranging from zero to the total number of open valences on the aromatic ring system; and where R', R'', R''', and R'''' are preferably independently selected from hydrogen, substituted or unsubstituted alkyl, substituted or unsubstituted heteroalkyl, substituted or unsubstituted cycloalkyl, substituted or unsubstituted heterocycloalkyl, substituted or unsubstituted aryl, and substituted or unsubstituted heteroaryl. When a compound described herein includes more than one R group, for example, each of the R groups is independently selected as are each R', R'', R''', and R'''' groups when more than one of these groups is present.

[0081] Substituents for rings (e.g. cycloalkyl, heterocycloalkyl, aryl, heteroaryl, cycloalkylene, heterocycloalkylene, arylene, or heteroarylene) may be depicted as substituents on the ring rather than on a specific atom of a ring (commonly referred to as a floating substituent). In such a case, the substituent may be attached to any of the ring atoms (obeying the rules of chemical valency) and in the case of fused rings or spirocyclic rings, a substituent depicted as associated with one member of the fused rings or spirocyclic rings (a floating substituent on a single ring), may be a substituent on any of the fused rings or spirocyclic rings (a floating substituent on multiple rings). When a substituent is attached to a ring, but not a specific atom (a floating substituent), and a subscript for the substituent is an integer greater than one, the multiple substituents may be on the same atom, same ring, different atoms, different fused rings, different spirocyclic rings, and each substituent may optionally be different. Where a point of attachment of a ring to the remainder of a molecule is not limited to a single atom (a floating substituent), the attachment point may be any atom of the ring and in the case of a fused ring or spirocyclic ring, any atom of any of the fused rings or spirocyclic rings while obeying the rules of chemical valency. Where a ring, fused rings, or spirocyclic rings contain one or more ring heteroatoms and the ring, fused rings, or spirocyclic rings are shown with one more floating substituents (including, but not limited to, points of attachment to the remainder of the molecule), the floating substituents may be bonded to the heteroatoms. Where the ring heteroatoms are shown bound to one or more hydrogens (e.g. a ring nitrogen with two bonds to ring atoms and a third bond to a hydrogen) in the structure or formula with the floating substituent, when the heteroatom is bonded to the floating substituent, the substituent will beunderstood to replace the hydrogen, while obeying the rules of chemical valency.

[0082] Two or more substituents may optionally be joined to form aryl, heteroaryl, cycloalkyl, or heterocycloalkyl groups. Such so-called ring-forming substituents are typically, though not necessarily, found attached to a cyclic base structure. In embodiments, the ring-forming substituents are attached to adjacent members of the base structure. For example, two ring- forming substituents attached to adjacent members of a cyclic base structure create a fused ring structure. In embodiments, the ring-forming substituents are attached to a single member of the base structure. For example, two ring-forming substituents attached to a single member of a cyclic base structure create a spirocyclic structure. In embodiments, the ring-forming substituents are attached to non-adjacent members of the base structure.

[0083] Two of the substituents on adjacent atoms of the aryl or heteroaryl ring may optionally form a ring of the formula -T-C(O)-(CRR')q-U-, wherein T and U are independently -NR-, -O-, -CRR'-, or a single bond, and q is an integer of from 0 to 3. Alternatively, two of the substituents on adjacent atoms of the aryl or heteroaryl ring may optionally be replaced with a substituent of the formula -A-(CH2)r-B-, wherein A and B are independently -CRR'-, -O-, -NR-, -S-, -S(O) -, -S(O)2-, -S(O)2NR'-, or a single bond, and r is an integer of from 1 to 4. One of the single bonds of the new ring so formed may optionally be replaced with a double bond. Alternatively, two of the substituents on adjacent atoms of the aryl or heteroaryl ring may optionally be replaced with a substituent of the formula -(CRR')s-X'- (C''R''R''')d-, where s and d are independently integers of from 0 to 3, and X' is -O-, -NR'-, -S-, -S(O)-, -S(O)2-, or -S(O)2NR'-. The substituents R, R', R'', and R''' are preferably independently selected from hydrogen, substituted or unsubstituted alkyl, substituted or unsubstituted heteroalkyl, substituted or unsubstituted cycloalkyl, substituted or unsubstituted heterocycloalkyl, substituted or unsubstituted aryl, and substituted or unsubstituted heteroaryl.

[0084] As used herein, the terms “heteroatom” or “ring heteroatom” are meant to include oxygen (O), nitrogen (N), sulfur (S), phosphorus (P), and silicon (Si).

[0085] A “substituent group,” as used herein, means a group selected from the following moieties:

[0086] (A) oxo, halogen, -CCl3, -CBr3, -CF3, -CI3,-CN, -OH, -NH2, -COOH, -CONH2, -NO2, -SH, -SO3H, -SO4H, -SO2NH2, ^NHNH2, ^ONH2, ^NHC(O)NHNH2, -NHC(O)NH2, -NHSO2H, -NHC(O)H, -NHC(O)OH, -NHOH, -OCCl3, -OCF3, -OCBr3, -OCI3,-OCHCl2, -OCHBr2, -OCHI2, -OCHF2, unsubstituted alkyl (e.g., C1-C8alkyl, C1-C6alkyl, or C1-C4alkyl), unsubstituted heteroalkyl (e.g., 2 to 8 membered heteroalkyl, 2 to 6 membered heteroalkyl, or 2 to 4 membered heteroalkyl), unsubstituted cycloalkyl (e.g., C3-C8cycloalkyl, C3-C6cycloalkyl,or C5-C6cycloalkyl), unsubstituted heterocycloalkyl (e.g., 3 to 8 membered heterocycloalkyl, 3 to 6 membered heterocycloalkyl, or 5 to 6 membered heterocycloalkyl), unsubstituted aryl (e.g., C6-C10aryl, C10aryl, or phenyl), or unsubstituted heteroaryl (e.g., 5 to 10 membered heteroaryl, 5 to 9 membered heteroaryl, or 5 to 6 membered heteroaryl), and

[0087] (B) alkyl, heteroalkyl, cycloalkyl, heterocycloalkyl, aryl, heteroaryl, substituted with at least one substituent selected from:

[0088] (i) oxo, halogen, -CCl3, -CBr3, -CF3, -CI3,-CN, -OH, -NH2, -COOH, -CONH2, -NO2, -SH, -SO3H, -SO4H, -SO2NH2, ^NHNH2, ^ONH2, ^NHC(O)NHNH2, -NHOH, -OCCl3, -OCF3, -NHC(O)NH2, -NHSO2H, -NHC(O)H, -NHC(O)OH, -OCBr3, -OCI3, -OCHCl2, -OCHBr2, -OCHI2, -OCHF2, unsubstituted alkyl (e.g., C1-C8 alkyl, C1-C6 alkyl, or C1-C4 alkyl), unsubstituted heteroalkyl (e.g., 2 to 8 membered heteroalkyl, 2 to 6 membered heteroalkyl, or 2 to 4 membered heteroalkyl), unsubstituted cycloalkyl (e.g., C3-C8 cycloalkyl, C3-C6 cycloalkyl, or C5-C6cycloalkyl), unsubstituted heterocycloalkyl (e.g., 3 to 8 membered heterocycloalkyl, 3 to 6 membered heterocycloalkyl, or 5 to 6 membered heterocycloalkyl), unsubstituted aryl (e.g., C6-C10aryl, C10aryl, or phenyl), or unsubstituted heteroaryl (e.g., 5 to 10 membered heteroaryl, 5 to 9 membered heteroaryl, or 5 to 6 membered heteroaryl), and

[0089] (ii) alkyl, heteroalkyl, cycloalkyl, heterocycloalkyl, aryl, heteroaryl, substituted with at least one substituent selected from:

[0090] (a) oxo, halogen, -CCl3, -CBr3, -CF3, -CI3,-CN, -OH, -NH2, -COOH, -CONH2, -NO2, -SH, -SO3H, -SO4H, -SO2NH2, ^NHNH2, ^ONH2, ^NHC(O)NHNH2, -NHC(O)NH2, -NHSO2H, -NHC(O)H, -NHC(O)OH, -NHOH, -OCCl3, -OCF3, -OCBr3, -OCI3, -OCHCl2, -OCHBr2, -OCHI2, -OCHF2, unsubstituted alkyl (e.g., C1-C8 alkyl, C1-C6 alkyl, or C1-C4 alkyl), unsubstituted heteroalkyl (e.g., 2 to 8 membered heteroalkyl, 2 to 6 membered heteroalkyl, or 2 to 4 membered heteroalkyl), unsubstituted cycloalkyl (e.g., C3-C8 cycloalkyl, C3-C6 cycloalkyl, or C5-C6 cycloalkyl), unsubstituted heterocycloalkyl (e.g., 3 to 8 membered heterocycloalkyl, 3 to 6 membered heterocycloalkyl, or 5 to 6 membered heterocycloalkyl), unsubstituted aryl (e.g., C6-C10 aryl, C10 aryl, or phenyl), or unsubstituted heteroaryl (e.g., 5 to 10 membered heteroaryl, 5 to 9 membered heteroaryl, or 5 to 6 membered heteroaryl), and

[0091] (b) alkyl, heteroalkyl, cycloalkyl, heterocycloalkyl, aryl, heteroaryl, substituted with at least one substituent selected from: oxo, halogen, -CCl3, -CBr3, -CF3, -CI3,-CN, -OH, -NH2, -COOH, -CONH2, -NO2, -SH, -SO3H, -SO4H, -SO2NH2, ^NHNH2, ^ONH2, ^NHC(O)NHNH2, -NHC(O)NH2, -NHSO2H, -NHC(O)H, -NHC(O)OH, -NHOH, -OCCl3, -OCF3, -OCBr3, -OCI3, -OCHCl2, -OCHBr2, -OCHI2, -OCHF2, unsubstituted alkyl (e.g., C1-C8alkyl, C1-C6alkyl, or C1- C4 alkyl), unsubstituted heteroalkyl (e.g., 2 to 8 membered heteroalkyl, 2 to 6 memberedheteroalkyl, or 2 to 4 membered heteroalkyl), unsubstituted cycloalkyl (e.g., C3-C8cycloalkyl, C3-C6 cycloalkyl, or C5-C6 cycloalkyl), unsubstituted heterocycloalkyl (e.g., 3 to 8 membered heterocycloalkyl, 3 to 6 membered heterocycloalkyl, or 5 to 6 membered heterocycloalkyl), unsubstituted aryl (e.g., C6-C10 aryl, C10 aryl, or phenyl), or unsubstituted heteroaryl (e.g., 5 to 10 membered heteroaryl, 5 to 9 membered heteroaryl, or 5 to 6 membered heteroaryl).

[0092] A “size-limited substituent” or “ size-limited substituent group,” as used herein, means a group selected from all of the substituents described above for a “substituent group,” wherein each substituted or unsubstituted alkyl is a substituted or unsubstituted C1-C20 alkyl, each substituted or unsubstituted heteroalkyl is a substituted or unsubstituted 2 to 20 membered heteroalkyl, each substituted or unsubstituted cycloalkyl is a substituted or unsubstituted C3-C8 cycloalkyl, each substituted or unsubstituted heterocycloalkyl is a substituted or unsubstituted 3 to 8 membered heterocycloalkyl, each substituted or unsubstituted aryl is a substituted or unsubstituted C6-C10 aryl, and each substituted or unsubstituted heteroaryl is a substituted or unsubstituted 5 to 10 membered heteroaryl.

[0093] A “lower substituent” or “ lower substituent group,” as used herein, means a group selected from all of the substituents described above for a “substituent group,” wherein each substituted or unsubstituted alkyl is a substituted or unsubstituted C1-C8 alkyl, each substituted or unsubstituted heteroalkyl is a substituted or unsubstituted 2 to 8 membered heteroalkyl, each substituted or unsubstituted cycloalkyl is a substituted or unsubstituted C3-C7 cycloalkyl, each substituted or unsubstituted heterocycloalkyl is a substituted or unsubstituted 3 to 7 membered heterocycloalkyl, each substituted or unsubstituted aryl is a substituted or unsubstituted C6-C10 aryl, and each substituted or unsubstituted heteroaryl is a substituted or unsubstituted 5 to 9 membered heteroaryl.

[0094] In embodiments, each substituted group described in the compounds herein is substituted with at least one substituent group. More specifically, in embodiments, each substituted alkyl, substituted heteroalkyl, substituted cycloalkyl, substituted heterocycloalkyl, substituted aryl, substituted heteroaryl, substituted alkylene, substituted heteroalkylene, substituted cycloalkylene, substituted heterocycloalkylene, substituted arylene, and / or substituted heteroarylene described in the compounds herein are substituted with at least one substituent group. In embodiments, at least one or all of these groups are substituted with at least one size-limited substituent group. In embodiments, at least one or all of these groups are substituted with at least one lower substituent group.

[0095] In embodiments of the compounds herein, each substituted or unsubstituted alkyl may be a substituted or unsubstituted C1-C20 alkyl, each substituted or unsubstituted heteroalkyl is asubstituted or unsubstituted 2 to 20 membered heteroalkyl, each substituted or unsubstituted cycloalkyl is a substituted or unsubstituted C3-C8 cycloalkyl, each substituted or unsubstituted heterocycloalkyl is a substituted or unsubstituted 3 to 8 membered heterocycloalkyl, each substituted or unsubstituted aryl is a substituted or unsubstituted C6-C10 aryl, and / or each substituted or unsubstituted heteroaryl is a substituted or unsubstituted 5 to 10 membered heteroaryl. In embodiments of the compounds herein, each substituted or unsubstituted alkylene is a substituted or unsubstituted C1-C20alkylene, each substituted or unsubstituted heteroalkylene is a substituted or unsubstituted 2 to 20 membered heteroalkylene, each substituted or unsubstituted cycloalkylene is a substituted or unsubstituted C3-C8cycloalkylene, each substituted or unsubstituted heterocycloalkylene is a substituted or unsubstituted 3 to 8 membered heterocycloalkylene, each substituted or unsubstituted arylene is a substituted or unsubstituted C6-C10 arylene, and / or each substituted or unsubstituted heteroarylene is a substituted or unsubstituted 5 to 10 membered heteroarylene.

[0096] In embodiments, each substituted or unsubstituted alkyl is a substituted or unsubstituted C1-C8 alkyl, each substituted or unsubstituted heteroalkyl is a substituted or unsubstituted 2 to 8 membered heteroalkyl, each substituted or unsubstituted cycloalkyl is a substituted or unsubstituted C3-C7 cycloalkyl, each substituted or unsubstituted heterocycloalkyl is a substituted or unsubstituted 3 to 7 membered heterocycloalkyl, each substituted or unsubstituted aryl is a substituted or unsubstituted C6-C10 aryl, and / or each substituted or unsubstituted heteroaryl is a substituted or unsubstituted 5 to 9 membered heteroaryl. In embodiments, each substituted or unsubstituted alkylene is a substituted or unsubstituted C1-C8 alkylene, each substituted or unsubstituted heteroalkylene is a substituted or unsubstituted 2 to 8 membered heteroalkylene, each substituted or unsubstituted cycloalkylene is a substituted or unsubstituted C3-C7cycloalkylene, each substituted or unsubstituted heterocycloalkylene is a substituted or unsubstituted 3 to 7 membered heterocycloalkylene, each substituted or unsubstituted arylene is a substituted or unsubstituted C6-C10arylene, and / or each substituted or unsubstituted heteroarylene is a substituted or unsubstituted 5 to 9 membered heteroarylene.

[0097] In embodiments, a substituted or unsubstituted moiety (e.g., substituted or unsubstituted alkyl, substituted or unsubstituted heteroalkyl, substituted or unsubstituted cycloalkyl, substituted or unsubstituted heterocycloalkyl, substituted or unsubstituted aryl, substituted or unsubstituted heteroaryl, substituted or unsubstituted alkylene, substituted or unsubstituted heteroalkylene, substituted or unsubstituted cycloalkylene, substituted or unsubstituted heterocycloalkylene, substituted or unsubstituted arylene, and / or substituted or unsubstituted heteroarylene) is unsubstituted (e.g., is an unsubstituted alkyl, unsubstitutedheteroalkyl, unsubstituted cycloalkyl, unsubstituted heterocycloalkyl, unsubstituted aryl, unsubstituted heteroaryl, unsubstituted alkylene, unsubstituted heteroalkylene, unsubstituted cycloalkylene, unsubstituted heterocycloalkylene, unsubstituted arylene, and / or unsubstituted heteroarylene, respectively). In embodiments, a substituted or unsubstituted moiety (e.g., substituted or unsubstituted alkyl, substituted or unsubstituted heteroalkyl, substituted or unsubstituted cycloalkyl, substituted or unsubstituted heterocycloalkyl, substituted or unsubstituted aryl, substituted or unsubstituted heteroaryl, substituted or unsubstituted alkylene, substituted or unsubstituted heteroalkylene, substituted or unsubstituted cycloalkylene, substituted or unsubstituted heterocycloalkylene, substituted or unsubstituted arylene, and / or substituted or unsubstituted heteroarylene) is substituted (e.g., is a substituted alkyl, substituted heteroalkyl, substituted cycloalkyl, substituted heterocycloalkyl, substituted aryl, substituted heteroaryl, substituted alkylene, substituted heteroalkylene, substituted cycloalkylene, substituted heterocycloalkylene, substituted arylene, and / or substituted heteroarylene, respectively).

[0098] In embodiments, a substituted moiety (e.g., substituted alkyl, substituted heteroalkyl, substituted cycloalkyl, substituted heterocycloalkyl, substituted aryl, substituted heteroaryl, substituted alkylene, substituted heteroalkylene, substituted cycloalkylene, substituted heterocycloalkylene, substituted arylene, and / or substituted heteroarylene) is substituted with at least one substituent group, wherein if the substituted moiety is substituted with a plurality of substituent groups, each substituent group may optionally be different. In embodiments, if the substituted moiety is substituted with a plurality of substituent groups, each substituent group is different.

[0099] In embodiments, a substituted moiety (e.g., substituted alkyl, substituted heteroalkyl, substituted cycloalkyl, substituted heterocycloalkyl, substituted aryl, substituted heteroaryl, substituted alkylene, substituted heteroalkylene, substituted cycloalkylene, substituted heterocycloalkylene, substituted arylene, and / or substituted heteroarylene) is substituted with at least one size-limited substituent group, wherein if the substituted moiety is substituted with a plurality of size-limited substituent groups, each size-limited substituent group may optionally be different. In embodiments, if the substituted moiety is substituted with a plurality of size- limited substituent groups, each size-limited substituent group is different.

[0100] In embodiments, a substituted moiety (e.g., substituted alkyl, substituted heteroalkyl, substituted cycloalkyl, substituted heterocycloalkyl, substituted aryl, substituted heteroaryl, substituted alkylene, substituted heteroalkylene, substituted cycloalkylene, substituted heterocycloalkylene, substituted arylene, and / or substituted heteroarylene) is substituted with atleast one lower substituent group, wherein if the substituted moiety is substituted with a plurality of lower substituent groups, each lower substituent group may optionally be different. In embodiments, if the substituted moiety is substituted with a plurality of lower substituent groups, each lower substituent group is different.

[0101] In embodiments, a substituted moiety (e.g., substituted alkyl, substituted heteroalkyl, substituted cycloalkyl, substituted heterocycloalkyl, substituted aryl, substituted heteroaryl, substituted alkylene, substituted heteroalkylene, substituted cycloalkylene, substituted heterocycloalkylene, substituted arylene, and / or substituted heteroarylene) is substituted with at least one substituent group, size-limited substituent group, or lower substituent group; wherein if the substituted moiety is substituted with a plurality of groups selected from substituent groups, size-limited substituent groups, and lower substituent groups; each substituent group, size- limited substituent group, and / or lower substituent group may optionally be different. In embodiments, if the substituted moiety is substituted with a plurality of groups selected from substituent groups, size-limited substituent groups, and lower substituent groups; each substituent group, size-limited substituent group, and / or lower substituent group is different.

[0102] Certain compounds of the present disclosure possess asymmetric carbon atoms (optical or chiral centers) or double bonds; the enantiomers, racemates, diastereomers, tautomers, geometric isomers, stereoisometric forms that may be defined, in terms of absolute stereochemistry, as (R)-or (S)- or, as (D)- or (L)- for amino acids, and individual isomers are encompassed within the scope of the present disclosure. The compounds of the present disclosure do not include those that are known in art to be too unstable to synthesize and / or isolate. The present disclosure is meant to include compounds in racemic and optically pure forms. Optically active (R)- and (S)-, or (D)- and (L)-isomers may be prepared using chiral synthons or chiral reagents, or resolved using conventional techniques. When the compounds described herein contain olefinic bonds or other centers of geometric asymmetry, and unless specified otherwise, it is intended that the compounds include both E and Z geometric isomers. As used herein, the term “isomers” refers to compounds having the same number and kind of atoms, and hence the same molecular weight, but differing in respect to the structural arrangement or configuration of the atoms. The term “tautomer,” as used herein, refers to one of two or more structural isomers which exist in equilibrium and which are readily converted from one isomeric form to another. It will be apparent to one skilled in the art that certain compounds of this disclosure may exist in tautomeric forms, all such tautomeric forms of the compounds being within the scope of the disclosure. Unless otherwise stated, structures depicted herein are also meant to include all stereochemical forms of the structure; i.e., the R and S configurationsfor each asymmetric center. Therefore, single stereochemical isomers (stereoisomers) as well as enantiomeric and diastereomeric mixtures of the present compounds are within the scope of the disclosure.

[0103] The compounds described herein may also contain unnatural proportions of atomic isotopes at one or more of the atoms that constitute such compounds. For example, the compounds may be radiolabeled with radioactive isotopes, such as for example tritium (3H), iodine-125 (125I), or carbon-14 (14C). All isotopic variations of the compounds described herein, whether radioactive or not, are encompassed within the scope of the present disclosure.

[0104] It should be noted that throughout the application that alternatives are written in Markush groups, for example, each amino acid position that contains more than one possible amino acid. It is specifically contemplated that each member of the Markush group should be considered separately, thereby comprising another embodiment, and the Markush group is not to be read as a single unit.

[0105] “Analog,” or “analogue” is used in accordance with its plain ordinary meaning within Chemistry and Biology and refers to a chemical compound that is structurally similar to another compound (i.e., a so-called “reference” compound) but differs in composition, e.g., in the replacement of one atom by an atom of a different element, or in the presence of a particular functional group, or the replacement of one functional group by another functional group, or the absolute stereochemistry of one or more chiral centers of the reference compound. Accordingly, an analog is a compound that is similar or comparable in function and appearance but not in structure or origin to a reference compound.

[0106] The terms “a” or “an,” as used in herein means one or more. In addition, the phrase “substituted with a[n],” as used herein, means the specified group may be substituted with one or more of any or all of the named substituents. For example, where a group, such as an alkyl or heteroaryl group, is “substituted with an unsubstituted C1-C20alkyl, or unsubstituted 2 to 20 membered heteroalkyl,” the group may contain one or more unsubstituted C1-C20alkyls, and / or one or more unsubstituted 2 to 20 membered heteroalkyls.

[0107] Where a moiety is substituted with an R substituent, the group may be referred to as “R-substituted.” Where a moiety is R-substituted, the moiety is substituted with at least one R substituent and each R substituent is optionally different. Where a particular R group is present in the description of a chemical genus (such as Formula (1)), a Roman alphabetic symbol may be used to distinguish each appearance of that particular R group. For example, where multiple R3substituents are present, each R3substituent may be distinguished as R3A, R3B, wherein each of R3A, R3B, is defined within the scope of the definition of R3and optionally differently.

[0108] A person of ordinary skill in the art will understand when a variable (e.g., moiety orlinker) of a compound or of a compound genus (e.g., a genus described herein) is described by aname or formula of a standalone compound with all valencies filled, the unfilled valence(s) of the variable will be dictated by the context in which the variable is used. For example, when a variable of a compound as described herein is connected (e.g., bonded) to the remainder of the compound through a single bond, that variable is understood to represent a monovalent form (i.e., capable of forming a single bond due to an unfilled valence) of a standalone compound (e.g., if the variable is named “methane” in an embodiment but the variable is known to be attached by a single bond to the remainder of the compound, a person of ordinary skill in the art would understand that the variable is actually a monovalent form of methane, i.e., methyl or – CH3). Likewise, for a linker variable (e.g., L1, L2, or L3as described herein), a person of ordinary skill in the art will understand that the variable is the divalent form of a standalone compound (e.g., if the variable is assigned to “PEG” or “polyethylene glycol” in an embodiment but the variable is connected by two separate bonds to the remainder of the compound, a person of ordinary skill in the art would understand that the variable is a divalent (i.e., capable of forming two bonds through two unfilled valences) form of PEG instead of the standalone compound PEG).

[0109] The term “bond” or “bonded” refers to direct bonds, such as covalent bonds (e.g., direct or a linking group), or indirect bonds, such as non-covalent bond (e.g., electrostatic interactions (e.g., ionic bond, hydrogen bond, halogen bond), van der Waals interactions (e.g., dipole-dipole, dipole-induced dipole, London dispersion), ring stacking (pi effects), hydrophobic interactions, and the like).

[0110] The term “electron-withdrawing group” refers to a chemical moiety or substituent that removes electron density from a conjugated pi-electron system, thereby making the pi electron system less electrophilic.

[0111] The term “electron-donating group” refers to a chemical moiety or substituent that can donate electron density into a conjugated pi-electron system, thereby making the pi electron system more nucleophilic.

[0112] The terms “bind” and “bound” as used herein is used in accordance with its plain and ordinary meaning and refers to the association between atoms or molecules. The association can be direct or indirect. For example, bound atoms or molecules may be bound, e.g., by covalent bond, linker (e.g. a first linker or second linker), or non-covalent bond (e.g. electrostatic interactions (e.g. ionic bond, hydrogen bond, halogen bond), van der Waals interactions (e.g. dipole-dipole, dipole-induced dipole, London dispersion), ring stacking (pi effects), hydrophobicinteractions and the like).

[0113] The term “capable of binding” as used herein refers to a moiety (e.g., a single-domain antibody or a recombinant protein as described herein, i.e., comprising an unnatural amino acid side chain that is capable of binding to an amino acid residue on a different protein) that is able to measurably bind to a target. In aspects, where a moiety is capable of binding a target, the moiety is capable of binding with a Kd of less than about 10 µM, 5 µM, 1 µM, 500 nM, 250 nM, 100 nM, 75 nM, 50 nM, 25 nM, 15 nM, 10 nM, 5 nM, 1 nM, or about 0.1 nM.

[0114] “Arginine” refers to the amino acid having the following structure: arginine” and a

[0115] The term “naturally-occurring arginine” refers to an arginine that naturally occurs at a given position in a protein. In embodiments, the “naturally occurring arginine” is proximal in a three-dimensional space to an unnatural amino acid described herein, such that the non-naturally occurring arginine can interact (e.g., through van der walls forces) with the –L4P(=O)(F)(NR2R3) group in the unnatural amino acid side chain.

[0116] The term “non-naturally occurring arginine” refers to an arginine that is a point mutation of a different amino acid (e.g., Ala, Ile, Leu, Met, Val, Phe, Trp, Tyr, Asn, Cys, Gln, Ser, Thr, Asp, Glu, His, Lys, Gly, Pro) that naturally occurs in a protein. In embodiments, the “non-naturally occurring arginine” is proximal in a three-dimensional space to an unnatural amino acid described herein, such that the non-naturally occurring arginine can interact (e.g., through van der walls forces) with the –L4P(=O)(F)(NR2R3) group in the unnatural amino acid side chain.

[0117] The term “proximal” means a spatial distance that an arginine (either non-naturally occurring arginine or naturally-occurring arginine) can interact through van der Waals forces with the –L4P(=O)(F)(NR2R3) group of an unnatural amino acid as described herein (including embodiments thereof). In embodiments, “proximal” means that an arginine (either non-naturally occurring arginine or naturally occurring arginine) can facilitate the –F departure from the –L4P(=O)(F)(NR2R3) group in the unnatural amino acid when the –L4P(=O)(F)(NR2R3) groupreacts with the lysine, histidine, tyrosine, or cysteine of a target protein. The skilled artisan will readily appreciate that “proximal” is based on the three-dimensional structure of the protein. In embodiments, “proximal” means up to about 25 angstroms. In embodiments, “proximal” means up to about 20 angstroms. In embodiments, “proximal” means up to about 15 angstroms. In embodiments, “proximal” means up to about 10 angstroms. In embodiments, “proximal” means from about 1 angstrom to about 25 angstroms. In embodiments, “proximal” means from about 1 angstrom to about 20 angstroms. In embodiments, “proximal” means from about 1 angstrom to about 15 angstroms. In embodiments, “proximal” means from about 1 angstrom to about 12 angstroms. In embodiments, “proximal” means from about 1 angstrom to about 10 angstroms. In embodiments, “proximal” means from about 1 angstrom to about 8 angstroms. In embodiments, “proximal” means from about 1 angstrom to about 6 angstroms. In embodiments, “proximal” means from about 1 angstrom to about 5 angstroms. In embodiments, “proximal” means from about 1 angstroms to about 4 angstroms.

[0118] In embodiments, the term “proximal” means that the naturally or non-naturally occurring arginine is within 1 to about 10 amino acid residues of the unnatural amino acid. In embodiments, the naturally or non-naturally occurring arginine is within 1 to about 9 amino acid residues of the unnatural amino acid. In embodiments, the naturally or non-naturally occurring arginine is within 1 to about 8 amino acid residues of the unnatural amino acid. In embodiments, the naturally or non-naturally occurring arginine is within 1 to about 7 amino acid residues of the unnatural amino acid. In embodiments, the naturally or non-naturally occurring arginine is within 1 to about 6 amino acid residues of the unnatural amino acid. In embodiments, the naturally or non-naturally occurring arginine is within 1 to about 5 amino acid residues of the unnatural amino acid. In embodiments, the naturally or non-naturally occurring arginine is within 1 to about 4 amino acid residues of the unnatural amino acid. In embodiments, the naturally or non-naturally occurring arginine is within 1 to about 3 amino acid residues of the unnatural amino acid. In embodiments, the naturally or non-naturally occurring arginine is within 1 to about 2 amino acid residues of the unnatural amino acid. In embodiments, the naturally or non-naturally occurring arginine is within 2 to about 6 amino acid residues of the unnatural amino acid. In embodiments, the naturally or non-naturally occurring arginine is within 2 to about 5 amino acid residues of the unnatural amino acid. In embodiments, the naturally or non-naturally occurring arginine is within 2 to about 4 amino acid residues of the unnatural amino acid. In embodiments, the naturally or non-naturally occurring arginine is within 2 to about 3 amino acid residues of the unnatural amino acid. In embodiments, the naturally or non-naturally occurring arginine is within 1 amino acid residue of the unnaturalamino acid. In embodiments, the naturally or non-naturally occurring arginine is within 2 amino acid residues of the unnatural amino acid. In embodiments, the naturally or non-naturally occurring arginine is within 3 amino acid residues of the unnatural amino acid. In embodiments, the naturally or non-naturally occurring arginine is within 4 amino acid residues of the unnatural amino acid. In embodiments, the naturally or non-naturally occurring arginine is within 5 amino acid residues of the unnatural amino acid. The phrase “within 1” means that the naturally or non- naturally-occurring arginine is next to the unnatural amin acid. The phrase “within 2” means that there is a single amino acid residue between the naturally or non-naturally-occurring arginine and the unnatural amino acid.

[0119] Compounds

[0120] Provided herein are compounds of Formula (1) or a stereoisomer thereof: O wherein L4is a to 8; L1is a bond,substituted or unsubstituted alkylene, or substituted or unsubstituted heteroalkylene; R1is hydrogen or an electron withdrawing group; and R2and R3are each independently substituted or unsubstituted C1-5 alkyl. In embodiments, L4is a bond. In embodiments, L4is -(CH2)1-5-. In embodiments, L4is –O-. In embodiments, L4is -O-(CH2)1-5-. When L4is -O-(CH2)1-5-, the oxygen atom is adjacent to the phenyl ring and the alkylene group is adjacent to the phosphorous atom. In embodiments, L1is a bond or substituted or unsubstituted heteroalkylene. In embodiments, R2and R3are each independently unsubstituted C1-5 alkyl. The substituents are described in more detail below.

[0121] In embodiments, the compound of Formula (1) is a compound of Formula (IA) or a stereoisomer thereof: O , wherein x is analkylene, or substituted or unsubstituted heteroalkylene; R1is hydrogen or an electron withdrawing group; and R2and R3are each independently substituted or unsubstituted C1-5alkyl. In embodiments, L1is a bond or substituted or unsubstituted heteroalkylene. In embodiments, R2and R3are eachindependently unsubstituted C1-5alkyl. The substituents are described in more detail below.

[0122] In embodiments, the compound of Formula (1) is a compound of Formula (1B) or a stereoisomer thereof: , wherein x is an alkylene, orsubstituted or group; and R2and R3are each independently substituted or unsubstituted C1-5 alkyl. In embodiments, L1is a bond or substituted or unsubstituted heteroalkylene. In embodiments, R2and R3are each independently unsubstituted C1-5 alkyl. The substituents are described in more detail below.

[0123] In embodiments, the compound of Formula (1) is a compound of Formula (1C) or a stereoisomer thereof: , wherein x is analkylene, or substituted or unsubstituted heteroalkylene; and R1is hydrogen or an electron withdrawing group. In embodiments, L1is a bond or substituted or unsubstituted heteroalkylene. The substituents are described in more detail below.

[0124] In embodiments, the compound of Formula (1) is a compound of Formula (1D) or a stereoisomer thereof: , wherein x is analkylene, or substituted or unsubstituted heteroalkylene; and R2and R3are each independently substituted orunsubstituted C1-5alkyl. In embodiments, L1is a bond or substituted or unsubstituted heteroalkylene. In embodiments, R2and R3are each independently unsubstituted C1-5 alkyl. The substituents are described in more detail below.

[0125] In embodiments, the compound of Formula (1) is a compound of Formula (1E) or a stereoisomer thereof: , wherein x is an alkylene, orsubstituted or a or substituted or unsubstituted heteroalkylene. The substituents are described in more detail below.

[0126] In embodiments, the compound of Formula (1) is PFY: .

[0127] InH2N .

[0128] , wherein x is ansubstituted or unsubstituted alkylene, or substituted or unsubstituted heteroalkylene; R1is hydrogen or anelectron withdrawing group; and R2is substituted or unsubstituted C1-5alkyl. In embodiments, x1 is 0. In embodiments, x1 is 1. In embodiments, x1 is 2. In embodiments, x1 is 3. In embodiments, x1 is 4. In embodiments, x1 is 5. In embodiments, L1is a bond or substituted or unsubstituted heteroalkylene. In embodiments, R2is unsubstituted C1-5 alkyl. The substituents are described in more detail below.

[0129] In embodiments, the compound of Formula (5) is a compound of Formula (5A) or a stereoisomer thereof: , wherein x is an substituted orunsubstituted or or R2is substituted or unsubstituted C1-5alkyl. In embodiments, x1 is 0. In embodiments, x1 is 1. In embodiments, x1 is 2. In embodiments, x1 is 3. In embodiments, x1 is 4. In embodiments, x1 is 5. In embodiments, L1is a bond or substituted or unsubstituted heteroalkylene. In embodiments, R2is unsubstituted C1-5 alkyl. The substituents are described in more detail below.

[0130] In embodiments, the compound of Formula (5) is a compound of Formula (5B) or a stereoisomer thereof: R1O O , wherein x is ansubstituted or unsubstituted alkylene, or substituted or unsubstituted heteroalkylene; and R1is hydrogen or an electron withdrawing group. In embodiments, x1 is 0. In embodiments, x1 is 1. In embodiments, x1 is 2. In embodiments, x1 is 3. In embodiments, x1 is 4. In embodiments, x1 is 5. In embodiments, L1is a bond or substituted or unsubstituted heteroalkylene. The substituents are described in more detail below.

[0131] In embodiments, the compound of Formula (5) is a compound of Formula (5C) or a stereoisomer thereof:C), wherein x is an int nd, substituted or unsubstituted alkylene, or substituted or unsubstituted heteroalkylene; and R1is hydrogen or an electron withdrawing group. In embodiments, x1 is 0. In embodiments, x1 is 1. In embodiments, x1 is 2. In embodiments, x1 is 3. In embodiments, x1 is 4. In embodiments, x1 is 5. In embodiments, L1is a bond or substituted or unsubstituted heteroalkylene. The substituents are described in more detail below.

[0132] In embodiments, the compound of Formula (5) is a compound of Formula (5D) or a stereoisomer thereof: .

[0133] Inof Formula (5E) or a stereoisomer thereof: .

[0134]

[0135] herein are proteins comprising an unnatural amino acid, wherein the unnatural amino comprises a side chain of Formula (2): wherein L4is a bond,1 to 8; L1is a bond, substituted or unsubstituted alkylene, or substituted or unsubstituted heteroalkylene; R1is hydrogen or an electron withdrawing group; and R2and R3are each independently substituted or unsubstituted C1-5 alkyl. In embodiments, L4is a bond. In embodiments, L4is -(CH2)1-5-. Inembodiments, L4is –O-. In embodiments, L4is -O-(CH2)1-5-. When L4is -O-(CH2)1-5-, the oxygen atom is adjacent to the phenyl ring and the alkylene group is adjacent to the phosphorous atom. In embodiments, L1is a bond or substituted or unsubstituted heteroalkylene. In embodiments, R2and R3are each independently unsubstituted C1-5 alkyl. The substituents are described in more detail below. In embodiments, the protein further comprises a non-naturally occurring arginine proximal to the unnatural amino acid and / or a naturally occurring arginine proximal to the unnatural amino acid. In embodiments, the protein further comprises a non- naturally occurring arginine proximal to the unnatural amino acid. In embodiments, the protein further comprises a naturally occurring arginine proximal to the unnatural amino acid.

[0136] In embodiments, the proteins comprises an unnatural amino acid, wherein the unnatural amino comprises a side chain of: , wherein x is an alkylene, orsubstituted or unsubstituted heteroalkylene; R1is hydrogen or an electron withdrawing group; and R2and R3are each independently substituted or unsubstituted C1-5 alkyl. In embodiments, L1is a bond or substituted or unsubstituted heteroalkylene. In embodiments, R2and R3are each independently unsubstituted C1-5 alkyl. The substituents are described in more detail below. In embodiments, the protein further comprises a non-naturally occurring arginine proximal to the unnatural amino acid and / or a naturally occurring arginine proximal to the unnatural amino acid. In embodiments, the protein further comprises a non-naturally occurring arginine proximal to the unnatural amino acid. In embodiments, the protein further comprises a naturally occurring arginine proximal to the unnatural amino acid.

[0137] In embodiments, the proteins comprises an unnatural amino acid, wherein the unnatural amino comprises a side chain of: , wherein x is analkylene, or substituted or unsubstituted heteroalkylene; R1is hydrogen or an electron withdrawing group;and R2and R3are each independently substituted or unsubstituted C1-5alkyl. In embodiments, L1is a bond or substituted or unsubstituted heteroalkylene. In embodiments, R2and R3are each independently unsubstituted C1-5alkyl. The substituents are described in more detail below. In embodiments, the protein further comprises a non-naturally occurring arginine proximal to the unnatural amino acid and / or a naturally occurring arginine proximal to the unnatural amino acid. In embodiments, the protein further comprises a non-naturally occurring arginine proximal to the unnatural amino acid. In embodiments, the protein further comprises a naturally occurring arginine proximal to the unnatural amino acid.

[0138] In embodiments, the proteins comprises an unnatural amino acid, wherein the unnatural amino comprises a side chain of: , wherein x is an alkylene, orsubstituted or unsubstituted heteroalkylene; and R1is hydrogen or an electron withdrawing group. In embodiments, L1is a bond or substituted or unsubstituted heteroalkylene. The substituents are described in more detail below. In embodiments, the protein further comprises a non-naturally occurring arginine proximal to the unnatural amino acid and / or a naturally occurring arginine proximal to the unnatural amino acid. In embodiments, the protein further comprises a non-naturally occurring arginine proximal to the unnatural amino acid. In embodiments, the protein further comprises a naturally occurring arginine proximal to the unnatural amino acid.

[0139] In embodiments, the proteins comprises an unnatural amino acid, wherein the unnatural amino comprises a side chain of: , wherein x is analkylene, or substituted or unsubstituted heteroalkylene; and R2and R3are each independently substituted or unsubstituted C1-5alkyl. In embodiments, L1is a bond or substituted or unsubstitutedheteroalkylene. In embodiments, R2and R3are each independently unsubstituted C1-5alkyl. The substituents are described in more detail below. In embodiments, the protein further comprises a non-naturally occurring arginine proximal to the unnatural amino acid and / or a naturally occurring arginine proximal to the unnatural amino acid. In embodiments, the protein further comprises a non-naturally occurring arginine proximal to the unnatural amino acid. In embodiments, the protein further comprises a naturally occurring arginine proximal to the unnatural amino acid.

[0140] In embodiments, the proteins comprises an unnatural amino acid, wherein the unnatural amino comprises a side chain of: , wherein x is an alkylene, orsubstituted or unsubstituted heteroalkylene. In embodiments, L1is a bond or substituted or unsubstituted heteroalkylene. The substituents are described in more detail below. In embodiments, the protein further comprises a non-naturally occurring arginine proximal to the unnatural amino acid and / or a naturally occurring arginine proximal to the unnatural amino acid. In embodiments, the protein further comprises a non-naturally occurring arginine proximal to the unnatural amino acid. In embodiments, the protein further comprises a naturally occurring arginine proximal to the unnatural amino acid.

[0141] In embodiments, the proteins comprises an unnatural amino acid, wherein the unnatural amino comprises a side chain of: . In embodiments, thearginine proximal to the unnatural amino acid and / or a naturally occurring arginine proximal to the unnatural amino acid. In embodiments, the protein further comprises a non-naturally occurring arginine proximal to the unnatural amino acid. In embodiments, the protein further comprises a naturally occurring arginine proximal to the unnatural amino acid.

[0142] In embodiments, the proteins comprises an unnatural amino acid, wherein the unnatural amino comprises a side chain of: . In embodiments, proximal tothe unnatural unnatural amino acid. In embodiments, the protein further comprises a non-naturally occurring arginine proximal to the unnatural amino acid. In embodiments, the protein further comprises a naturally occurring arginine proximal to the unnatural amino acid.

[0143] Provided herein are proteins comprising an unnatural amino acid, wherein the unnatural amino comprises a side chain of Formula (6): , wherein x is an integera bond, substituted or unsubstituted alkylene, or substituted or unsubstituted heteroalkylene; R1is hydrogen or an electron withdrawing group; and R2is substituted or unsubstituted C1-5 alkyl. In embodiments, x1 is 0. In embodiments, x1 is 1. In embodiments, x1 is 2. In embodiments, x1 is 3. In embodiments, x1 is 4. In embodiments, x1 is 5. In embodiments, L1is a bond or substituted or unsubstituted heteroalkylene. In embodiments, R2is unsubstituted C1-5alkyl. The substituents are described in more detail below. In embodiments, the protein further comprises a non- naturally occurring arginine proximal to the unnatural amino acid and / or a naturally occurring arginine proximal to the unnatural amino acid. In embodiments, the protein further comprises a non-naturally occurring arginine proximal to the unnatural amino acid. In embodiments, the protein further comprises a naturally occurring arginine proximal to the unnatural amino acid.

[0144] In embodiments, the proteins comprises an unnatural amino acid, wherein the unnatural amino comprises a side chain of:A), wherein x is an integer f a bond, substituted or unsubstituted alkylene, or substituted or unsubstituted heteroalkylene; and R2is substituted or unsubstituted C1-5 alkyl. In embodiments, x1 is 0. In embodiments, x1 is 1. In embodiments, x1 is 2. In embodiments, x1 is 3. In embodiments, x1 is 4. In embodiments, x1 is 5. In embodiments, L1is a bond or substituted or unsubstituted heteroalkylene. In embodiments, R2is unsubstituted C1-5alkyl. The substituents are described in more detail below. In embodiments, the protein further comprises a non-naturally occurring arginine proximal to the unnatural amino acid and / or a naturally occurring arginine proximal to the unnatural amino acid. In embodiments, the protein further comprises a non-naturally occurring arginine proximal to the unnatural amino acid. In embodiments, the protein further comprises a naturally occurring arginine proximal to the unnatural amino acid.

[0145] In embodiments, the proteins comprises an unnatural amino acid, wherein the unnatural amino comprises a side chain of: , wherein x is an integerbond, substituted or unsubstituted alkylene, or substituted or unsubstituted heteroalkylene; and R1is hydrogen or an electron withdrawing group. In embodiments, x1 is 0. In embodiments, x1 is 1. In embodiments, x1 is 2. In embodiments, x1 is 3. In embodiments, x1 is 4. In embodiments, x1 is 5. In embodiments, L1is a bond or substituted or unsubstituted heteroalkylene. The substituents are described in more detail below. In embodiments, the protein further comprises a non-naturally occurring arginine proximal to the unnatural amino acid and / or a naturally occurring arginine proximal to the unnatural amino acid. In embodiments, the protein further comprises a non- naturally occurring arginine proximal to the unnatural amino acid. In embodiments, the protein further comprises a naturally occurring arginine proximal to the unnatural amino acid.

[0146] In embodiments, the proteins comprises an unnatural amino acid, wherein the unnatural amino comprises a side chain of:C), wherein x is an integer a bond, substituted or unsubstituted alkylene, or substituted or unsubstituted heteroalkylene; and R1is hydrogen or an electron withdrawing group. In embodiments, x1 is 0. In embodiments, x1 is 1. In embodiments, x1 is 2. In embodiments, x1 is 3. In embodiments, x1 is 4. In embodiments, x1 is 5. In embodiments, L1is a bond or substituted or unsubstituted heteroalkylene. The substituents are described in more detail below. In embodiments, the protein further comprises a non-naturally occurring arginine proximal to the unnatural amino acid and / or a naturally occurring arginine proximal to the unnatural amino acid. In embodiments, the protein further comprises a non- naturally occurring arginine proximal to the unnatural amino acid. In embodiments, the protein further comprises a naturally occurring arginine proximal to the unnatural amino acid.

[0147] In embodiments, the proteins comprises an unnatural amino acid, wherein the unnatural amino comprises a side chain of: . In embodiments, thearginine proximal to the unnatural amino acid and / or a naturally occurring arginine proximal to the unnatural amino acid. In embodiments, the protein further comprises a non-naturally occurring arginine proximal to the unnatural amino acid. In embodiments, the protein further comprises a naturally occurring arginine proximal to the unnatural amino acid.

[0148] In embodiments, the proteins comprises an unnatural amino acid, wherein the unnatural amino comprises a side chain of: . Inproximal to the unnatural amino acid and / or a naturally occurring arginine proximal to the unnatural aminoacid. In embodiments, the protein further comprises a non-naturally occurring arginine proximal to the unnatural amino acid. In embodiments, the protein further comprises a naturally occurring arginine proximal to the unnatural amino acid.

[0149] Provided herein are proteins of Formula (3) or a stereoisomer thereof: wherein W is –H, moiety, oran amino acid - 1-5-, - 1-5-, ; integer from 1 to 8; L1is a bond, substituted or unsubstituted alkylene, or substituted or unsubstituted heteroalkylene; R1is hydrogen or an electron withdrawing group; and R2and R3are each independently substituted or unsubstituted C1-5 alkyl. In embodiments, L4is a bond. In embodiments, L4is -(CH2)1-5-. In embodiments, L4is –O-. In embodiments, L4is -O-(CH2)1-5-. When L4is -O-(CH2)1-5-, the oxygen atom is adjacent to the phenyl ring and the alkylene group is adjacent to the phosphorous atom. In embodiments, L1is a bond or substituted or unsubstituted heteroalkylene. In embodiments, R2and R3are each independently unsubstituted C1-5 alkyl. The substituents are described in more detail below. In embodiments, W and / or Y comprises a non-naturally occurring arginine proximal to the unnatural amino acid and / or a naturally occurring arginine proximal to the unnatural amino acid. In embodiments, W or Y further comprises a non-naturally occurring arginine proximal to the unnatural amino acid. In embodiments, W and / or Y further comprises a naturally occurring arginine proximal to the unnatural amino acid.

[0150] In embodiments, the protein of Formula (3) is a protein of Formula (3A) or a stereoisomer thereof: , wherein W ismoiety, or an amino acid moiety; x is an integer from 1 to 8; L1is a bond, substituted or unsubstitutedalkylene, or substituted or unsubstituted heteroalkylene; R1is hydrogen or an electron withdrawing group; and R2and R3are each independently substituted or unsubstituted C1-5 alkyl. In embodiments, L1is a bond or substituted or unsubstituted heteroalkylene. In embodiments, R2and R3are each independently unsubstituted C1-5 alkyl. In embodiments, W is not –H when Y is –OH. The substituents are described in more detail below. In embodiments, W and / or Y comprises a non-naturally occurring arginine proximal to the unnatural amino acid and / or a naturally occurring arginine proximal to the unnatural amino acid. In embodiments, W or Y further comprises a non-naturally occurring arginine proximal to the unnatural amino acid. In embodiments, W and / or Y further comprises a naturally occurring arginine proximal to the unnatural amino acid.

[0151] In embodiments, the protein of Formula (3) is a protein of Formula (3B) or a stereoisomer thereof: , wherein W ismoiety, or an amino acid moiety; x is an integer from 1 to 8; L1is a bond, substituted or unsubstituted alkylene, or substituted or unsubstituted heteroalkylene; R1is hydrogen or an electron withdrawing group; and R2and R3are each independently substituted or unsubstituted C1-5 alkyl. In embodiments, L1is a bond or substituted or unsubstituted heteroalkylene. In embodiments, R2and R3are each independently unsubstituted C1-5 alkyl. In embodiments, W is not –H when Y is –OH. The substituents are described in more detail below. In embodiments, W and / or Y comprises a non-naturally occurring arginine proximal to the unnatural amino acid and / or a naturally occurring arginine proximal to the unnatural amino acid. In embodiments, W or Y further comprises a non-naturally occurring arginine proximal to the unnatural amino acid. In embodiments, W and / or Y further comprises a naturally occurring arginine proximal to the unnatural amino acid.

[0152] In embodiments, the protein of Formula (3) is a protein of Formula (3C) or a stereoisomer thereof:C), wherein W is – ptidyl moiety, or an amino acid moiety; x is an integer from 1 to 8; L is a bond, substituted or unsubstituted alkylene, or substituted or unsubstituted heteroalkylene; and R1is hydrogen or an electron withdrawing group. In embodiments, L1is a bond or substituted or unsubstituted heteroalkylene. In embodiments, W is not –H when Y is –OH. The substituents are described in more detail below. In embodiments, W and / or Y comprises a non-naturally occurring arginine proximal to the unnatural amino acid and / or a naturally occurring arginine proximal to the unnatural amino acid. In embodiments, W or Y further comprises a non-naturally occurring arginine proximal to the unnatural amino acid. In embodiments, W and / or Y further comprises a naturally occurring arginine proximal to the unnatural amino acid.

[0153] In embodiments, the protein of Formula (3) is a protein of Formula (3D) or a stereoisomer thereof: , wherein W ismoiety, or an amino acid moiety; x is an integer from 1 to 8; L1is a bond, substituted or unsubstituted alkylene, or substituted or unsubstituted heteroalkylene; and R2and R3are each independently substituted or unsubstituted C1-5 alkyl. In embodiments, L1is a bond or substituted or unsubstituted heteroalkylene. In embodiments, R2and R3are each independently unsubstituted C1-5 alkyl. In embodiments, W is not –H when Y is –OH. The substituents are described in more detail below. In embodiments, W and / or Y comprises a non-naturally occurring arginine proximal to the unnatural amino acid and / or a naturally occurring arginine proximal to the unnatural amino acid. In embodiments, W or Y further comprises a non-naturally occurringarginine proximal to the unnatural amino acid. In embodiments, W and / or Y further comprises a naturally occurring arginine proximal to the unnatural amino acid.

[0154] In embodiments, the protein of Formula (3) is a protein of Formula (3E) or a stereoisomer thereof: , wherein W is moiety, oran amino acid x an a or unsubstituted alkylene, or substituted or unsubstituted heteroalkylene. In embodiments, L1is a bond or substituted or unsubstituted heteroalkylene. In embodiments, W is not –H when Y is –OH. The substituents are described in more detail below. In embodiments, W and / or Y comprises a non- naturally occurring arginine proximal to the unnatural amino acid and / or a naturally occurring arginine proximal to the unnatural amino acid. In embodiments, W or Y further comprises a non-naturally occurring arginine proximal to the unnatural amino acid. In embodiments, W and / or Y further comprises a naturally occurring arginine proximal to the unnatural amino acid.

[0155] In embodiments, the protein of Formula (3) is a protein of Formula (3F) or a stereoisomer thereof: , W is –H, a peptidylmoiety, or an amino acid moiety. In embodiments, W is not –H when Y is –OH. The substituents are described in more detail below. In embodiments, W and / or Y comprises a non-naturally occurring arginine proximal to the unnatural amino acid and / or a naturally occurring arginine proximal to the unnatural amino acid. In embodiments, W or Y further comprises a non-naturally occurringarginine proximal to the unnatural amino acid. In embodiments, W and / or Y further comprises a naturally occurring arginine proximal to the unnatural amino acid.

[0156] In embodiments, the protein of Formula (3) is a protein of Formula (3G) or a stereoisomer thereof: : , W is –H, a or an aminoacid moiety. are in more detail below. In embodiments, W and / or Y comprises a non-naturally occurring arginine proximal to the unnatural amino acid and / or a naturally occurring arginine proximal to the unnatural amino acid. In embodiments, W or Y further comprises a non-naturally occurring arginine proximal to the unnatural amino acid. In embodiments, W and / or Y further comprises a naturally occurring arginine proximal to the unnatural amino acid.

[0157] Provided herein are proteins of Formula (7) or a stereoisomer thereof: , wherein W is –H, apeptidyl moiety, or an amino acid moiety; x is an integer from 1 to 8; x1 is an integer from 0 to 5; L1is a bond, substituted or unsubstituted alkylene, or substituted or unsubstituted heteroalkylene; R1is hydrogen or an electron withdrawing group; and R2is substituted or unsubstituted C1-5 alkyl. In embodiments, x1 is 0. In embodiments, x1 is 1. In embodiments, x1 is 2. In embodiments, x1 is 3. In embodiments, x1 is 4. In embodiments, x1 is 5. In embodiments, L1is a bond or substituted or unsubstituted heteroalkylene. In embodiments, R2is unsubstituted C1-5alkyl. In embodiments, W is not –H when Y is –OH. The substituents are described in more detail below. In embodiments, W and / or Y comprises a non-naturally occurring arginine proximal to the unnatural amino acid and / or a naturally occurring arginine proximal to the unnatural amino acid. In embodiments, W or Y further comprises a non-naturally occurring arginine proximal to theunnatural amino acid. In embodiments, W and / or Y further comprises a naturally occurring arginine proximal to the unnatural amino acid.

[0158] In embodiments, the protein of Formula (7) is a protein of Formula (7A) or a stereoisomer thereof: , wherein W is –H, a peptidyl moiety, oran amino acid 5; L1is a bond, substituted or unsubstituted alkylene, or substituted or unsubstituted heteroalkylene; and R2is substituted or unsubstituted C1-5 alkyl. In embodiments, x1 is 0. In embodiments, x1 is 1. In embodiments, x1 is 2. In embodiments, x1 is 3. In embodiments, x1 is 4. In embodiments, x1 is 5. In embodiments, L1is a bond or substituted or unsubstituted heteroalkylene. In embodiments, R2is unsubstituted C1-5 alkyl. In embodiments, W is not –H when Y is –OH. The substituents are described in more detail below. In embodiments, W and / or Y comprises a non-naturally occurring arginine proximal to the unnatural amino acid and / or a naturally occurring arginine proximal to the unnatural amino acid. In embodiments, W or Y further comprises a non- naturally occurring arginine proximal to the unnatural amino acid. In embodiments, W and / or Y further comprises a naturally occurring arginine proximal to the unnatural amino acid.

[0159] In embodiments, the protein of Formula (7) is a protein of Formula (7B) or a stereoisomer thereof: , wherein W is –H,peptidyl moiety, or an amino acid moiety; x is an integer from 1 to 8; x1 is an integer from 0 to 5; L1is a bond, substituted or unsubstituted alkylene, or substituted or unsubstituted heteroalkylene; and R1is hydrogen or an electron withdrawing group. In embodiments, x1 is 0. In embodiments, x1 is 1. In embodiments, x1 is 2. In embodiments, x1 is 3. In embodiments, x1 is 4. In embodiments, x1is 5. In embodiments, L1is a bond or substituted or unsubstituted heteroalkylene. In embodiments, W is not –H when Y is –OH. The substituents are described in more detail below. In embodiments, W and / or Y comprises a non-naturally occurring arginine proximal to the unnatural amino acid and / or a naturally occurring arginine proximal to the unnatural amino acid. In embodiments, W or Y further comprises a non-naturally occurring arginine proximal to the unnatural amino acid. In embodiments, W and / or Y further comprises a naturally occurring arginine proximal to the unnatural amino acid.

[0160] In embodiments, the protein of Formula (7) is a protein of Formula (7C) or a stereoisomer thereof: , wherein W is –H, peptidyl moiety, oran amino acid moiety; x is an integer from 1 to 8; x1 is an integer from 0 to 5; L1is a bond, substituted or unsubstituted alkylene, or substituted or unsubstituted heteroalkylene; and R1is hydrogen or an electron withdrawing group. In embodiments, x1 is 0. In embodiments, x1 is 1. In embodiments, x1 is 2. In embodiments, x1 is 3. In embodiments, x1 is 4. In embodiments, x1 is 5. In embodiments, L1is a bond or substituted or unsubstituted heteroalkylene. In embodiments, W is not –H when Y is –OH. The substituents are described in more detail below. In embodiments, W and / or Y comprises a non-naturally occurring arginine proximal to the unnatural amino acid and / or a naturally occurring arginine proximal to the unnatural amino acid. In embodiments, W or Y further comprises a non-naturally occurring arginine proximal to the unnatural amino acid. In embodiments, W and / or Y further comprises a naturally occurring arginine proximal to the unnatural amino acid.

[0161] In embodiments, the protein of Formula (7) is a protein of Formula (7D) or a stereoisomer thereof:W is –H, a peptidyl moiety, or an amino acid moiety; Y is –OH, a peptidyl moiety, or an amino acid moiety. In embodiments, W is not –H when Y is –OH. The substituents are described in more detail below. In embodiments, W and / or Y comprises a non-naturally occurring arginine proximal to the unnatural amino acid and / or a naturally occurring arginine proximal to the unnatural amino acid. In embodiments, W or Y further comprises a non-naturally occurring arginine proximal to the unnatural amino acid. In embodiments, W and / or Y further comprises a naturally occurring arginine proximal to the unnatural amino acid.

[0162] In embodiments, the protein of Formula (7) is a protein of Formula (7E) or a stereoisomer thereof: , W is –H, or an aminoacid moiety. In embodiments, W is not –H when Y is –OH. The substituents are described in more detail below. In embodiments, W and / or Y comprises a non-naturally occurring arginine proximal to the unnatural amino acid and / or a naturally occurring arginine proximal to the unnatural amino acid. In embodiments, W or Y further comprises a non-naturally occurring arginine proximal to the unnatural amino acid. In embodiments, W and / or Y further comprises a naturally occurring arginine proximal to the unnatural amino acid.

[0163] In embodiments of the compounds described herein, the protein is an antibody, an antibody variant, or a receptor protein. In embodiments, the protein is an antibody. In embodiments, the protein is an antibody variant. In embodiments, the protein is a receptor protein. In embodiments, the antibody variant is a variant as defined herein. In embodiments, the antibody variant is a single-chain variable fragment, a single-domain antibody, an affibody, or an antigen-binding fragment. In embodiments, the antibody variant is a single-chain variable fragment. In embodiments, the antibody variant is a single-domain antibody. In embodiments, the antibody variant is an affibody. In embodiments, the antibody variant is or an antigen- binding fragment. In embodiments, the receptor protein is any receptor protein described herein.

[0164] In embodiments of the compounds described herein, the protein is a receptor protein. In embodiments, the receptor protein is a programmed death-ligand 1 (PD-L1) receptor, a programmed cell death protein 1 (PD-1) receptor, a 5-hydroxytryptamine receptor, anacetylcholine receptor, an adenosine receptor, an adenosine A2A receptor, an adenosine A2B receptor, an angiotensin receptor, an apelin receptor, a bile acid receptor, a bombesin receptor, a bradykinin receptor, a cannabinoid receptor, a chemerin receptor, a chemokine receptor, a cholecystokinin receptor, a Class A Orphan receptor, a dopamine receptor, an endothelin receptor, an epidermal growth factor receptor (EGFR), a formyl peptide receptor, a free fatty acid receptor, a galanin receptor, a ghrelin receptor, a glycoprotein hormone receptor, a gonadotrophin-releasing hormone receptor, a G protein-coupled receptor, a G protein-coupled estrogen receptor, a histamine receptor, a hydroxycarboxylic acid receptor, a kisspeptin receptor, a leukotriene receptor, a lysophospholipid receptor, a lysophospholipid S1P receptor, a melanin- concentrating hormone receptor, a melanocortin receptor, a melatonin receptor, a motilin receptor, a neuromedin U receptor, a neuropeptide FF / neuropeptide AF receptor, a neuropeptide S receptor, a neuropeptide W / neuropeptide B receptor, a neuropeptide Y receptor, a neurotensin receptor, an opioid receptor, an opsin receptor, an orexin receptor, an oxoglutarate receptor, a P2Y receptor, a platelet-activating factor receptor, a prokineticin receptor, a prolactin-releasing peptide receptor, a prostanoid receptor, a proteinase-activated receptor, a QRFP receptor, a relaxin family peptide receptor, a somatostatin receptor, a succinate receptor, a tachykinin receptor, a thyrotropin-releasing hormone receptor, a trace amine receptor, a urotensin receptor, a vasopressin receptor, or a combination of two or more thereof.

[0165] In embodiments of the compounds described herein, the protein is a receptor protein. In embodiments, the receptor protein is a programmed death-ligand 1 (PD-L1) receptor, a programmed cell death protein 1 (PD-1) receptor, a 5-hydroxytryptamine receptor, an acetylcholine receptor, an adenosine receptor, an adenosine A2A receptor, an adenosine A2B receptor, an angiotensin receptor, an apelin receptor, a bile acid receptor, a bombesin receptor, a bradykinin receptor, a cannabinoid receptor, a chemerin receptor, a chemokine receptor, a cholecystokinin receptor, a Class A Orphan receptor, a dopamine receptor, an endothelin receptor, an epidermal growth factor receptor (EGFR), a formyl peptide receptor, a free fatty acid receptor, a galanin receptor, a ghrelin receptor, a glycoprotein hormone receptor, a gonadotrophin-releasing hormone receptor, a G protein-coupled receptor, a G protein-coupled estrogen receptor, a histamine receptor, a hydroxycarboxylic acid receptor, a kisspeptin receptor, a leukotriene receptor, a lysophospholipid receptor, a lysophospholipid S1P receptor, a melanin- concentrating hormone receptor, a melanocortin receptor, a melatonin receptor, a motilin receptor, a neuromedin U receptor, a neuropeptide FF / neuropeptide AF receptor, a neuropeptide S receptor, a neuropeptide W / neuropeptide B receptor, a neuropeptide Y receptor, a neurotensin receptor, an opioid receptor, an opsin receptor, an orexin receptor, an oxoglutarate receptor, aP2Y receptor, a platelet-activating factor receptor, a prokineticin receptor, a prolactin-releasing peptide receptor, a prostanoid receptor, a proteinase-activated receptor, a QRFP receptor, a relaxin family peptide receptor, a somatostatin receptor, a succinate receptor, a tachykinin receptor, a thyrotropin-releasing hormone receptor, a trace amine receptor, a urotensin receptor, a vasopressin receptor, or a combination of two or more thereof. In embodiments, the receptor protein is an integrin. In embodiments, the receptor protein is a somatostatin receptor. In embodiments, the receptor protein is a gonadotropin-releasing hormone receptor. In embodiments, the receptor protein is a bombesin receptor. In embodiments, the receptor protein is a vasoactive intestinal peptide receptor. In embodiments, the receptor protein is a neurotensin receptor. In embodiments, the receptor protein is a cholecystokinin 2 receptor. In embodiments, the receptor protein is a melanocortin receptor. In embodiments, the receptor protein is a ghrelin receptor.

[0166] In embodiments, the receptor protein is a PD-L1 receptor or a PD-1 receptor. In embodiments, the receptor protein is a PD-L1 receptor. In embodiments, the receptor protein is a PD-1 receptor.

[0167] In embodiments, the receptor protein is a receptor expressed on a cancer cell. In embodiments, the receptor protein is a receptor overexpressed on a cancer cell relative to a control.

[0168] In embodiments, the receptor protein is a G protein-coupled receptor. In embodiments, the receptor protein is a receptor tyrosine kinase. In embodiments, the receptor protein is a an ErbB receptor. In embodiments, the receptor protein is an epidermal growth factor receptor (EGFR). In embodiments, the receptor protein is epidermal growth factor receptor 1 (HER1). In embodiments, the receptor protein is epidermal growth factor receptor 2 (HER2). In embodiments, the receptor protein is epidermal growth factor receptor 3 (HER3). In embodiments, the receptor protein is epidermal growth factor receptor 4 (HER4).

[0169] In embodiments, the proteins comprise an unnatural amino acid within CDR-L1, CDR- L2, CDR-L3, CDR-H1, CDR-H2, or CDR-H3. In embodiments, the protein is an antigen- binding fragment, an antibody, or an antibody variant. In embodiments, the protein is an antigen-binding fragment. In embodiments, the protein is a single-chain variable fragment. In embodiments, the protein is an antibody. In embodiments, the protein has one unnatural amino acid within CDR-L1. In embodiments, the protein has one unnatural amino acid within CDR-L2. In embodiments, the protein has one unnatural amino acid within CDR-L3. In embodiments, the protein has one unnatural amino acid within CDR-H1. In embodiments, the protein has one unnatural amino acid within CDR-H2. In embodiments, the protein has one unnatural aminoacid within CDR-H3. In embodiments, the protein has two or more unnatural amino acids within CDR-L1, CDR-L2, CDR-L3, CDR-H1, CDR-H2, or CDR-H3. The two or more unnatural acids can be in the same or different CDR, and can be in the same or different chain (i.e., light or heavy).

[0170] In embodiments, the proteins described herein further comprise a detectable agent. In embodiments, the detectable agent is a radioisotope.

[0171] The proteins described herein are capable of forming protein conjugates. Provided herein are methods of binding a target protein on a cell or within a cell comprising contacting the cell with a protein described herein (e.g., peptidyl moiety), wherein the protein is capable of specifically binding to a target protein (e.g., peptidyl moiety) on the surface of the cell or within a cell, whereby the protein forms a covalent bond with the target protein, thereby forming a protein conjugate. In embodiments, the covalent bond is formed through a phorphorus fluoride exchange reaction. In embodiments, the covalent bond is formed through a proximity-enabled, phorphorus fluoride exchange (PFEx) reaction.

[0172] In embodiments, the proximity-enabled PFEx reaction occurs at an acidic pH. In embodiments, the acidic pH is from 5 to 6.9. In embodiments, the acidic pH is from 5.5 to 6.5.

[0173] “Contacting” is used in accordance with its plain ordinary meaning and refers to the process of allowing at least two distinct species (e.g., chemical compounds including proteins described herein and targets on a cell or within a cell) to become sufficiently proximal to react, interact or physically touch. It should be appreciated; however, the resulting reaction product can be produced directly from a reaction between the added reagents or from an intermediate from one or more of the added reagents that can be produced in the reaction mixture.

[0174] Conjugates

[0175] Provided herein are protein conjugates comprising a first peptidyl moiety linked to a second peptidyl moiety by a bioconjugate liker of the following formula: or wherein thefirst peptidyl moiety canbe a protein as described herein and the second peptidyl moiety can be a target protein as described herein.

[0176] Provided herein are protein conjugates of Formula (4) or a stereoisomer thereof: Wherein R4and R5as defined herein; L4is a bond, -(CH2)1-5-, - 1-5-, a bond, substituted or unsubstituted alkylene, or substituted or unsubstituted heteroalkylene; R1is hydrogen or an electron withdrawing group; and R2and R3are each independently substituted or unsubstituted C1-5alkyl. In embodiments, L4is a bond. In embodiments, L4is -(CH2)1-5-. In embodiments, L4is –O-. In embodiments, L4is -O-(CH2)1-5-. When L4is -O-(CH2)1-5-, the oxygen atom is adjacent to the phenyl ring and the alkylene group is adjacent to the phosphorous atom. In embodiments, L1is a bond or substituted or unsubstituted heteroalkylene. In embodiments, R2and R3are each independently unsubstituted C1-5alkyl. In embodiments, L2is a bond and L3is N N ,more detail below.

[0177] In embodiments, the protein conjugate of Formula (4) is a protein conjugate of Formula (4A): , wherein R4and R5as defined herein; x is an integer from 1 to 8; L1is a bond, substituted or unsubstituted alkylene, or substituted or unsubstituted heteroalkylene; R1is hydrogen or an electron withdrawing group; and R2and R3are each independently substituted or unsubstituted C1-5 alkyl. In embodiments, L1is a bond or substituted or unsubstituted heteroalkylene. In embodiments, R2and R3are each independently unsubstituted C1-5 alkyl. In embodiments, L2is a bond and L3isN N , more

[0178] In embodiments, the protein conjugate of Formula (4) is a protein conjugate of Formula (4B): , wherein R4and R5as defined herein; xis an integer from a or or substituted or unsubstituted heteroalkylene; R1is hydrogen or an electron withdrawing group; and R2and R3are each independently substituted or unsubstituted C1-5alkyl. In embodiments, L1is a bond or substituted or unsubstituted heteroalkylene. In embodiments, R2and R3are each independently unsubstituted C1-5alkyl. In embodiments, L2is a bond and L3is N N ,more detail below.

[0179] In embodiments, the protein conjugate of Formula (4) is a protein conjugate of Formula (4C): , wherein R4anddefined herein; x is an integer from 1 to 8; L1is a bond, substituted or unsubstituted alkylene, or substituted or unsubstituted heteroalkylene; and R1is hydrogen or an electron withdrawing group. In embodiments, L1is a bond or substituted or unsubstituted heteroalkylene. In embodiments, L2is a bond and L3isN N , more

[0180] In embodiments, the protein conjugate of Formula (4) is a protein conjugate of Formula (4D): , wherein R4and as defined herein; xis an integer from a or or substituted or unsubstituted heteroalkylene; and R2and R3are each independently substituted or unsubstituted C1-5alkyl. In embodiments, L1is a bond or substituted or unsubstituted heteroalkylene. In embodiments, R2and R3are each independently unsubstituted C1-5 alkyl. In embodiments, L2is a bond and L3is N N ,more detail below.

[0181] In embodiments, the protein conjugate of Formula (4) is a protein conjugate of Formula (4A): O N(CH3)2, wherein R4andas defined herein; x is an integer from 1 to 8; and L1is a bond, substituted or unsubstituted alkylene, or substituted or unsubstituted heteroalkylene. In embodiments, L1is a bond or substituted or unsubstituted heteroalkylene. In embodiments, L2is a bond and L3isN N , more

[0182] In embodiments, the protein conjugate of Formula (4) is a protein conjugate of Formula (4F): , wherein R4and R5are L3are as definedherein. In a N N , moredetail below.

[0183] In embodiments, the protein conjugate of Formula (4) is a protein conjugate of Formula (4G): , wherein R4defined herein. In embodiments, L2is a bond and L3is N ,more detail below.

[0184] Provided herein are protein conjugates of Formula (8) or a stereoisomer thereof:8), wherein R4and R5e as defined herein; x is an integer from 1 to 8; x1 is an integer from 0 to 5; L is a bond, substituted or unsubstituted alkylene, or substituted or unsubstituted heteroalkylene; R1is hydrogen or an electron withdrawing group; and R2is substituted or unsubstituted C1-5 alkyl. In embodiments, x1 is 0. In embodiments, x1 is 1. In embodiments, x1 is 2. In embodiments, x1 is 3. In embodiments, x1 is 4. In embodiments, x1 is 5. In embodiments, L1is a bond or substituted or unsubstituted heteroalkylene. In embodiments, R2is unsubstituted C1-5alkyl. In embodiments, L2is a bond and L3is N N , moredetail below.

[0185] In embodiments, the protein conjugate of Formula (8) is a protein conjugate of Formula (8A): , wherein R4and R5as defined herein; x is an integer from 1 to 8; x1 is an integer from 0 to 5; L1is a bond, substituted or unsubstituted alkylene, or substituted or unsubstituted heteroalkylene; and R2is substituted or unsubstituted C1-5 alkyl. In embodiments, x1 is 0. In embodiments, x1 is 1. In embodiments, x1 is 2. In embodiments, x1 is 3. In embodiments, x1 is 4. In embodiments, x1 is 5. In embodiments, L1is a bond or substituted or unsubstituted heteroalkylene. In embodiments, R2is unsubstituted C1-5 alkyl. In embodiments, L2is a bond and L3is N ,more detail below.

[0186] In embodiments, the protein conjugate of Formula (8) is a protein conjugate of Formula (8B): , wherein R4and as defined herein; x is an integer fromor unsubstituted alkylene, or substituted or unsubstituted heteroalkylene; and R1is hydrogen or an electron withdrawing group. In embodiments, x1 is 0. In embodiments, x1 is 1. In embodiments, x1 is 2. In embodiments, x1 is 3. In embodiments, x1 is 4. In embodiments, x1 is 5. In embodiments, L1is a bond or substituted or unsubstituted heteroalkylene. In embodiments, L2is a bond and L3is N N , moredetail below.

[0187] In embodiments, the protein conjugate of Formula (8) is a protein conjugate of Formula (8C): , wherein R4andas defined herein; x is an integer from 1 to 8; x1 is an integer from 0 to 5; L1is a bond, substituted or unsubstituted alkylene, or substituted or unsubstituted heteroalkylene; and R1is hydrogen or an electron withdrawing group. In embodiments, x1 is 0. In embodiments, x1 is 1. In embodiments, x1 is 2. In embodiments, x1 is 3. In embodiments, x1 is 4. In embodiments, x1 is 5. In embodiments, L1is a bond or substituted or unsubstituted heteroalkylene. In embodiments, L2is a bond and L3is N ,more detail below.

[0188] In embodiments, the protein conjugate of Formula (8) is a protein conjugate ofFormula (8D): O O P L3-R5, wherein R4and R5are L3are as defined herein. InN N , more

[0189] In embodiments, the protein conjugate of Formula (8) is a protein conjugate of Formula (8E): , wherein R4defined herein / . In embodiments, L2is a bond and L3is N ,more detail below.

[0190] In the compounds of Formulae (4), (4A)-(4G), and embodiments thereof, and Formulae (8), (8A)-(8H), and embodiments thereof, L2is a bond, -NR2A-, -S-, -S(O)2-, -O-, -C(O)-, -C(O)O-, -OC(O)-, -N(R2A)C(O)-, -C(O)N(R2A)-, -NR2AC(O)NR2B-, -NR2AC(NH)NR2B-, -SO2N(R2A)-, -N(R2A)SO2-, -C(S)-, substituted or unsubstituted alkylene, substituted or unsubstituted heteroalkylene, substituted or unsubstituted cycloalkylene, substituted or unsubstituted heterocycloalkylene, substituted or unsubstituted arylene, substituted or unsubstituted heteroarylene, a combination of two thereof, or a combination of three thereof; and R2Aand R2Bare independently hydrogen, substituted or unsubstituted alkyl, substituted or unsubstituted heteroalkyl, substituted or unsubstituted cycloalkyl, substituted or unsubstitutedheterocycloalkyl, substituted or unsubstituted aryl, substituted or unsubstituted heteroaryl, a combination of two thereof, or a combination of three thereof. In embodiments, L2is a bond.

[0191] In the compounds of Formulae (4), (4A)-(4G), and embodiments thereof, and Formulae (8), (8A)-(8H), and embodiments thereof, L3is a bond, -N(R3A)-, -S-, -S(O)2-, -O-, -C(O)-, -C(O)O-, -OC(O)-, -N(R3A)C(O)-, -C(O)N(R3A)-, -NR3AC(O)NR3B-, -NR3AC(NH)NR3B-, -SO2N(R3A)-, -N(R3A)SO2-, -C(S)-, substituted or unsubstituted alkylene, substituted or unsubstituted heteroalkylene, substituted or unsubstituted cycloalkylene, substituted or unsubstituted heterocycloalkylene, substituted or unsubstituted arylene, substituted or unsubstituted heteroarylene, a combination of two thereof, or a combination of three thereof; and R3Aand R3Bare independently hydrogen, substituted or unsubstituted alkyl, substituted or unsubstituted heteroalkyl, substituted or unsubstituted cycloalkyl, substituted or unsubstituted heterocycloalkyl, substituted or unsubstituted aryl, substituted or unsubstituted heteroaryl, a combination of two thereof, or a combination of three thereof. In embodiments, L3is N N ,In embodiments, L3, wherein –NH- is adjacent to the phosphorous.In embodiments, L3, wherein the N heteroatom is adjacent to thephosphorous. In embodiments, L3, wherein the O heteroatom isadjacent to the phosphorous. In embodiments, L3, wherein the –S- is adjacent to the phosphorous.

[0192] In the compounds of Formulae (4), (4A)-(4G), and embodiments thereof, and Formulae (8), (8A)-(8H), and embodiments thereof, L2is a bond and L3is N ,In embodiments, L2is a bond and L3, wherein –NH- is adjacent toN N the phosphorous. In embodiments, L2is a bond and L3, wherein the heteroatom is adjacent to the phosphorous. In L3is, wherein the heteroatom is adjacent to the phosphorous. Inand L3is , wherein the –S- is adjacent to the phosphorous.

[0193] Substituents

[0194] With reference to the compounds described herein, L4is a bond, -(CH2)1-5-, -O-(CH2)1-5-, or –O-. In embodiments, L4is -(CH2)1-5-, -O-(CH2)1-5-, or –O-. In embodiments, L4is -O-(CH2)1-5- or –O-. In embodiments, L4is–O-. In embodiments, L4is a bond. In embodiments, L4is -(CH2)1-5-. In embodiments, L4is -(CH2)1-4-. In embodiments, L4is -(CH2)1-3-. In embodiments, L4is -(CH2)1-2. In embodiments, L4is -CH2-. In embodiments, L4is -(CH2)2-. In embodiments, L4is -(CH2)3-. In embodiments, L4is -(CH2)4-. In embodiments, L4is -(CH2)5-. In embodiments, L4is -O-(CH2)1-5-. In embodiments, L4is –O-(CH2)1-4-. In embodiments, L4is –O-(CH2)1-3-. In embodiments, L4is –O-(CH2)1-2-. In embodiments, L4is –O-CH2-. In embodiments, L4is –O-(CH2)2-. In embodiments, L4is –O-(CH2)3-. In embodiments, L4is –O-(CH2)4-. In embodiments, L4is –O-(CH2)5-. In embodiments when L4is -O-(CH2)1-5-, the oxygen atom is adjacent the phenyl ring.

[0195] With reference to the compounds described herein, x is an integer from 0 to 8. In embodiments, x is an integer from 1 to 8. In embodiments, x is an integer from 1 to 7. In embodiments, x is an integer from 1 to 6. In embodiments, x is an integer from 1 to 5. In embodiments, x is an integer from 1 to 4. In embodiments, x is an integer from 1 to 3. In embodiments, x is an integer of 1 or 2. In embodiments, x is 1. In embodiments, x is 2. In embodiments, x is 3. In embodiments, x is 4. In embodiments, x is 5. In embodiments, x is 6. In embodiments, x is 7. In embodiments, x is 8. In embodiments, x is 0.

[0196] With reference to the compounds described herein, R2and R3are each independently substituted or unsubstituted C1-5alkyl group. In embodiments, R2and R3are each independently unsubstituted C1-5 alkyl group. In embodiments, R2and R3are each independently unsubstituted C1-4alkyl group. In embodiments, R2and R3are each independently unsubstituted C1-3alkyl group. In embodiments, R2and R3are each independently unsubstituted C1-2 alkyl group. In embodiments, R2and R3are methyl. In embodiments, R2and R3are ethyl. In embodiments, R2and R3are propyl. In embodiments, R2and R3are isopropyl. In embodiments, R2and R3are butyl. In embodiments, R2and R3are isobutyl. In embodiments where R2and R3are substituted C1-5alkyl group, the substituents are one or more of halogen, -CF3, -CBr3, -CCl3, -CI3, -CHF2, -CHBr2, -CHCl2, -CHI2, -CH2F, -CH2Br, -CH2Cl, -CH2I, -OCF3, -OCBr3, -OCCl3, -OCI3, -OCHF2, -OCHBr2, -OCHCl2, -OCHI2, -OCH2F, -OCH2Br, -OCH2Cl, -OCH2I, -CN, -OH, -NH2, -COOH, -CONH2, -NO2, -SH, -SO3H, -SO4H, -SO2NH2, -NHNH2, -ONH2, -N(O)2, -NHSO2H, -NHC(O)NHNH2, -N(O)2, -NHSO2H, -NHC(O)H, -NHC(O)OH, -NHOH, -N3,unsubstituted C1-4alkyl, and unsubstituted 2 to 4 membered heteroalkyl.

[0197] With reference to the compounds described herein, R1is hydrogen or an electron withdrawing group. In embodiments, R1is hydrogen. In embodiments, R1is hydrogen, halogen, -CX13, -CHX12, -CH2X1, -OCX13, -OCH2X1, -OCHX12, -CN, -SOn1R1A, -SOv1NR1AR1B, -NHC(O)NR1AR1B, -N(O)m1, -NR1AR1B, -C(O)R1A, -C(O)-OR1A, -C(O)NR1AR1B, -OR1A, -NR1ASO2R1B, -NR1AC(O)R1B, -NR1AC(O)OR1B, -NR1AOR1B, -NR3+, substituted or unsubstituted alkyl, substituted or unsubstituted heteroalkyl, substituted or unsubstituted aryl, or substituted or unsubstituted heteroaryl. In embodiments, R1is hydrogen, halogen, -CX13, -CHX12, -CH2X1, -OCX13, -OCH2X1, -OCHX12, -CN, -SOn1R1A, -SOv1NR1AR1B, -NHC(O)NR1AR1B, -N(O)m1, -NR1AR1B, -C(O)R1A, -C(O)-OR1A, -C(O)NR1AR1B, -OR1A, -NR1ASO2R1B, -NR1AC(O)R1B, -NR1AC(O)OR1B, -NR1AOR1B, -NR3+, substituted or unsubstituted alkyl, or substituted or unsubstituted heteroalkyl. In embodiments, R1is halogen, -CX13, -CHX12, -CH2X1, -OCX13, -OCH2X1, -OCHX12, -CN, -SOn1R1A, -SOv1NR1AR1B, -NHC(O)NR1AR1B, -N(O)m1, -NR1AR1B, -C(O)R1A, -C(O)-OR1A, -C(O)NR1AR1B, -OR1A, -NR1ASO2R1B, -NR1AC(O)R1B, -NR1AC(O)OR1B, -NR1AOR1B, -NR3+, substituted or unsubstituted alkyl, or substituted or unsubstituted heteroalkyl. In embodiments, R1is halogen, -CX13, -CHX12, -CH2X1, -OCX13, -OCH2X1, -OCHX12, -CN, -SOn1R1A, -N(O)m1, -C(O)R1A, -C(O)-OR1A, -C(O)NR1AR1B, -OR1A, -NR1ASO2R1B, -NR1AC(O)R1B, -NR1AC(O)OR1B, or -NR1AOR1B. In embodiments, R1is halogen, -CX13, -CHX12, -CH2X1, -OCX13, -OCH2X1, -OCHX12, -CN, or -N(O)m1.

[0198] In embodiments, R1is an electron-donating group or an electron-withdrawing group.

[0199] In embodiments, R1is an electron-withdrawing group. In embodiments, the electron- withdrawing group is halogen, -CX13, -CHX12, -CH2X1, , -CN, -SOn1R1A, -SOv1NR1AR1B, -N(O)m1, -C(O)R1A, -C(O)OR1A, -C(O)NR1AR1B, -NR1AOR1B, -NR3+, substituted or unsubstituted alkyl, or substituted or unsubstituted heteroalkyl; wherein X1, R1A, R1B, n1, v1, and m1 are as defined herein. In embodiments, R1Aand R1Bare hydrogen.

[0200] In embodiments, R1is an electron-donating group. In embodiments, the electron-donating group is –Cl, -Br, -I, -CX23, -CHX22, -OCX13, -OCH2X1, -OCHX12, , -OCOR1A,-OC(O)R1A, -OC(O)NR1AR1B, -SR1A, -PR1AR1B-NHC(O)NR1AR1B, -NR1AR1B, -OR1A, -NR1ASO2R1B, -NR1AC(O)R1B, -NR1AC(O)OR1B, substituted or unsubstituted alkyl, substituted or unsubstituted heteroalkyl, substituted or unsubstituted aryl, or substituted or unsubstituted heteroaryl. In embodiments, the substituted or unsubstituted alkyl is substituted or unsubstituted alkene. In embodiments, the electron-donating group is unsubstituted alkene. In embodiments, the substituted or unsubstituted alkyl is substituted or unsubstituted alkyne. In embodiments, R1Aand R1Bare hydrogen. In embodiments, the electron-donating group is unsubstituted alkyne.

[0201] In embodiments, R1is substituted or unsubstituted heteroalkyl. In embodiments, R1is unsubstituted heteroalkyl. In embodiments, R1is unsubstituted 2 to 8 membered heteroalkyl. In embodiments, R1is unsubstituted 2 to 6 membered heteroalkyl. In embodiments, R1is unsubstituted 2 to 4 membered heteroalkyl. In embodiments, R1is –O(CH2)mCH3, and m is an integer from 0 to 6. In embodiments, R1is –O(CH2)mCH3, and m is an integer from 0 to 4. In embodiments, R1is –O(CH2)mCH3, and m is an integer from 0 to 3. In embodiments, R1is – O(CH2)mCH3, and m is an integer from 0 to 2. In embodiments, R1is –O(CH2)mCH3, and m is 0 or 1. In embodiments, R1is –OCH3. In embodiments, R1is –OCH2CH3, In embodiments, R1is – O(CH2)2CH3, In embodiments, R1is –O(CH2)3CH3.

[0202] In embodiments, R1is halogen. In embodiments, R1is fluorine, chlorine, bromine, or iodine. In embodiments, R1is fluorine, chlorine, or bromine. In embodiments, R1is fluorine or chlorine. In embodiments, R1is fluorine or bromine. In embodiments, R1is chlorine or bromine. In embodiments, R1is fluorine. In embodiments, R1is chlorine. In embodiments, R1is bromine. In embodiments, R1is iodine.

[0203] In embodiments, R1is -CX13, -CHX12, or -CH2X1, wherein X1is halogen. In embodiments, R1is -CH2X1. In embodiments, R1is -CHX12. In embodiments, R1is -CX13. In embodiments, R1is -CF3. In embodiments, R1is -CHF2. In embodiments, R1is -CH2F. In embodiments, R1is -CCl3. In embodiments, R1is -CHCl2. In embodiments, R1is -CH2Cl. In embodiments, R1is -CBr3. In embodiments, R1is -CHBr2. In embodiments, R1is -CH2Br. In embodiments, R1is –CN. In embodiments, R1is -N(O)m1. In embodiments, R1is -NO2. In embodiments, R1is -SOn1R1A. In embodiments, R1is -SO2H. In embodiments, R1is -SOv1NR1AR1B. In embodiments, R1is -SO2NH2. In embodiments, R1is -NR3+.

[0204] In embodiments, R1is an alkyl group substituted with an electron-withdrawing group. In embodiments, R1is a halogen-substituted alkyl group. In embodiments, –(CH2)wCX13, - (CH2)wCHX12, or -(CH2)wCH2X1, wherein w is an integer from 1 to 5, and X1is halogen. Inembodiments, w is 1. In embodiments, w is 2. In embodiments, w is 3. In embodiments, w is 4. In embodiments, w is 5.

[0205] With reference to the compounds described herein, R1is ortho, para, or meta to the -L4-P(=O)(-F)(-NR2R3) group. In embodiments, R1is para to the -L4-P(=O)(-F)(-NR2R3) group. In embodiments, R1is meta to the -L4-P(=O)(-F)(-NR2R3) group. In embodiments, R1is ortho to the -L4-P(=O)(-F)(-NR2R3) group.

[0206] With reference to the compounds described herein, R1Ais hydrogen, substituted or unsubstituted alkyl, or substituted or unsubstituted heteroalkyl. In embodiments, R1Ais hydrogen, unsubstituted alkyl, or unsubstituted heteroalkyl. In embodiments, R1Ais hydrogen, substituted or unsubstituted C1-4alkyl, or substituted or unsubstituted 2 to 4 membered heteroalkyl. In embodiments, R1Ais hydrogen, unsubstituted C1-4 alkyl, or unsubstituted 2 to 4 membered heteroalkyl. In embodiments, R1Ais hydrogen. In embodiments, R1Ais unsubstituted C1-4 alkyl. In embodiments, R1Ais unsubstituted 2 to 4 membered heteroalkyl. In embodiments, R1Ais hydrogen and R1Bis hydrogen.

[0207] With reference to the compounds described herein, R1Bis hydrogen, substituted or unsubstituted alkyl, or substituted or unsubstituted heteroalkyl. In embodiments, R1Bis hydrogen, unsubstituted alkyl, or unsubstituted heteroalkyl. In embodiments, R1Bis hydrogen, substituted or unsubstituted C1-4alkyl, or substituted or unsubstituted 2 to 4 membered heteroalkyl. In embodiments, R1Bis hydrogen, unsubstituted C1-4 alkyl, or unsubstituted 2 to 4 membered heteroalkyl. In embodiments, R1Bis hydrogen. In embodiments, R1Bis unsubstituted C1-4 alkyl. In embodiments, R1Bis unsubstituted 2 to 4 membered heteroalkyl. In embodiments, R1Ais hydrogen and R1Bis hydrogen.

[0208] With reference to the compounds described herein, X1is independently –F, -Cl, -Br, or –I. In embodiments, X1is independently –F, -Cl, or -Br. In embodiments, X1is independently – F or -Cl. In embodiments, X1is –F. In embodiments, X1is -Cl. In embodiments, X1is -Br. In embodiments, X1is –I.

[0209] With reference to the compounds described herein, n1 is an integer from 0 to 4. In embodiments n1 is an integer from 0 to 3. In embodiments n1 is an integer from 0 to 2. In embodiments n1 is 0. In embodiments n1 is 1. In embodiments n1 is 2. In embodiments n1 is 3. In embodiments n1 is 4.

[0210] With reference to the compounds described herein, m1 is 1 or 2. In embodiments, m1 is 1. In embodiments, m1 is 2.

[0211] With reference to the compounds described herein, v1 is 1 or 2. In embodiments, v1 is 1. In embodiments, v1 is 2.

[0212] With reference to the compounds described herein, L1is a bond, substituted or unsubstituted alkylene, or substituted or unsubstituted heteroalkylene. In embodiments, L1is a bond. In embodiments, L1is substituted or unsubstituted alkylene. In embodiments, L1is substituted or unsubstituted C1-6 alkylene. In embodiments, L1is substituted or unsubstituted C1-4alkylene. In embodiments, L1is unsubstituted alkylene. In embodiments, L1is unsubstituted C1-6 alkylene. In embodiments, L1is unsubstituted C1-4 alkylene. In embodiments, L1is methylene. In embodiments, L1is ethylene. In embodiments, L1is propylene. In embodiments, L1is substituted or unsubstituted heteroalkylene. In embodiments, L1is substituted or unsubstituted 2 to 8 membered heteroalkylene. In embodiments, L1is substituted or unsubstituted 2 to 6 membered heteroalkylene.

[0213] In embodiments, L1is –NH-C(O)-(CH2)y- or –NH-C(O)-O-(CH2)y-, and y is an integer from 0 to 6. In embodiments, y is an integer from 0 to 5. In embodiments, y is an integer from 0 to 4. In embodiments, y is an integer from 0 to 3. In embodiments, y is an integer from 0 to 2. In embodiments, y is an integer from 0 to 3. In embodiments, L1is –NH-C(O)-. In embodiments, L1is –NH-C(O)-(CH2)- In embodiments, L1is –NH-C(O)-(CH2)2-. In embodiments, L1is –NH- C(O)-(CH2)3-. In embodiments, L1is –NH-C(O)-O-(CH2)y-, and y is an integer from 0 to 3. In embodiments, L1is –NH-C(O)-O-. In embodiments, L1is –NH-C(O)-O-(CH2)-. In embodiments, L1is –NH-C(O)-O-(CH2)2-. In embodiments, L1is –NH-C(O)-O-(CH2)3-. In embodiments, the –NH- is adjacent the phenyl ring.

[0214] In embodiments, L1is –C(O)NH-(CH2)y- or –C(O)-O-NH-(CH2)y-, and y is an integer from 0 to 6, wherein the –C(O)- is adjacent the phenyl ring. In embodiments, y is an integer from 0 to 5. In embodiments, y is an integer from 0 to 4. In embodiments, y is an integer from 0 to 3. In embodiments, y is an integer from 0 to 2. In embodiments, y is an integer from 0 to 3. In embodiments, L1is -C(O)NH-. In embodiments, L1is –C(O)NH-(CH2)- In embodiments, L1is – C(O)-NH-(CH2)2-. In embodiments, L1is –C(O)-NH-(CH2)3-. In embodiments, L1is -C(O)-O- NH-(CH2)y-, and y is an integer from 0 to 3. In embodiments, L1is –C(O)-O-NH-. In embodiments, L1is -C(O)-O-NH-(CH2)-. In embodiments, L1is -C(O)-O-NH-(CH2)2-. In embodiments, L1is -C(O)-O-NH(CH2)3-. In embodiments, the –C(O)- is adjacent the phenyl ring.

[0215] With reference to the compounds described herein, L2is a bond, -NR2A-, -S-, -S(O)2-, -O-, -C(O)-, -C(O)O-, -OC(O)-, -N(R2A)C(O)-, -C(O)N(R2A)-, -NR2AC(O)NR2B-, -NR2AC(NH)NR2B-, -SO2N(R2A)-, -N(R2A)SO2-, -C(S)-, substituted or unsubstituted alkylene, substituted or unsubstituted heteroalkylene, substituted or unsubstituted cycloalkylene, substituted or unsubstituted heterocycloalkylene, substituted or unsubstituted arylene, orsubstituted or unsubstituted heteroarylene. In embodiments, L2is a bond, -NH-, -S-, -S(O)2-, -O-, -C(O)-, -C(O)O-, -OC(O)-, -NHC(O)-, -C(O)NH-, -NHC(O)NH-, -NHC(NH)NH-, -SO2NH-, -NHSO2-, -C(S)-, L12-substituted or unsubstituted alkylene, L12-substituted or unsubstituted heteroalkylene, L12-substituted or unsubstituted cycloalkylene, L12-substituted or unsubstituted heterocycloalkylene, L12-substituted or unsubstituted arylene, or L12-substituted or unsubstituted heteroarylene. In embodiments, L2is a bond, -NH-, -S-, -S(O)2-, -O-, -C(O)-, -C(O)O-, -OC(O)-, -NHC(O)-, -C(O)NH-, -NHC(O)NH-, -NHC(NH)NH-, -SO2NH-, -NHSO2-, -C(S)-, unsubstituted alkylene, unsubstituted heteroalkylene, unsubstituted cycloalkylene, unsubstituted heterocycloalkylene, unsubstituted arylene, or unsubstituted heteroarylene. In embodiments, L2is a bond. In embodiments, the alkylene is a C1-6 alkylene. In embodiments, the alkylene is a C1-4alkylene. In embodiments, the heteroalkylene is a 2 to 6 membered heteroalkylene. In embodiments, the heteroalkylene is a 2 to 4 membered heteroalkylene. In embodiments, the cycloalkylene is a C5-C6cycloalkylene. In embodiments, the heterocycloalkylene is a 5 or 6 membered heterocycloalkylene. In embodiments, the arylene is a C5-6arylene. In embodiments, the heteroarylene is a 5 or 6 membered heteroarylene.

[0216] With reference to the compounds described herein, R2Aand R2Bare independently hydrogen, substituted or unsubstituted alkyl, substituted or unsubstituted heteroalkyl, substituted or unsubstituted cycloalkyl, substituted or unsubstituted heterocycloalkyl, substituted or unsubstituted aryl, or substituted or unsubstituted heteroaryl. In embodiments, the alkylene is a C1-4alkylene. In embodiments, the heteroalkylene is a 2 to 6 membered heteroalkylene. In embodiments, the heteroalkylene is a 2 to 4 membered heteroalkylene. In embodiments, the cycloalkylene is a C5-C6cycloalkylene. In embodiments, the heterocycloalkylene is a 5 or 6 membered heterocycloalkylene. In embodiments, the arylene is a C5-6 arylene. In embodiments, the heteroarylene is a 5 or 6 membered heteroarylene. In embodiments, R2Aand R2Bare hydrogen.

[0217] With reference to the compounds described herein, L12is halogen, -CF3, -CBr3, -CCl3, -CI3, -CHF2, -CHBr2, -CHCl2, -CHI2, -CH2F, -CH2Br, -CH2Cl, -CH2I, -OCF3, -OCBr3, -OCCl3, -OCI3, -OCHF2, -OCHBr2, -OCHCl2, -OCHI2, -OCH2F, -OCH2Br, -OCH2Cl, -OCH2I, -CN, -OH, -NH2, -COOH, -CONH2, -NO2, -SH, -SO3H, -SO4H, -SO2NH2, -NHNH2, -ONH2, -NHC(O)NHNH2, -N(O)2, -NHSO2H, -NHC(O)H, -NHC(O)OH, -NHOH, -N3, unsubstituted alkyl, unsubstituted heteroalkyl, unsubstituted cycloalkyl, unsubstituted heterocycloalkyl, unsubstituted aryl, or unsubstituted heteroaryl. In embodiments, the alkylene is a C1-4alkylene. In embodiments, the heteroalkylene is a 2 to 6 membered heteroalkylene. In embodiments, the heteroalkylene is a 2 to 4 membered heteroalkylene. In embodiments, the cycloalkylene is a C5-C6cycloalkylene. In embodiments, the heterocycloalkylene is a 5 or 6 membered heterocycloalkylene. In embodiments, the arylene is a C5-6 arylene. In embodiments, the heteroarylene is a 5 or 6 membered heteroarylene.

[0218] With reference to the compounds described herein, L3is a bond, -N(R3A)-, -S-, -S(O)2-, -O-, -C(O)-, -C(O)O-, -OC(O)-, -N(R3A)C(O)-, -C(O)N(R3A)-, -NR3AC(O)NR3B-, -C(S)-, -NR3AC(NH)NR3B-, -SO2N(R3A)-, -N(R3A)SO2-, substituted or unsubstituted alkylene, substituted or unsubstituted heteroalkylene, substituted or unsubstituted cycloalkylene, substituted or unsubstituted heterocycloalkylene, substituted or unsubstituted arylene, substituted or unsubstituted heteroarylene, a combination of two thereof, or a combination of three thereof. In embodiments, L3is a bond, -NH-, -S-, -S(O)2-, -O-, -C(O)-, -C(O)O-, -OC(O)-, -NHC(O)-, -C(O)NH-, -NHC(O)NH-, -NHC(NH)NH-, -SO2NH-, -NHSO2-, -C(S)-, L13-substituted or unsubstituted alkylene, L13-substituted or unsubstituted heteroalkylene, L13- substituted or unsubstituted cycloalkylene, L13-substituted or unsubstituted heterocycloalkylene, L13-substituted or unsubstituted arylene, L13-substituted or unsubstituted heteroarylene, a combination of two thereof, or a combination of three thereof. In embodiments, the alkylene is a C1-4 alkylene. In embodiments, the heteroalkylene is a 2 to 6 membered heteroalkylene. In embodiments, the heteroalkylene is a 2 to 4 membered heteroalkylene. In embodiments, the cycloalkylene is a C5-C6 cycloalkylene. In embodiments, the heterocycloalkylene is a 5 or 6 membered heterocycloalkylene. In embodiments, the arylene is a C5-6 arylene. In embodiments, the heteroarylene is a 5 or 6 membered heteroarylene.

[0219] In embodiments where L3is a combination of two thereof or a combination of three thereof, at least one of the combination is C5-C6cycloalkylene, 5 or 6 membered heterocycloalkylene, C5-6 arylene, or 5 or 6 membered heteroarylene. In embodiments where L3is a combination of two thereof or a combination of three thereof, then only one of the combination is C5-C6 cycloalkylene, 5 or 6 membered heterocycloalkylene, C5-6 arylene, or 5 or 6 membered heteroarylene.

[0220] With reference to the compounds described herein, R3Aand R3Bare independently hydrogen, substituted or unsubstituted alkyl, substituted or unsubstituted heteroalkyl, substituted or unsubstituted cycloalkyl, substituted or unsubstituted heterocycloalkyl, substituted or unsubstituted aryl, or substituted or unsubstituted heteroaryl. In embodiments, the alkylene is a C1-4 alkylene. In embodiments, the heteroalkylene is a 2 to 6 membered heteroalkylene. In embodiments, the heteroalkylene is a 2 to 4 membered heteroalkylene. In embodiments, the cycloalkylene is a C5-C6 cycloalkylene. In embodiments, the heterocycloalkylene is a 5 or 6 membered heterocycloalkylene. In embodiments, the arylene is a C5-6 arylene. In embodiments,the heteroarylene is a 5 or 6 membered heteroarylene.

[0221] With reference to the compounds described herein, L13is halogen, -CF3, -CBr3, -CCl3, -CI3, -CHF2, -CHBr2, -CHCl2, -CHI2, -CH2F, -CH2Br, -CH2Cl, -CH2I, -OCF3, -OCBr3, -OCCl3, -OCI3, -OCHF2, -OCHBr2, -OCHCl2, -OCHI2, -OCH2F, -OCH2Br, -OCH2Cl, -OCH2I, -CN, -OH, -NH2, -COOH, -CONH2, -NO2, -SH, -SO3H, -SO4H, -SO2NH2, -NHNH2, -ONH2, -NHC(O)NHNH2, -N(O)2, -NHSO2H, -NHC(O)H, -NHC(O)OH, -NHOH, -N3, unsubstituted alkyl, unsubstituted heteroalkyl, unsubstituted cycloalkyl, unsubstituted heterocycloalkyl, unsubstituted aryl, or unsubstituted heteroaryl. In embodiments, the alkylene is a C1-4 alkylene. In embodiments, the heteroalkylene is a 2 to 6 membered heteroalkylene. In embodiments, the heteroalkylene is a 2 to 4 membered heteroalkylene. In embodiments, the cycloalkylene is a C5- C6cycloalkylene. In embodiments, the heterocycloalkylene is a 5 or 6 membered heterocycloalkylene. In embodiments, the arylene is a C5-6 arylene. In embodiments, the heteroarylene is a 5 or 6 membered heteroarylene.

[0222] With reference to the compounds of Formula (3) (3A)-(3G), (7), (7A)-(7E) and embodiments thereof, W is –H, a peptidyl moiety, or an amino acid moiety; and Y is –OH, a peptidyl moiety, or an amino acid moiety. In embodiments, W is –H and Y is –OH. In embodiments, W is not –H when Y is –OH. In embodiments, W is a peptidyl moiety and Y is peptidyl moiety. In embodiments, W is a an amino acid moiety and Y is an amino acid moiety. In embodiments, W is a peptidyl moiety or an amino acid moiety and Y is a peptidyl moiety or an amino acid moiety. In embodiments, W is a peptidyl moiety and Y is –OH, a peptidyl moiety, or an amino acid moiety. In embodiments, W is –H, a peptidyl moiety, or an amino acid moiety; and Y is a peptidyl moiety.

[0223] In embodiments of the compounds described herein, the peptidyl moiety of R5is a target protein. In embodiments, the peptidyl moiety of R5is a receptor protein, a cytosolic protein, a transcriptional factor, or an enzyme. In embodiments, the peptidyl moiety of R5is a receptor protein. In embodiments, the receptor protein is an extracellular domain receptor protein, a transmembrane domain receptor protein, or an intracellular domain receptor protein. In embodiments, the peptidyl moiety of R5is a cytosolic protein. In embodiments, the peptidyl moiety of R5is a transcriptional factor. In embodiments, the peptidyl moiety of R5is an enzyme.

[0224] In embodiments of the compounds described herein, the peptidyl moiety of R4comprises an antibody or an antibody variant; and the peptidyl moiety of R5comprises a target protein. In embodiments, the peptidyl moiety of R4comprises an antibody or an antibody variant; and the peptidyl moiety of R5comprises a target protein, wherein the target protein comprises a lysine, histidine, tyrosine, or cysteine bonded to L3, where L3is a bond. Inembodiments, R4comprises an antibody. In embodiments, R4comprises an antibody variant. In embodiments, the antibody variant is a variant as defined herein. In embodiments, the antibody variant is a single-chain variable fragment, a single-domain antibody, an affibody, or an antigen- binding fragment. In embodiments, the antibody variant is a single-chain variable fragment. In embodiments, the antibody variant is a single-domain antibody. In embodiments, the antibody variant is an affibody. In embodiments, the antibody variant is an antigen-binding fragment. In embodiments, the target protein is a receptor protein, a cytosolic protein, a transcriptional factor, or an enzyme. In embodiments, R5is a receptor protein. In embodiments, R5is an extracellular domain receptor protein. In embodiments, R5is a transmembrane domain receptor protein. In embodiments, R5is an intracellular domain receptor protein. In embodiments, R5is a cytosolic protein. In embodiments, R5is a transcriptional factor. In embodiments, R5is an enzyme.

[0225] In embodiments of the compounds described herein, the peptidyl moiety of R4comprises an antibody or an antibody variant; and the peptidyl moiety of R5comprises a receptor protein. In embodiments, the peptidyl moiety of R4comprises an antibody or an antibody variant; and the peptidyl moiety of R5comprises a receptor protein, wherein the receptor protein comprises a lysine, histidine, tyrosine, or cysteine bonded to L3, where L3is a bond. In embodiments, R4comprises an antibody. In embodiments, R4comprises an antibody variant. In embodiments, the antibody variant is a variant as defined herein. In embodiments, the antibody variant is a single-chain variable fragment, a single-domain antibody, an affibody, or an antigen-binding fragment. In embodiments, the antibody variant is a single-chain variable fragment. In embodiments, the antibody variant is a single-domain antibody. In embodiments, the antibody variant is an affibody. In embodiments, the antibody variant is an antigen-binding fragment. In embodiments, the receptor protein is any receptor protein described herein.

[0226] In embodiments of the compounds described herein, the peptidyl moiety of R4comprises a receptor protein; and the peptidyl moiety of R5comprises an antibody or an antibody variant. In embodiments, the peptidyl moiety of R4comprises a receptor protein; and the peptidyl moiety of R5comprises an antibody or an antibody variant; wherein the antibody or antibody variant comprises a lysine, histidine, tyrosine or cysteine bonded to L3, where L3is a bond. In embodiments, R5comprises an antibody. In embodiments, R5comprises an antibody variant. In embodiments, the antibody variant is a variant as defined herein. In embodiments, the antibody variant is a single-chain variable fragment, a single-domain antibody, an affibody, or an antigen-binding fragment. In embodiments, the antibody variant is a single-chain variable fragment. In embodiments, the antibody variant is a single-domain antibody. In embodiments, the antibody variant is an affibody. In embodiments, the antibody variant is an antigen-bindingfragment. In embodiments, the receptor protein is any receptor protein described herein.

[0227] In embodiments, the biomolecules, proteins, and peptidyl moieties described herein comprise a receptor protein. In embodiments, the receptor protein is a 5-hydroxytryptamine receptor, an acetylcholine receptor, an adenosine receptor, an adenosine A2A receptor, an adenosine A2B receptor, an angiotensin receptor, an apelin receptor, a bile acid receptor, a bombesin receptor, a bradykinin receptor, a cannabinoid receptor, a chemerin receptor, a chemokine receptor, a cholecystokinin receptor, a Class A Orphan receptor, a dopamine receptor, an endothelin receptor, an epidermal growth factor receptor (EGFR), a formyl peptide receptor, a free fatty acid receptor, a galanin receptor, a ghrelin receptor, a glycoprotein hormone receptor, a gonadotrophin-releasing hormone receptor, a G protein-coupled receptor, a G protein-coupled estrogen receptor, a histamine receptor, a hydroxycarboxylic acid receptor, a kisspeptin receptor, a leukotriene receptor, a lysophospholipid receptor, a lysophospholipid S1P receptor, a melanin-concentrating hormone receptor, a melanocortin receptor, a melatonin receptor, a motilin receptor, a neuromedin U receptor, a neuropeptide FF / neuropeptide AF receptor, a neuropeptide S receptor, a neuropeptide W / neuropeptide B receptor, a neuropeptide Y receptor, a neurotensin receptor, an opioid receptor, an opsin receptor, an orexin receptor, an oxoglutarate receptor, a P2Y receptor, a platelet-activating factor receptor, a prokineticin receptor, a prolactin-releasing peptide receptor, a prostanoid receptor, a proteinase-activated receptor, a QRFP receptor, a relaxin family peptide receptor, a somatostatin receptor, a succinate receptor, a tachykinin receptor, a thyrotropin-releasing hormone receptor, a trace amine receptor, a urotensin receptor, a vasopressin receptor, or a combination of two or more thereof. In embodiments, the receptor protein is an integrin. In embodiments, the receptor protein is a somatostatin receptor. In embodiments, the receptor protein is a gonadotropin-releasing hormone receptor. In embodiments, the receptor protein is a bombesin receptor. In embodiments, the receptor protein is a vasoactive intestinal peptide receptor. In embodiments, the receptor protein is a neurotensin receptor. In embodiments, the receptor protein is a cholecystokinin 2 receptor. In embodiments, the receptor protein is a melanocortin receptor. In embodiments, the receptor protein is a ghrelin receptor.

[0228] In embodiments, the receptor protein is a receptor expressed on a cancer cell. In embodiments, the receptor protein is a receptor overexpressed on a cancer cell relative to a control.

[0229] In embodiments, the receptor protein is a G protein-coupled receptor. In embodiments, the receptor protein is a receptor tyrosine kinase. In embodiments, the receptor protein is a an ErbB receptor. In embodiments, the receptor protein is an epidermal growth factor receptor(EGFR). In embodiments, the receptor protein is epidermal growth factor receptor 1 (HER1). In embodiments, the receptor protein is epidermal growth factor receptor 2 (HER2). In embodiments, the receptor protein is epidermal growth factor receptor 3 (HER3). In embodiments, the receptor protein is epidermal growth factor receptor 4 (HER4).

[0230] Cellular Compositions

[0231] The disclosure provides cells comprising the compounds, compositions and complexes provided herein, including embodiments thereof. In embodiments, the cell comprises the compound of Formula (1), including any embodiment thereof. In embodiments, the cell comprises the compound of Formula (1A), including any embodiment thereof. In embodiments, the cell comprises the compound of Formula (1B), including any embodiment thereof. In embodiments, the cell comprises the compound of Formula (1C), including any embodiment thereof. In embodiments, the cell comprises the compound of Formula (1D), including any embodiment thereof. In embodiments, the cell comprises the compound of Formula (1E), including any embodiment thereof. In embodiments, the cell comprises PFY. In embodiments, the cell comprises PFK. In embodiments, the cell comprises the compound of Formula (5), including any embodiment thereof. In embodiments, the cell comprises the compound of Formula (5A), including any embodiment thereof. In embodiments, the cell comprises the compound of Formula (5B), including any embodiment thereof. In embodiments, the cell comprises the compound of Formula (5C), including any embodiment thereof. In embodiments, the cell comprises the compound of Formula (5D), including any embodiment thereof. In embodiments, the cell comprises the compound of Formula (5E), including any embodiment thereof. In embodiments, the cell further includes a mutant pyrrolysyl-tRNA synthetase as described herein, including embodiments thereof. In embodiments, the cell further includes a vector as described herein, including embodiments thereof. In embodiments, the cell further includes a tRNAPyl.

[0232] In embodiments, the compound of Formula (1) (including embodiments (1A)-(1E), PFY, PFK) is biosynthesized inside the cell, thereby generating a cell containing the compound of Formula (IV). In embodiments, the compound of Formula (1) is contained in the medium outside the cell and penetrates into the cell, thereby generating a cell containing the compound of Formula (1). In embodiments, the cell comprises the compound of Formula (1). In embodiments, the cell comprises the compound of Formula (1) that is synthesized inside the cell. In embodiments, the cell comprises the compound of Formula (1) that is synthesized outside a cell, and that penetrates into the cell.

[0233] In embodiments, the compound of Formula (2) (including embodiments (2A)-(2G)) isbiosynthesized inside the cell, thereby generating a cell containing the compound of Formula (2). In embodiments, the compound of Formula (2) is contained in the medium outside the cell and penetrates into the cell, thereby generating a cell containing the compound of Formula (2). In embodiments, the cell comprises the compound of Formula (2). In embodiments, the cell comprises the compound of Formula (2) that is synthesized inside the cell. In embodiments, the cell comprises the compound of Formula (2) that is synthesized outside a cell, and that penetrates into the cell.

[0234] In embodiments, the compound of Formula (3) (including embodiments (3A)-(3G)) is biosynthesized inside the cell, thereby generating a cell containing the compound of Formula (3). In embodiments, the compound of Formula (3) is contained in the medium outside the cell and penetrates into the cell, thereby generating a cell containing the compound of Formula (3). In embodiments, the cell comprises the compound of Formula (3). In embodiments, the cell comprises the compound of Formula (3) that is synthesized inside the cell. In embodiments, the cell comprises the compound of Formula (3) that is synthesized outside a cell, and that penetrates into the cell.

[0235] In embodiments, the compound of Formula (5) (including embodiments (5A)-(5E)) is biosynthesized inside the cell, thereby generating a cell containing the compound of Formula (5). In embodiments, the compound of Formula (5) is contained in the medium outside the cell and penetrates into the cell, thereby generating a cell containing the compound of Formula (5). In embodiments, the cell comprises the compound of Formula (5). In embodiments, the cell comprises the compound of Formula (5) that is synthesized inside the cell. In embodiments, the cell comprises the compound of Formula (5) that is synthesized outside a cell, and that penetrates into the cell.

[0236] In embodiments, the compound of Formula (6) (including embodiments (6A)-(6E)) is biosynthesized inside the cell, thereby generating a cell containing the compound of Formula (6). In embodiments, the compound of Formula (6) is contained in the medium outside the cell and penetrates into the cell, thereby generating a cell containing the compound of Formula (6). In embodiments, the cell comprises the compound of Formula (6). In embodiments, the cell comprises the compound of Formula (6) that is synthesized inside the cell. In embodiments, the cell comprises the compound of Formula (6) that is synthesized outside a cell, and that penetrates into the cell.

[0237] In embodiments, the compound of Formula (7) (including embodiments (7A)-(7E)) is biosynthesized inside the cell, thereby generating a cell containing the compound of Formula (7). In embodiments, the compound of Formula (7) is contained in the medium outside the celland penetrates into the cell, thereby generating a cell containing the compound of Formula (7). In embodiments, the cell comprises the compound of Formula (7). In embodiments, the cell comprises the compound of Formula (7) that is synthesized inside the cell. In embodiments, the cell comprises the compound of Formula (7) that is synthesized outside a cell, and that penetrates into the cell.

[0238] In embodiments, the cell comprises the biomolecule conjugates described herein. In embodiments, the cell comprises biomolecule conjugate of Formula (4) (including embodiments (4A)-(4G)). In embodiments, the cell comprises biomolecule conjugate of Formula (8) (including embodiments (8A)-(8E)).

[0239] In embodiments, the disclosure provides a cell comprising the protein described herein. In aspects, the cell further includes a vector as described herein. In embodiments, the protein is biosynthesized inside the cell, thereby generating a cell containing the recombinant protein. In aspects, the protein is contained in the medium outside the cell and penetrates into the cell, thereby generating a cell containing the protein. In aspects, the cell comprises a recombinant protein that is synthesized inside the cell. In aspects, the cell comprises a protein that is synthesized outside a cell, and that penetrates into the cell. A cell can be any prokaryotic or eukaryotic cell. For example, any of the compounds (e.g., recombinant proteins) and compositions described herein can be expressed in bacterial cells such as E. coli, insect cells, yeast or mammalian cells (such as Hela cells, Chinese hamster ovary cells (CHO) or COS cells). In aspects, a cell can be a premature mammalian cell, i.e., pluripotent stem cell. In aspects, a cell can be derived from other human tissue. Other suitable cells are known to those skilled in the art.

[0240] The protein provided herein may be delivered to cells using methods well known in the art. Thus, in an aspect is provided a nucleic acid sequence encoding the protein described herein, including embodiments thereof. Thus, in an aspect is provided a vector including a nucleic acid sequence encoding the protein described herein, including embodiments thereof.

[0241] A cell can be any prokaryotic or eukaryotic cell. In aspects, the cell is prokaryotic. In aspects, the cell is eukaryotic. In aspects, the cell is a bacterial cell, a fungal cell, a plant cell, an archael cell, or an animal cell. In aspects, the animal cell is an insect cell or a mammalian cell. In aspects, the cell is a bacterial cell. In aspects, the cell is a fungal cell. In aspects, the cell is a plant cell. In aspects, the cell is an archael cell. In aspects, the cell is an animal cell. In aspects, the cell is an insect cell. In aspects, the cell is a mammalian cell. In aspects, the cell is a human cell. For example, any of the compositions described herein can be expressed in bacterial cells such as E. coli, insect cells, yeast or mammalian cells (such as Hela cells, Chinese hamster ovarycells (CHO) or COS cells). In aspects, the cell is a premature mammalian cell, i.e., a pluripotent stem cell. In aspects, the cell is derived from other human tissue. Other suitable cells are known to those skilled in the art.

[0242] Pyrrolysyl-tRNA Synthetase

[0243] As described herein, an unnatural amino acid (e.g., of Formula (1), Formula (5) and embodiments thereof) may be inserted into or replace a naturally occurring amino acid in a protein. In order for the unnatural amino acid to be inserted or replace an amino acid in a protein, it must be capable of being incorporated during proteinogenesis. Thus, the unnatural amino acid must be present on a transfer RNA molecule (tRNA) such that it may be used in translation. Loading of amino acids occurs via an aminoacyl-tRNA synthetase, which is an enzyme that facilitates the attachment of appropriate amino acids to tRNA molecules. However, the attachment of unnatural amino acids to tRNA may not necessarily be accomplished by the naturally occurring aminoacyl-tRNA synthetase. Engineered aminoacyl-tRNA synthetases (e.g., mutant pyrrolysyl-tRNA synthetase (PyIRS)) may be useful for attaching unnatural amino acids to tRNA. A PyIRS mutant library was generated. Mutant pyrrolysyl-tRNA synthetases and methods for making them are described, for example, in US 2021 / 0002325, WO 2020 / 072674, WO 2020 / 206341, and WO 2022 / 232377, the disclosures of which are incorporated by reference herein in their entirety.

[0244] In embodiments, the disclosure provides a pyrrolysyl-tRNA synthetases having at least 85% sequence identity to the amino acid sequence of SEQ ID NO:9. In embodiments, the disclosure provides a pyrrolysyl-tRNA synthetases having at least 90% sequence identity to the amino acid sequence of SEQ ID NO:9. In embodiments, the disclosure provides a pyrrolysyl- tRNA synthetases having at least 95% sequence identity to the amino acid sequence of SEQ ID NO:9. In embodiments, the disclosure provides a pyrrolysyl-tRNA synthetases comprising the amino acid sequence of SEQ ID NO:9. In embodiments, the disclosure provides a pyrrolysyl- tRNA synthetases as set forth in SEQ ID NO:9.

[0245] Vectors

[0246] The compositions (e.g., mutant pyrrolysyl-tRNA synthetase, tRNAPyl) provided hereinmay be delivered to cells using methods well known in the art. Thus, in an embodiment is provided a vector including a nucleic acid sequence encoding a mutant pyrrolysyl-tRNA synthetase as described herein, including embodiments thereof. In embodiments, the vector further includes a nucleic acid sequence encoding tRNAPyl. In embodiments, the vector comprises a nucleic acid sequence encoding a mutant pyrrolysyl-tRNA synthetase as described herein. In embodiments, the vector further includes a nucleic acid sequence encoding tRNAPyl.

[0247] Methods of Forming a Biomolecule or Biomolecule Conjugate

[0248] The compositions provided herein are useful for forming a protein comprising an unnatural amino acid (e.g., a compound of Formula (3) or embodiments thereof). In embodiments, the method of forming a protein comprising an unnatural amino acid comprises contacting a protein, a mutant pyrrolysyl-tRNA synthetase, a tRNAPyl, and a compound of Formula (1) (including embodiments thereof), thereby producing the protein comprising the unnatural amino acid of Formula (1) (including embodiments thereof). The protein produced by the method will comprise the unnatural amino acid side chain of Formula (2) (including embodiments thereof). The mutant pyrrolysyl-tRNA synthetase used in the method of producing the biomolecule is any described herein or known in the art (e.g., SEQ ID NO:9). The tRNAPylused in the method of producing the protein is any described herein. In embodiments, the reaction is performed in vitro. In embodiments, the reaction is performed in vivo. In embodiments, the reaction is performed in one or more living cells. In embodiments, the reaction is performed in one or more living bacterial cells. In embodiments, the reaction is performed in one or more living mammalian cells.

[0249] The compositions provided herein are useful for forming a protein comprising an unnatural amino acid (e.g., a compound of Formula (3) or embodiments thereof). In embodiments, the method of forming a protein of Formula (3) (including embodiments thereof) comprises contacting a protein, a mutant pyrrolysyl-tRNA synthetase, a tRNAPyl, and a compound of Formula (1) (including embodiments thereof), thereby producing the protein of Formula (3) (including embodiments thereof). The mutant pyrrolysyl-tRNA synthetase used in the method of producing the biomolecule is any described herein or known in the art (e.g., SEQ ID NO:9). The tRNAPylused in the method of producing the protein is any described herein. In embodiments, the reaction is performed in vitro. In embodiments, the reaction is performed in vivo. In embodiments, the reaction is performed in one or more living cells. In embodiments, the reaction is performed in one or more living bacterial cells. In embodiments, the reaction is performed in one or more living mammalian cells.

[0250] The compositions provided herein are useful for forming a protein comprising an unnatural amino acid (e.g., a compound of Formula (7) or embodiments thereof). In embodiments, the method of forming a protein comprising an unnatural amino acid comprises contacting a protein, a mutant pyrrolysyl-tRNA synthetase, a tRNAPyl, and a compound of Formula (5) (including embodiments thereof), thereby producing the protein comprising the unnatural amino acid of Formula (5) (including embodiments thereof). The protein produced by the method will comprise the unnatural amino acid side chain of Formula (6) (includingembodiments thereof). The mutant pyrrolysyl-tRNA synthetase used in the method of producing the biomolecule is any described herein or known in the art (e.g., SEQ ID NO:9). The tRNAPylused in the method of producing the protein is any described herein. In embodiments, the reaction is performed in vitro. In embodiments, the reaction is performed in vivo. In embodiments, the reaction is performed in one or more living cells. In embodiments, the reaction is performed in one or more living bacterial cells. In embodiments, the reaction is performed in one or more living mammalian cells.

[0251] The compositions provided herein are useful for forming a protein comprising an unnatural amino acid (e.g., a compound of Formula (7) or embodiments thereof). In embodiments, the method of forming a protein of Formula (7) (including embodiments thereof) comprises contacting a protein, a mutant pyrrolysyl-tRNA synthetase, a tRNAPyl, and a compound of Formula (5) (including embodiments thereof), thereby producing the protein of Formula (7) (including embodiments thereof). The mutant pyrrolysyl-tRNA synthetase used in the method of producing the biomolecule is any described herein or known in the art (e.g., SEQ ID NO:9). The tRNAPylused in the method of producing the protein is any described herein. In embodiments, the reaction is performed in vitro. In embodiments, the reaction is performed in vivo. In embodiments, the reaction is performed in one or more living cells. In embodiments, the reaction is performed in one or more living bacterial cells. In embodiments, the reaction is performed in one or more living mammalian cells.

[0252] Provided herein is a method of enhancing the bioreactivity and / or binding efficacy of a protein comprising: (i) mutating a first amino acid to an unnatural amino acid as described herein (including embodiments thereof), and (ii) mutating a second amino acid proximal to the first amino acid to arginine; thereby enhancing the bioreactivity and / or binding efficacy of the protein. In embodiments, the method of enhancing the bioreactivity and / or binding efficacy of a protein comprises: (i) mutating a first amino acid to an unnatural amino acid, and (ii) mutating a second amino acid proximal to the first amino acid to arginine; wherein the first amino acid is Arg, Ala, Ile, Leu, Met, Val, Phe, Trp, Tyr, Asn, Cys, Gln, Ser, Thr, Asp, Glu, His, Lys, Gly, or Pro; wherein the second amino acid is Ala, Ile, Leu, Met, Val, Phe, Trp, Tyr, Asn, Cys, Gln, Ser, Thr, Asp, Glu, His, Lys, Gly, or Pro; thereby enhancing the bioreactivity and / or binding efficacy of the protein. Mutating a second amino acid to arginine results in the “non-naturally occurring arginine.” The protein can be any protein described herein, including a protein having the amino acid side chain of Formula (2) or any embodiment thereof, a protein of Formula (3) or any embodiment thereof, a protein having the amino acid side chain of Formula (6) or any embodiment thereof, a protein of Formula (7) or any embodiment thereof. In embodiments, theprotein is an antibody or an antibody variant. In embodiments, the antibody variant is a single- chain variable fragment, a single-domain antibody, an affibody, or an antigen-binding fragment. In embodiments, the protein is a single-chain variable fragment. In embodiments, the protein is a single-domain antibody. In embodiments, the protein is an affibody. In embodiments, the protein is an antigen-binding fragment. In embodiments, the method further comprises contacting the protein with a target protein, thereby covalently bonding the protein to the target protein.

[0253] Provided herein is a method of enhancing the bioreactivity and / or binding efficacy of a protein comprising an unnatural amino acid, the method comprising mutating an amino acid proximal to the unnatural amino acid to arginine; wherein the unnatural amino acid is any unnatural amino acid described herein (including embodiments thereof); thereby enhancing the bioreactivity and / or binding efficacy of the protein. In embodiments, the method of enhancing the bioreactivity and / or binding efficacy of a protein comprising an unnatural amino acid comprises mutating an amino acid proximal to the unnatural amino acid to arginine; wherein the amino acid is Ala, Ile, Leu, Met, Val, Phe, Trp, Tyr, Asn, Cys, Gln, Ser, Thr, Asp, Glu, His, Lys, Gly, or Pro; thereby enhancing the bioreactivity and / or binding efficacy of the protein. Mutating the amino acid to arginine results in the “non-naturally occurring arginine.” The protein can be any protein described herein, including a protein having the amino acid side chain of Formula (2) or any embodiment thereof, a protein of Formula (3) or any embodiment thereof, a protein having the amino acid side chain of Formula (6) or any embodiment thereof, a protein of Formula (7) or any embodiment thereof. In embodiments, the protein is an antibody or an antibody variant. In embodiments, the antibody variant is a single-chain variable fragment, a single-domain antibody, an affibody, or an antigen-binding fragment. In embodiments, the protein is a single-chain variable fragment. In embodiments, the protein is a single-domain antibody. In embodiments, the protein is an affibody. In embodiments, the protein is an antigen- binding fragment. In embodiments, the method further comprises contacting the protein with a target protein, thereby covalently bonding the protein to the target protein.

[0254] Compounds for Diagnostic Imaging

[0255] Provided herein are proteins comprising a detectable label. The protein can be any protein described herein, including a protein having the amino acid side chain of Formula (2) or any embodiment thereof, a protein of Formula (3) or any embodiment thereof, a protein having the amino acid side chain of Formula (6) or any embodiment thereof, a protein of Formula (7) or any embodiment thereof. The detectable label can be any known in the art. In embodiments, the detectable label is a detectable label that can be used in medical imaging. In embodiments, the detectable label is a label that can be used for radiography, magnetic resonance imaging, nuclearmedicine, ultrasound elastography, photoacoustic imaging, tomography, echocardiography, functional near-infrared spectroscopy, magnetic particle imaging. In embodiments, the detectable label is a label that can be use for tomography. In embodiments, the detectable label is a label that can be used for positron emission tomography.

[0256] A “detectable agent” or “detectable moiety” is a composition detectable by appropriate means such as spectroscopic, photochemical, biochemical, immunochemical, chemical, magnetic resonance imaging, or other physical means. In embodiments, the proteins described herein are bonded to a detectable agent. In embodiments, an antibody or antibody variant is bonded to a detectable agent. In embodiments, a nanobody is bonded to a detectable agent. In embodiments, the bond is noncovalent or covalent. In embodiments, the bond is covalent. In embodiments, the protein is covalently bonded to a detectable agent. In embodiments, the antibody or antibody variant is covalently bonded to a detectable agent. In embodiments, a nanobody is covalently bonded to a detectable agent. In embodiments when the protein is covalently bonded to a detectable agent, the covalent bond is between the detectable agent and a naturally-occurring amino acid in the protein. In embodiments when the nanobody is covalently bonded to a detectable agent, the covalent bond is between the detectable agent and a naturally- occurring amino acid in the nanobody. Methods for covalently bonding detectable agents to proteins are well-known in the art. Detectable agents include18F,32P,33P,45Ti,47Sc,52Fe,59Fe,62Cu,64Cu,67Cu,67Ga,68Ga,77As,86Y,90Y.89Sr,89Zr,94Tc,94Tc,99mTc,99Mo,105Pd,105Rh,111Ag,111In,123I,124I,125I,131I,142Pr,143Pr,149Pm,153Sm,154-1581Gd,161Tb,166Dy,166Ho,169Er,fluorophore (e.g., fluorescent dyes), electron-dense reagents, enzymes (e.g., as commonly used in an ELISA), biotin, digoxigenin, paramagnetic molecules, paramagnetic nanoparticles, ultrasmall superparamagnetic iron oxide (“USPIO”) nanoparticles, USPIO nanoparticle aggregates, superparamagnetic iron oxide (“SPIO”) nanoparticles, SPIO nanoparticle aggregates, monocrystalline iron oxide nanoparticles, monochrystalline iron oxide, nanoparticle contrast agents, liposomes or other delivery vehicles containing Gadolinium chelate (“Gd- chelate”) molecules, Gadolinium, radioisotopes, radionuclides (e.g., carbon-11, nitrogen-13, oxygen-15, fluorine-18, rubidium-82), fluorodeoxyglucose (e.g., fluorine-18 labeled), any gamma ray emitting radionuclides, positron-emitting radionuclide, radiolabeled glucose, radiolabeled water, radiolabeled ammonia, biocolloids, microbubbles (e.g. including microbubble shells including albumin, galactose, lipid, and / or polymers; microbubble gas core including air, heavy gases, perfluorcarbon, nitrogen, octafluoropropane, perflexane lipidmicrosphere, perflutren, etc.), iodinated contrast agents (e.g., iohexol, iodixanol, ioversol, iopamidol, ioxilan, iopromide, diatrizoate, metrizoate, ioxaglate), barium sulfate, thorium dioxide, gold, gold nanoparticles, gold nanoparticle aggregates, fluorophores, two-photon fluorophores, or haptens and proteins or other entities which can be made detectable, e.g., by incorporating a radiolabel into a peptide or antibody specifically reactive with a target peptide. A detectable moiety is a monovalent detectable agent or a detectable agent capable of forming a bond with another composition. In embodiments, paramagnetic ions that may be used as imaging agents in accordance with the embodiments of the disclosure include, e.g., ions of transition and lanthanide metals (e.g., metals having atomic numbers of 21-29, 42, 43, 44, or 57- 71). These metals include ions of Cr, V, Mn, Fe, Co, Ni, Cu, La, Ce, Pr, Nd, Pm, Sm, Eu, Gd, Tb, Dy, Ho, Er, Tm, Yb and Lu.

[0257] A “radioisotope” that may be used as imaging and / or labeling agents in accordance with the embodiments of the disclosure include, but are not limited to,18F,32P,33P,45Ti,47Sc,52Fe,59Fe,62Cu,64Cu,67Cu,67Ga,68Ga,77As,86Y,90Y.89Sr,89Zr,94Tc,94Tc,99mTc,99Mo,105Pd,105Rh,111Ag,111In,123I,124I,125I,131I,142Pr,143Pr,149Pm,153Sm,154-1581Gd,161Tb,166Dy,166Ho,169Er,175Lu,177Lu,186Re,188Re,189Re,194Ir,198Au,199Au,211At,211Pb,212Bi,212Pb,213Bi,223Ra and225Ac. In embodiments, the proteins described herein are bonded to a radioisotope. In embodiments, an antibody or antibody variant is bonded to a radioisotope. In embodiments, a nanobody is bonded to a radioisotope. In embodiments, the bond is noncovalent or covalent. In embodiments, the bond is covalent. In embodiments, the protein is covalently bonded to a radioisotope. In embodiments, the antibody or antibody variant is covalently bonded to a radioisotope. In embodiments, a nanobody is covalently bonded to a radioisotope. In embodiments when the protein is covalently bonded to a radioisotope, the covalent bond is between the radioisotope and a naturally-occurring amino acid in the protein. In embodiments when the nanobody is covalently bonded to a radioisotope, the covalent bond is between the radioisotope and a naturally-occurring amino acid in the nanobody. Methods for covalently bonding radioisotopes to proteins are well-known in the art.

[0258] In embodiments, the detectable label is a radioisotope. In embodiments, the detectable label is an idoine radioisotope. In embodiments, the radioisotope is123I,124I,125I, or131I. In embodiments, the radioisotope is123I. In embodiments, the radioisotope is124I. In embodiments, the radioisotope is125I. In embodiments, the radioisotope is131I. In embodiments, the radioisotope is a positron-emitting radioisotope. In embodiments, the positron-emitting radioisotope is11C,13N,15O,18F,64Cu,68Ga,78Br,82Rb,86Y,89Zr,90Y,22Na,26Al,40K,83Sr, or124I. In embodiments, the positron-emitting radioisotope is11C. In embodiments, the positron-emitting radioisotope is13N. In embodiments, the positron-emitting radioisotope is15O. In embodiments, the positron-emitting radioisotope is18F. In embodiments, the positron-emitting radioisotope is64Cu. In embodiments, the positron-emitting radioisotope is168Ga. In embodiments, the positron-emitting radioisotope is78Br. In embodiments, the positron-emitting radioisotope is82Rb. In embodiments, the positron-emitting radioisotope is86Y. In embodiments, the positron-emitting radioisotope is89Zr. In embodiments, the positron-emitting radioisotope is90Y. In embodiments, the positron-emitting radioisotope is22Na. In embodiments, the positron- emitting radioisotope is26Al. In embodiments, the positron-emitting radioisotope is40K. In embodiments, the positron-emitting radioisotope is83Sr. In embodiments, the positron-emitting radioisotope is124I. In embodiments, the radioisotope is an alpha-emitting radioisotope. In embodiments, the alpha-emitting radioisotope is211At,227Th,225Ac,223Ra,213Bi, or212Bi. In embodiments, the alpha-emitting radioisotope is211At. In embodiments, the alpha-emitting radioisotope is227Th. In embodiments, the alpha-emitting radioisotope is225Ac. In embodiments, the alpha-emitting radioisotope is223Ra. In embodiments, the alpha-emitting radioisotope is213Bi. In embodiments, the alpha-emitting radioisotope is212Bi.

[0259] Pharmaceutical Compositions

[0260] Any of the proteins described herein may be administered to a subject in a pharmaceutical composition further comprising a pharmaceutically acceptable excipient. The compositions are suitable for formulation and administration in vitro or in vivo. Suitable carriers and excipients and their formulations are known in the art and described, e.g., Remington: The Science and Practice of Pharmacy, 21st Ed, Lippicott Williams & Wilkins (2005). The term “pharmaceutical composition” encompasses compositions administered to a patient for therapeutic purposes (e.g., treating a disease) and / or diagnostic purposes (e.g., medical imaging). Medical imagining includes, without limitation, radiography, magnetic resonance imaging, nuclear medicine, ultrasound elastography, photoacoustic imaging, tomography (e.g., positron emission tomography), echocardiography, functional near-infrared spectroscopy, magnetic particle imaging, and the like.

[0261] “Pharmaceutically acceptable excipient” and “pharmaceutically acceptable carrier” refer to a substance that aids the administration of an active agent to and absorption by a subject and can be included in the compositions of the disclosure without causing a significant adverse toxicological effect on the patient. Non-limiting examples of pharmaceutically acceptable excipients include water, NaCl, normal saline solutions, lactated Ringer’s, normal sucrose, normal glucose, binders, fillers, disintegrants, lubricants, coatings, sweeteners, flavors, salt solutions (such as Ringer's solution), alcohols, oils, gelatins, carbohydrates such as lactose,amylose or starch, fatty acid esters, hydroxymethycellulose, polyvinyl pyrrolidine, and colors, and the like. Such preparations can be sterilized and, if desired, mixed with auxiliary agents such as lubricants, preservatives, stabilizers, wetting agents, emulsifiers, salts for influencing osmotic pressure, buffers, coloring, and / or aromatic substances and the like that do not deleteriously react with the compounds of the disclosure. One of skill in the art will recognize that other pharmaceutical excipients are useful. Pharmaceutically acceptable excipients can be used in pharmaceutical compositions for therapeutic purposes (e.g., treating a disease) and / or diagnostic purposes (e.g., imaging, such as positron emission tomography).

[0262] Solutions of the pharmaceutical compositions can be prepared in water suitably mixed with a lipid or surfactant, such as hydroxypropylcellulose. Dispersions can also be prepared in glycerol, liquid polyethylene glycols, and mixtures thereof and in oils. Under ordinary conditions of storage and use, these preparations can contain a preservative to prevent the growth of microorganisms. Solutions can be administered, e.g., parenterally, such as subcutaneously or intravenously (e.g., infusion or bolus).

[0263] Pharmaceutical compositions can be delivered via intranasal or inhalable solutions. The intranasal composition can be a spray, aerosol, or inhalant. The inhalable composition can be a spray, aerosol, or inhalant. Nasal solutions can be aqueous solutions designed to be administered to the nasal passages in drops or sprays. Nasal solutions can be prepared so that they are similar in many respects to nasal secretions. Thus, the aqueous nasal solutions usually are isotonic and slightly buffered to maintain a pH of 5.5 to 6.5. In addition, antimicrobial preservatives, similar to those used in ophthalmic preparations and appropriate drug stabilizers, if required, may be included in the formulation. Various commercial nasal preparations are known in the art.

[0264] Oral formulations can include excipients as, for example, pharmaceutical grades of mannitol, lactose, starch, magnesium stearate, sodium saccharine, cellulose, magnesium carbonate and the like. These compositions take the form of solutions, suspensions, tablets, pills, capsules, sustained release formulations or powders. In aspects, oral pharmaceutical compositions will comprise an inert diluent or edible carrier, or they may be enclosed in hard or soft shell gelatin capsule, or they may be compressed into tablets, or they may be incorporated directly with the food. For oral therapeutic administration, the active compounds may be incorporated with excipients and used in the form of ingestible tablets, buccal tablets, troches, capsules, elixirs, suspensions, syrups, wafers, and the like. The percentage of the compositions and preparations may, of course, be varied and may be between about 1 to about 75% of the weight of the unit. The amount of nucleic acids in such compositions is such that a suitabledosage can be obtained.

[0265] For parenteral administration in an aqueous solution, for example, the solution should be suitably buffered and the liquid diluent first rendered isotonic with sufficient saline or glucose. Aqueous solutions, in particular, sterile aqueous media, are especially suitable for intravenous, intramuscular, subcutaneous and intraperitoneal administration. For example, one dosage could be dissolved in 1 ml of isotonic NaCl solution and either added to 1000 ml of hypodermoclysis fluid or injected at the proposed site of infusion.

[0266] Sterile injectable solutions can be prepared by incorporating the recombinant proteins in the required amount in the appropriate solvent followed by filtered sterilization. Generally, dispersions are prepared by incorporating the various sterilized active ingredients into a sterile vehicle which contains the basic dispersion medium. Vacuum-drying and freeze-drying techniques, which yield a powder of the active ingredient plus any additional desired ingredients, can be used to prepare sterile powders for reconstitution of sterile injectable solutions. The preparation of more, or highly, concentrated solutions for direct injection is also contemplated. Dimethyl sulfoxide can be used as solvent for rapid penetration, delivering high concentrations of the active agents to a small area.

[0267] For vaccination or immunization purposes the recombinant proteins or single-domain antibodies provided herein including embodiments thereof, may be formulated and introduced as a vaccine through oral, intradermal, intramuscular, intraperitoneal, intravenous, subcutaneous, intranasal, and via scarification (scratching through the top layers of skin, e.g., using a bifurcated needle) or any other standard route of immunization. Vaccine formulations suitable for oral administration may be in the form of capsules, cachets, pills, tablets, lozenges (using a flavored basis, usually sucrose and acacia or tragacanth), powders, granules, or as a solution or a suspension in an aqueous or non-aqueous liquid, or as an oil-in-water or water-in-oil liquid emulsion, or as an elixir or syrup, or as pastilles (using an inert base, such as gelatin and glycerin, or sucrose and acacia), each containing a predetermined amount of a subject composition thereof as an active ingredient or any other oral composition as listed above. Alternatively, the vaccines may be administered parenterally as injections (intravenous, intramuscular or subcutaneous). The amount of recombinant proteins used in a vaccine can depend upon a variety of factors including the route of administration, species, and use of booster administration. However, a person of ordinary skill in the art would immediately recognize appropriate and / or equivalent doses looking at dosages of approved whopping cough vaccines for guidance.

[0268] The dosage and frequency (single or multiple doses) of the proteins administered to asubject can vary depending upon a variety of factors, for example, whether the mammal suffers from another disease, and its route of administration; size, age, sex, health, body weight, body mass index, and diet of the recipient; nature and extent of symptoms of the disease being treated, kind of concurrent treatment, complications from the disease being treated or other health- related problems. Other therapeutic regimens or agents can be used in conjunction with the methods and proteins (e.g., recombinant proteins, antibodies, antibody variants, single-domain antibodies) described herein. Adjustment and manipulation of established dosages (e.g., frequency and duration) are within the ability of the skilled artisan.

[0269] For any composition, proteins described herein, the effective amount can be initially determined from cell culture assays. Target concentrations will be those concentrations of proteins that are capable of achieving the methods described herein, as measured using the methods described herein or known in the art. As is known in the art, effective amounts of proteins for use in humans can also be determined from animal models. For example, a dose for humans can be formulated to achieve a concentration that has been found to be effective in animals. The dosage in humans can be adjusted by monitoring effectiveness and adjusting the dosage upwards or downwards, as described above. Adjusting the dose to achieve maximal efficacy in humans based on the methods described above and other methods is well within the capabilities of the ordinarily skilled artisan.

[0270] Embodiments

[0271] Embodiment 1. A compound of Formula (1) or a stereoisomer thereof or a compoundof Formula (5) or a stereoisomer thereof: O wherein: L4is a1 to 8, x1 is an integer from 0 to 5, L1is a bond, substituted or unsubstituted alkylene, or substituted or unsubstituted heteroalkylene, R1is hydrogen or an electron withdrawing group, and R2and R3are each independently substituted or unsubstituted C1-5alkyl.

[0272] Embodiment 2. The compound of embodiment 1, wherein L1is a bond.

[0273] Embodiment 3. The compound of embodiment 1, wherein L1is substituted or unsubstituted 2 to 6 membered heteroalkylene.

[0274] Embodiment 4. The compound of embodiment 3, wherein L1is -C(O)-NH-(CH2)y-, -C(O)-O-NH-(CH2)y-, –NH-C(O)-(CH2)y- or –NH-C(O)-O-(CH2)y-, and y is an integer from 0 to 2.

[0275] Embodiment 5. The compound of embodiment 3, wherein L1is –NH-C(O)- or -C(O)- NH-

[0276] Embodiment 6. The compound of any one of embodiments 1 to 5, wherein R2and R3are each independently unsubstituted C1-4 alkyl.

[0277] Embodiment 7. The compound of embodiment 6, wherein R2and R3are methyl.

[0278] Embodiment 8. The compound of any one of embodiments 1 to 7, wherein L4is bonded to the carbon atom para to the carbon atom to which L1is bonded.

[0279] Embodiment 9. The compound of any one of embodiments 1 to 7, wherein L4is bonded to the carbon atom meta to the carbon atom to which L1is bonded.

[0280] Embodiment 10. The compound of any one of embodiments 1 to 9, wherein L4is –O-.

[0281] Embodiment 11. The compound of any one of embodiments 1 to 9, wherein L4is -O- (CH2)1-5-.

[0282] Embodiment 12. The compound of any one of embodiments 1 to 9, wherein L4is - (CH2)1-5-.

[0283] Embodiment 13. The compound of any one of embodiments 1 to 9, wherein L4is a bond.

[0284] Embodiment 14. The compound of any one of embodiments 1 to 13, wherein x is an integer from 1 to 4.

[0285] Embodiment 15. The compound of any one of embodiments 1 to 14, wherein x1 is 1.

[0286] Embodiment 16. The compound of any one of embodiments 1 to 15, wherein R1is hydrogen.

[0287] Embodiment 17. The compound of any one of embodiments 1 to 15, wherein: R1is hydrogen, halogen, -CX13, -CHX12, -CH2X1, -OCX13, -OCH2X1, -OCHX12, -CN, -SOn1R1A, -SOv1NR1AR1B, -NHC(O)NR1AR1B, -N(O)m1, -NR1AR1B, -C(O)R1A, -C(O)-OR1A, -C(O)NR1AR1B, -OR1A, -NR1ASO2R1B, -NR1AC(O)R1B, -NR1AC(O)OR1B, -NR1AOR1B, substituted or unsubstituted alkyl, or substituted or unsubstituted heteroalkyl, X1is independently –F, -Cl, -Br, or –I, R1Ais hydrogen, substituted or unsubstituted alkyl, or substituted or unsubstituted heteroalkyl, R1Bis hydrogen, substituted or unsubstituted alkyl, orsubstituted or unsubstituted heteroalkyl, n1 is an integer from 0 to 4, m1 is 1 or 2, and v1 is 1 or 2.

[0288] Embodiment 18. The compound of embodiment 17, wherein R1is halogen, -CX13, - CHX12, -CH2X1, -OCX13, -OCH2X1, -OCHX12, -CN, -SOn1R1A, -N(O)m1, -C(O)R1A, -C - -C - - -NR1AC -NR1AC1-4.

[0290] Embodiment 20. The compound of embodiment 17, wherein R1is halogen.

[0291] Embodiment 21. The compound of embodiment 1, wherein the compound of Formula (1) is: .

[0292] Embodiment 22.the compound of Formula (I) is: H2N 2 .

[0293] of Formula (II) is: .

[0294] is a bond and x is 1.

[0295] Embodiment 25. The compound of embodiment 23, wherein x is 4 and L1is–C(O)NH-, wherein–C(O)- is adjacent the phenyl ring.

[0296] Embodiment 26. A protein comprising an unnatural amino acid, wherein the unnatural amino comprises a side chain of Formula (III) or Formula (IV): O L4P NR2R3wherein L4is a bond, 1 to 8, x1 is an 1integer from 0 to 5, L a or or substituted or unsubstituted heteroalkylene, and R1is hydrogen or an electron withdrawing group, R2and R3are each independently substituted or unsubstituted C1-5 alkyl.

[0297] Embodiment 27. The protein of embodiment 26, wherein L1is a bond.

[0298] Embodiment 28. The protein of embodiment 26, wherein L1is substituted or unsubstituted 2 to 6 membered heteroalkylene.

[0299] Embodiment 29. The protein of embodiment 28, wherein L1is -C(O)-NH-(CH2)y-, –NH-C(O)-(CH2)y- or –NH-C(O)-O-(CH2)y-, and y is an integer from 0 to 2.

[0300] Embodiment 30. The protein of embodiment 28, wherein L1is –NH-C(O)- or -C(O)-NH-

[0301] Embodiment 31. The protein of any one of embodiments 26 to 30, wherein R2and R3are each independently unsubstituted C1-4 alkyl.

[0302] Embodiment 32. The protein of embodiment 31, wherein R2and R3are methyl.

[0303] Embodiment 33. The protein of any one of embodiments 26 to 32, wherein L4is bonded to the carbon atom para to the carbon atom to which L1is bonded.

[0304] Embodiment 34. The protein of any one of embodiments 26 to 32, wherein L4is bonded to the carbon atom meta to the carbon atom to which L1is bonded.

[0305] Embodiment 35. The protein of any one of embodiments 26 to 34, wherein L4is –O-.

[0306] Embodiment 36. The protein of any one of embodiments 26 to 34, wherein L4is -O- (CH2)1-5-.

[0307] Embodiment 37. The protein of any one of embodiments 26 to 34, wherein L4is -(CH2)1-5-.

[0308] Embodiment 38. The protein of any one of embodiments 26 to 34, wherein L4is a bond.

[0309] Embodiment 39. The protein of any one of embodiments 26 to 38, wherein x is an integer from 1 to 4.

[0310] Embodiment 40. The protein of any one of embodiments 26 to 39, wherein x1 is 1.

[0311] Embodiment 41. The protein of any one of embodiments 26 to 40, wherein R1is hydrogen.

[0312] Embodiment 42. The protein of any one of embodiments 26 to 40, wherein R1is hydrogen, halogen, -CX13, -CHX12, -CH2X1, -OCX13, -CH2X1, -OCHX12, -CN, -SOn1R1A, -SOv1NR1AR1B, -NHC(O)NR1AR1B, -N(O)m1, -NR1AR1B, -C(O)R1A, -C(O)-OR1A, -C(O)NR1AR1B, -OR1A, -NR1ASO2R1B, -NR1AC(O)R1B, -NR1AC(O)OR1B, -NR1AOR1B, substituted or unsubstituted alkyl, or substituted or unsubstituted heteroalkyl, X1is independently –F, -Cl, -Br, or –I, R1Ais hydrogen, substituted or unsubstituted alkyl, or substituted or unsubstituted heteroalkyl, R1Bis hydrogen, substituted or unsubstituted alkyl, or substituted or unsubstituted heteroalkyl, n1 is an integer from 0 to 4, m1 is 1 or 2, and v1 is 1 or 2.

[0313] Embodiment 43. The protein of embodiment 42, wherein R1is halogen, -CX13, - CHX12, -CH2X1, -OCX13, -OCH2X1, -OCHX12, -CN, -SOn1R1A, -N(O)m1, -C(O)R1A, -C(O)-OR1A, -C(O)NR1AR1B, -OR1A, -NR1ASO2R1B, -NR1AC(O)R1B, -NR1AC(O)OR1B, or -NR1AOR1B.

[0314] Embodiment 44. The protein of embodiment 42, wherein R1is unsubstituted heteroalkyl, wherein the heteroalkyl is –O(CH2)1-4.

[0315] Embodiment 45. The protein of embodiment 42, wherein R1is halogen.

[0316] Embodiment 46. The protein of embodiment 26, wherein the unnatural amino acid side chain is: .

[0317] Embodiment 47. the unnatural amino acid side chain is:.

[0318] Embodim natural amino acid side chain is: .

[0319] Embodiment1L is a bond and x is 1.

[0320] Embodiment 50. The protein of embodiment 48, wherein x is 4 and L1is –C(O)NH-, wherein the –C(O)- is adjacent the phenyl ring.

[0321] Embodiment 51. The protein of any one of embodiments 26 to 50, wherein the protein is an antibody.

[0322] Embodiment 52. The protein of any one of embodiments 26 to 50, wherein the protein is a single-chain variable fragment, a single-domain antibody, an affibody, or an antigen-binding fragment.

[0323] Embodiment 53. The protein of embodiment 51 or 52, wherein the unnatural amino acid is within a CDR region of the protein.

[0324] Embodiment 54. The protein of any one of embodiments 26 to 50, wherein the protein is a target protein.

[0325] Embodiment 55. The protein of any one of embodiments 26 to 54, wherein the protein comprises: (i) an arginine proximal to the unnatural amino acid or (ii) a non-naturally occurring arginine proximal to the unnatural amino acid.

[0326] Embodiment 56. The protein of any one of embodiments 26 to 54, wherein the protein comprises a non-naturally occurring arginine proximal to the unnatural amino acid.

[0327] Embodiment 57. The protein of any one of embodiments 26 to 54, wherein the protein comprises a non-naturally occurring arginine proximal to the unnatural amino acid.

[0328] Embodiment 58. The protein of any one of embodiments 26 to 57, further comprising a detectable agent.

[0329] Embodiment 59. The protein of embodiment 58, wherein the detectable agent is aradioisotope.

[0330] Embodiment 60. A nucleic acid encoding the protein of any one of embodiments 26 to 59.

[0331] Embodiment 61. A vector comprising a nucleic acid, wherein the nucleic acid encodes the protein of any one of embodiments 26 to 59.

[0332] Embodiment 62. A protein conjugate of Formula (4) or Formula (8): or , wherein: R4and R5-(CH2)1-5-, -O-(CH2)1-5-, or –O-, x is an integer from 1 to 8, x1 is an integer from 0 to 5, L1is a bond, substituted or unsubstituted alkylene, or substituted or unsubstituted heteroalkylene, and R1is hydrogen or an electron withdrawing group, R2and R3are each independently substituted or unsubstituted C1-5 alkyl, L2is a bond, -NR2A-, -S-, -S(O)2-, -O-, -C(O)-, -C(O)O-, -OC(O)-, -N(R2A)C(O)-, -C(O)N(R2A)-, -NR2AC(O)NR2B-, -NR2AC(NH)NR2B-, -SO2N(R2A)-, -N(R2A)SO2-, -C(S)-, substituted or unsubstituted alkylene, substituted or unsubstituted heteroalkylene, substituted or unsubstituted cycloalkylene, substituted or unsubstituted heterocycloalkylene, substituted or unsubstituted arylene, substituted or unsubstituted heteroarylene, a combination of two thereof, or a combination of three thereof, L3is a bond, -N(R3A)-, -S-, -S(O)2-, -O-, -C(O)-, -C(O)O-, -OC(O)-, -N(R3A)C(O)-, -C(O)N(R3A)-, -NR3AC(O)NR3B-, -NR3AC(NH)NR3B-, -SO2N(R3A)-, -N(R3A)SO2-, -C(S)-, substituted or unsubstituted alkylene, substituted or unsubstituted heteroalkylene, substituted or unsubstituted cycloalkylene, substituted or unsubstituted heterocycloalkylene, substituted or unsubstituted arylene, substituted or unsubstituted heteroarylene, a combination of two thereof, or a combination of three thereof, and R2A, R2B, R3A, and R3Bare independently hydrogen, substituted or unsubstituted alkyl, substituted or unsubstituted heteroalkyl, substituted or unsubstituted cycloalkyl, substituted or unsubstituted heterocycloalkyl, substituted or unsubstituted aryl, substituted or unsubstituted heteroaryl, a combination of two thereof, or a combination of three thereof.

[0333] Embodiment 63. The protein conjugate of embodiment 62, wherein L2is a bond.

[0334] Embodiment 64. The protein conjugate of embodiment 62 or 63, wherein L1is a bond.

[0335] Embodiment 65. The protein conjugate of embodiment 62 or 63, wherein L1is substituted or unsubstituted 2 to 6 membered heteroalkylene.

[0336] Embodiment 66. The protein conjugate of embodiment 65, wherein L1is -C(O)-NH-(CH2)y-, -C(O)-O-NH-(CH2)y-, –NH-C(O)-(CH2)y- or –NH-C(O)-O-(CH2)y-, and y is an integer from 0 to 2.

[0337] Embodiment 67. The protein conjugate of embodiment 65, wherein L1is –NH-C(O)- or -C(O)-NH-.

[0338] Embodiment 68. The protein conjugate of any one of embodiments 62 to 67, wherein R2and R3are each independently unsubstituted C1-4 alkyl.

[0339] Embodiment 69. The protein conjugate of embodiment 68, wherein R2and R3are methyl.

[0340] Embodiment 70. The protein conjugate of any one of embodiments 62 to 69, wherein L4is bonded to the carbon atom para to the carbon atom to which L1is bonded.

[0341] Embodiment 71. The protein conjugate of any one of embodiments 62 to 69, wherein L4is bonded to the carbon atom meta to the carbon atom to which L1is bonded.

[0342] Embodiment 72. The protein conjugate of any one of embodiments 62 to 71, wherein L4is –O-.

[0343] Embodiment 73. The protein conjugate of any one of embodiments 62 to 71, wherein L4is -O-(CH2)1-5-.

[0344] Embodiment 74. The protein conjugate of any one of embodiments 62 to 71, wherein L4is -(CH2)1-5-.

[0345] Embodiment 75. The protein conjugate of any one of embodiments 62 to 71, wherein L4is a bond.

[0346] Embodiment 76. The protein conjugate of any one of embodiments 62 to 75, wherein x is an integer from 1 to 4.

[0347] Embodiment 77. The protein conjugate of any one of embodiments 62 to 76, wherein x1 is 1.

[0348] Embodiment 78. The protein conjugate of any one of embodiments 62 to 77, wherein R1is hydrogen.

[0349] Embodiment 79. The protein conjugate of any one of embodiments 62 to 77, wherein: R1is hydrogen, halogen, -CX13, -CHX12, -CH2X1, -OCX13, -OCH2X1, -OCHX12, -CN, -SOn1R1A, -SOv1NR1AR1B, -NHC(O)NR1AR1B, -N(O)m1, -NR1AR1B, -C(O)R1A, -C(O)-OR1A, -C(O)NR1AR1B, -OR1A, -NR1ASO2R1B, -NR1AC(O)R1B, -NR1AC(O)OR1B, -NR1AOR1B,substituted or unsubstituted alkyl, or substituted or unsubstituted heteroalkyl, X1is independently –F, -Cl, -Br, or –I, R1Ais hydrogen, substituted or unsubstituted alkyl, or substituted or unsubstituted heteroalkyl, R1Bis hydrogen, substituted or unsubstituted alkyl, or substituted or unsubstituted heteroalkyl, n1 is an integer from 0 to 4, m1 is 1 or 2, and v1 is 1 or 2.

[0350] Embodiment 80. The protein conjugate of embodiment 79, wherein R1is halogen, -CX13, -CHX12, -CH2X1, -OCX13, -OCH2X1, -OCHX12, -CN, -SOn1R1A, -N(O)m1, - - - - - - -1-4.

[0352] Embodiment 82. The protein conjugate of embodiment 79, wherein R1is halogen.

[0353] Embodiment 83. The protein conjugate of embodiment 62, wherein the protein conjugate of Formula (4) is: .

[0354] Embodiment 84.62, wherein the protein conjugate of Formula (4) is: .

[0355] the protein conjugate of Formula (8) is: .

[0356] Embodiment wherein L1is a bond and x is 1.

[0357] Embodiment 87. The protein conjugate of embodiment 85, wherein x is 4 and L1is–C(O)NH-, wherein –C(O)- is adjacent the phenyl ring.

[0358] Embodiment 88. The protein conjugate of any one of embodiments 62 to 87, wherein L2is a bond and L3-R5is: .

[0359] Embodiment 89. The of embodiments 62 to 87, whereinL2is a bond and L3-R5is: .

[0360] Embodiment 90. Theof embodiments 62 to 87, wherein L2is a bond and L3-R5is: .

[0361] Embodiment 91. Theof embodiments 62 to 87, wherein L2is a bond and L3-R5is: .

[0362] Embodiment 92. The proteinone of embodiments 62 to 91, wherein the peptidyl moiety of R5comprises a target protein.

[0363] Embodiment 93. The protein conjugate of any one of embodiments 62 to 91, wherein the peptidyl moiety of R4comprises an antibody, and the peptidyl moiety of R5comprises a target protein.

[0364] Embodiment 94. The protein conjugate of any one of embodiments 62 to 91, wherein the peptidyl moiety of R4comprises an antibody variant, and the peptidyl moiety of R5comprises a target protein.

[0365] Embodiment 95. The protein conjugate of embodiment 94 wherein the antibody variant is a single-chain variable fragment, a single-domain antibody, an affibody, or an antigen-binding fragment.

[0366] Embodiment 96. The protein conjugate of any one of embodiments 62 to 91, wherein the peptidyl moiety of R4comprises a target protein and the peptidyl moiety of R5comprises an antibody.

[0367] Embodiment 97. The protein conjugate of any one of embodiments 62 to 91, wherein the peptidyl moiety of R4comprises a target protein and the peptidyl moiety of R5comprises anantibody variant.

[0368] Embodiment 98. The protein conjugate of embodiment 97, wherein the antibody variant is a single-chain variable fragment, a single-domain antibody, an affibody, or an antigen- binding fragment.

[0369] Embodiment 99. The protein conjugate of any one of embodiments 92 to 98, wherein the target protein is a receptor protein.

[0370] Embodiment 100. The protein conjugate embodiment 99, wherein the receptor protein is expressed on a cancer cell.

[0371] Embodiment 101. The protein conjugate of any one of embodiments 92 to 98, wherein the target protein is an extracellular domain receptor protein, a transmembrane domain receptor protein, or an intracellular domain receptor protein.

[0372] Embodiment 102. The protein conjugate of any one of embodiments 92 to 98, wherein the target protein is a cytosolic protein, a transcriptional factor, or an enzyme.

[0373] Embodiment 103. The protein conjugate of any one of embodiments 92 to 98, wherein the target protein is a programmed death-ligand 1 receptor, a programmed cell death protein 1 receptor, a 5-hydroxytryptamine receptor, an acetylcholine receptor, an adenosine receptor, an adenosine A2A receptor, an adenosine A2B receptor, an angiotensin receptor, an apelin receptor, a bile acid receptor, a bombesin receptor, a bradykinin receptor, a cannabinoid receptor, a chemerin receptor, a chemokine receptor, a cholecystokinin receptor, a Class A Orphan receptor, a dopamine receptor, an endothelin receptor, an epidermal growth factor receptor, a formyl peptide receptor, a free fatty acid receptor, a galanin receptor, a ghrelin receptor, a glycoprotein hormone receptor, a gonadotrophin-releasing hormone receptor, a G protein-coupled receptor, a G protein-coupled estrogen receptor, a histamine receptor, a hydroxycarboxylic acid receptor, a kisspeptin receptor, a leukotriene receptor, a lysophospholipid receptor, a lysophospholipid S1P receptor, a melanin-concentrating hormone receptor, a melanocortin receptor, a melatonin receptor, a motilin receptor, a neuromedin U receptor, a neuropeptide FF / neuropeptide AF receptor, a neuropeptide S receptor, a neuropeptide W / neuropeptide B receptor, a neuropeptide Y receptor, a neurotensin receptor, an opioid receptor, an opsin receptor, an orexin receptor, an oxoglutarate receptor, a P2Y receptor, a platelet-activating factor receptor, a prokineticin receptor, a prolactin-releasing peptide receptor, a prostanoid receptor, a proteinase-activated receptor, a QRFP receptor, a relaxin family peptide receptor, a somatostatin receptor, a succinate receptor, a tachykinin receptor, a thyrotropin-releasing hormone receptor, a trace amine receptor, a urotensin receptor, or a vasopressin receptor.

[0374] Embodiment 104. A complex comprising the compound of any one of embodiments 1to 25 and a pyrrolysyl-tRNA synthetase.

[0375] Embodiment 105. The complex of embodiment 104, further comprising a tRNAPyl.

[0376] Embodiment 106. A cell comprising: (i) the compound of any one of embodiments 1 to 25, (ii) the protein of any one of embodiments 26 to 59, (iii) the nucleic acid of embodiment 60, (iv) the vector of embodiment 61, (v) the protein conjugate of any one of embodiments 62 to 103, or (vi) the complex of embodiment 104 or 105.

[0377] Embodiment 107. The cell of embodiment 106, wherein the cell is a bacterial cell or a mammalian cell.

[0378] Embodiment 108. A method of enhancing the bioreactivity or binding efficacy of the protein of any one of embodiments 26 to 59, the method comprising mutating a naturally- occurring amino acid in the protein to arginine, thereby producing a non-naturally occurring arginine, wherein the non-naturally occurring arginine is proximal to the unnatural amino acid.

[0379] Embodiment 109. A method of enhancing the bioreactivity or binding efficacy of a protein, the method comprising: (i) mutating a first amino acid to an unnatural amino acid, and (ii) mutating a second amino acid proximal to the first amino acid to arginine, wherein the unnatural amino comprises a side chain of Formula (2) or Formula (6): O wherein: L4is a bond,1 to 8, x1 is an integer from 0 to 5, L1is a bond, substituted or unsubstituted alkylene, or substituted or unsubstituted heteroalkylene, and R1is hydrogen or an electron withdrawing group, R2and R3are each independently substituted or unsubstituted C1-5 alkyl.

[0380] Embodiment 110. The method of embodiment 108 or 109, wherein the naturally- occurring amino acid in the protein is Ala, Ile, Leu, Met, Val, Phe, Trp, Tyr, Asn, Cys, Gln, Ser, Thr, Asp, Glu, His, Lys, Gly, or Pro. EXAMPLES

[0381] The following examples are intended to further illustrate certain embodiments of the disclosure. The examples are put forth so as to provide one of ordinary skill in the art and are notintended to limit its scope.

[0382] Two new Uaas PFY and PFK bearing phosphoramidofluoridate have been designed and incorporated into proteins in E. coli and mammalian cells through genetic code expansion. Conditions enabling PFEx in proteins were explored to generate precise covalent linkages in proteins both in vitro and in cells.

[0383] We investigated whether PFEx could be introduced in proteins and work in biocompatible conditions to generate new covalent linkages in proteins. Through designing, synthesizing, and genetic encoding new latent bioreactive amino acids PFY and PFK that bear phosphoramidofluoridate, we succeeded in introducing the new click PFEx reaction into proteins. We demonstrated that the incorporated PFY or PFK was able to covalently target a nearby His, Tyr, Lys, or Cys residue in proteins through proximity-enabled PFEx reactivity without external reagents, allowing for precise covalent cross-linking of interacting proteins both in vitro and in cells.

[0384] Our study unveils interesting new aspects of PEFx click chemistry in protein context under biological conditions. In small molecule PFEx reaction, polar aprotic organic solvents are used, and typically a Lewis-base-amine catalyst combined with a silicon additive are required to accelerate the reaction. Sun et al, Chem 9, 2128–2143 (2023). Such solvents and reagents are generally incompatible with proteins and other biomolecules. Our results demonstrate that PFEx could also occur in protic aqueous solvents and even inside the complex cellular environments, thereby establishing a robust basis for its potential biological applications. In addition, we showed that PFEx reaction in proteins could proceed without an external catalyst, indicating that placing the reactants in proximity was sufficient to activate the PFEx reaction. Interestingly, we found that a water-soluble silicon reagent, Na2SiO3, was able to boost the PFEx reaction between PFY and Cys / Tyr in proteins; however, Na2SiO3diminished the PFEX reaction between PFY and His. Moreover, another valuable finding is the temperature impact difference between Tyr and His, representing P(V)-O and P(V)-N linkage, respectively. We showed that His reacted with PFY at low temperature (4 ^C) efficiently, but the resultant crosslink was virtually broken at 95 ^C, while the crosslink formed with Tyr remained stable. Furthermore, we found that Cys also reliably cross-linked with PFY, in particular when facilitated by Na2SiO3, indicating that thiol is a functionality for PFEx reaction as well.

[0385] Example 1

[0386] Design, synthesis, and genetic incorporation of PFY into proteins in E. coli and mammalian cells.

[0387] We designed PFY as a tyrosine analog to contain a phosphoramidofluoridate (FIG.1A), because phosphoramidic difluorides decompose at room temperature in several hours while phosphoramidofluoridates are bench stable. We hypothesized that the phosphoramidofluoridate side chain of PFY would react with nucleophilic side chains of natural amino acid residues via PFEx when they are brought into close proximity, effect of which would be able to activate the PFEx reactivity of PFY (FIG.1B).

[0388] To synthesize PFY, the dimethylphosphoramidic difluoride 2 was first prepared from 1 through fluoride-chloride halogen exchange, and was then directly used for PFEx reaction with a protected tyrosine 3 followed by deprotection to give PFY in overall yield of 57% (FIG.1C, Supplementary information). See Sun et al, Chem 9, 2128–2143 (2023). To evaluate PFY’s cytotoxicity, we added different concentrations of PFY to E. coli or mammalian cells. The growth of E. coli was not affected by 4 mM of PFY in the growth media (FIG.8A), and cell viability assay showed that HEK-293T cells could tolerate 2 mM PFY well (FIG.8B), indicating no toxicity of PFY to both bacterial and mammalian cells.

[0389] With reference to FIG.1C, Compound 4 was synthesized as follows. To a stirred solution of compound 1 (6.6 g, 40.8 mmol) in acetone (150 mL) was added KF (18.9 g, 325.9 mmol). The mixture was stirred at room temperature for 3 h and then filtered through Celite. The solvent was removed under reduced pressure to give phosphoramidic difluoride compound 2, which was directly used for the next step. To a stirred solution benzyl protected tyrosine 3 (11.5 g, 28.4 mmol) and the freshly prepared compound 2 in ACN (100 mL) was added bis(trimethylsilyl)amine (HMDS, 6.5 g, 40.7 mmol) and 2-tert-butyl-1,1,3,3- tetramethylguanidine (BTMG, 1.4 g, 8.1 mmol). The resulting reaction was stirred at room temperature for 1 h. Then 200 ml EtOAc was added to dilute the reaction mixture and the organic phase was washed sequentially with H2O (100 mL) and brine (100 mL). The organic phase was dried over anhydrous Na2SO4and evaporated under reduced pressure to give the crude product, which was then purified by column chromatography (silica gel, Hex: EA=7:3) to give compound 4 white solid (8.6 g, 58.8 %).

[0390] Synthesis of PFY. To a 50 mL round bottom flask was added 20 mg 10% Pd / C, followed by addition of 2 mL MeOH under inert atmosphere (Argon). Compound 4 (200 mg, 0.39 mmol) in MeOH (4 mL) was added to the flask and the flask was sealed by a rubber stopper. The round bottom flask was vacuumed and charged with H2 gas by an H2 balloon. The mixture was stirred vigorously under H2 for 1 h (monitored by LC-MS) and then filtered through Celite. The solvent was removed under reduced pressure to PFY as white solid (110 mg, 97.2%).1H NMR (D2O): δ 7.36 (d, J = 9.2 Hz, 2H), 7.26 (d, J = 9.2 Hz, 2H), 3.99 (dd, J = 7.6 Hz, 1H), 3.31 - 3.12 (m, 2H), 2.83 (d, J = 4.0 Hz, 3H), 2.81 (d, J = 4.0 Hz, 3H).13C NMR (100 MHz,D2O): δ 173.6, 148.6 (d, J = 7.0 Hz, C-F), 133.1, 131.2, 120.3 (d, J = 5.0 Hz, C-F), 55.9, 35.7,35.6, 35.6. 31P NMR (162 MHz, MeOD) δ 0.5 (d, J = 977.6 Hz, P-F); 19F NMR (376 MHz,MeOD) δ -78.0 (d, J = 977.6 Hz, P-F). MS calcd for C11H17FN2O4P [M+H]+291.0904, found: 291.0902.

[0391] Example 2

[0392] To genetically encode PFY into proteins in E. coli, we identified a mutant pyrrolysyl- tRNA synthetase (PylRS) specific for PFY to pair with the pyrrolysyl-tRNA (tRNAPyl) for PFY incorporation through genetic code expansion via amber suppression. As PFY is similar to Uaas NpY, FSY, and mFSY in structure, we tested if synthetases specific for these Uaas could incorporate PFY. Hoppmann et al, Nat. Chem. Biol.13, 842+ (2017); Wang et al, J. Am. Chem. Soc.140, 4995–4999 (2018); Klauser et al, Chem. Commun.58, 6861–6864 (2022). Using a GFP-based fluorescence assay we found that synthetases evolved for NpY and mFSY both were able to incorporate PFY, and the mFSY-specific synthetase showed higher incorporation efficiency (FIG.9 and structures below). We thus used this synthetase for subsequent experiments and named it PFYRS for clarity.we expressed the Zspa affibody (Afb) gene containing a TAG codon at site 36 (Afb-36TAG) together with the tRNAPyl / PFYRS genes in E. coli. No full-length Afb was detected when PFY was not added in the cell culture; in the presence of 2 mM PFY, full-length Afb36PFY protein was produced in the yield of 5.0 mg / L (FIG.1D). We then purified the Afb(36PFY) protein, digested it with proteases, and analyzed with tandem mass spectrometry (MS). A series of b and y ions clearly indicated that PFY was incorporated at the TAG-specified position 36, but the dimethylamino group of PFY was hydrolyzed. This observation is consistent with the known hydrolysis of phosphoramidate group under acidic conditions used in liquid chromatography during MS analysis. Hoppmann et al, Nat. Chem. Biol.13, 842+ (2017). To further verify PFY incorporation, we used the tRNAPyl / PFYRS pair to incorporate PFY into another protein, ubiquitin (Ub), at sites 51 and 54, respectively. The Ub proteins were purified and analyzed using electrospray ionization time-of-flight MS (ESI-TOF-MS). For the Ub(51PFY) intact protein, a peak observed at 9530 Da corresponds to Ub containing PFY at site 51 (expected9530.7 Da); the other peak observed at 9503 Da corresponds to Ub containing PFY at site 51 but with the dimethylamino group of PFY hydrolyzed (expected 9503.6 Da); no peaks corresponding to Ub containing other amino acids at site 51 were observed. Similar results were also obtained for the Ub(54PFY) intact protein, in which PFY was incorporated at a different site 54. These results demonstrate that the identified tRNAPyl / PFYRS pair incorporated PFY into proteins with high efficiency and specificity in E. coli.

[0394] We further evaluated PFY incorporation in mammalian cells. The tRNAPyl / PFYRS genes were co-transfected with the EGFP gene containing a TAG stop codon at the permissive site 182 [EGFP(182TAG) gene] into HEK-293T cells. Only when PFY was added to the growth media did cells show green GFP fluorescence, suggesting the expression of full-length GFP by incorporating PFY at the TAG codon (FIG.1E); Western blot analysis of the cell lysate further confirmed full-length GFP expression in the presence of PFY (FIG.1F). These cells were next analyzed by flow cytometry. The percentage of GFP-positive cells increased with PFY concentration and incubation time (FIG.1G). Cell fluorescence intensity also increased with PFY concentration (FIG.1H). These data indicated that the tRNAPyl / PFYRS pair was able to incorporate PFY in mammalian cells efficiently and specifically.

[0395] Example 3

[0396] PFY reacts with His, Tyr, Lys, and Cys in proteins through proximity-enabled PFEx reaction.

[0397] We assessed whether PFY could react with natural amino acid side chains in proteins through proximity-enabled PFEx reaction. The Afb-Z protein pair was used, which binds in moderate affinity (Kd6 uM). To place PFY and target residue in proximity upon Afb-Z binding, we incorporated PFY in Afb at site 36 and mutated different amino acids at site 6 of the Z protein on the basis of the Afb-Z complex structure (FIG.2A). See Högbom et al, Proc. Natl. Acad. Sci.100, 3191–3196 (2003). Considering that the phosphoramidofluoridate in PFY is a weak electrophile, we decided to test amino acid residues with nucleophilic side chains. Maltose binding protein (MBP) was fused to the N-terminus of the Z protein to better separate the Afb and Z proteins of similar molecular weights. The purified Afb(36PFY) and MBP-Z(6X) proteins were incubated in PBS buffer for 16 hours and then analyzed by SDS-PAGE and Western blot under denatured conditions (FIGS.2B-2C). Among 12 amino acid residues tested, Lys and Cys showed relatively weak crosslinking, while His and Tyr showed strong crosslinking. To validate the crosslink, the cross-linked protein samples were analyzed by tandem MS. Covalently crosslinked peptides of Afb(36PFY) and MBP-Z(6X) were clearly identified, and fragmented ions unambiguously indicate that the incorporated PFY specifically cross-linked with the targetHis, Tyr, Lys, or Cys placed in proximity.

[0398] To show PFEx reactivity in a different protein context, we incorporated PFY into a nanobody SR4 that binds the Spike protein of SARS-CoV-2. Based on the crystal structure of the SR4-Spike complex, we incorporated PFY at site 57 of SR4 to target a proximal Tyr505 of the Spike protein. PFY was separately incorporated at site 54 of SR4, which is further away from Tyr505, as a control. Indeed, SR4(57PFY) cross-linked the Spike protein efficiently, while SR4(54PFY) did not, confirming that PFY reactivity was proximity-driven (FIGS.10A-10C). Li et al, Nat. Commun.12, 4635.10.1038 (2021).

[0399] Besides in vitro cross-linking, we further evaluated PFY reactivity in proteins directly in mammalian cells. We incorporated PFY into E. coli glutathione transferase (ecGST), a homodimeric protein, at its dimer interface and expressed the mutant ecGST in HEK-293T cells to test PFY’s ability to crosslink ecGST into a covalent dimer in a mammalian cellular environment. Specifically, in light of the structure of ecGST (FIG.2D), we incorporated PFY at site 103 of ecGST to target the proximal His106 in the other monomer. Nishida et al, J. Mol. Biol.281, 135–147 (1998). The nearby Lys107 was mutated to Ala so that His106 was the only reachable target. His106 was also mutated to Tyr, Lys, or Cys for testing, and to Ala as the negative control. The mutant ecGST was expressed in HEK-293T cells, and cell lysates were analyzed by Western blot to detect ecGST dimeric cross-linking (FIG.2E). As expected, no dimeric cross-linking was observed for Ala. Strong dimeric cross-linking was observed for His and Tyr, while weak dimeric cross-linking was observed for Lys and Cys, consistent with what was observed in the Afb-Z protein system. Together, these results indicate that PFY was able to react with His, Tyr, Lys, and Cys placed in proximity in proteins both in vitro and in cells.

[0400] Example 4

[0401] pH effect on PFY reaction varies with target residues

[0402] We next studied the effect of pH on PFY reaction in proteins. Purified Afb(36PFY) was incubated with MBP-Z(6X) at pH 7.4 or 8.8 overnight, followed with SDS-PAGE analysis (FIG.3A). The cross-linking efficiency of the two proteins was determined by densitometric measurement of the crosslink band and the MBP-Z band. The cross-linking efficiency between PFY and Lys or Cys did not change significantly upon pH change. The cross-linking efficiency between PFY and Tyr increased by 11% when pH was shifted to basic (FIG.3B), which may be explained by the enhanced deprotonation of the phenol side chain of Tyr. Unexpectedly, the cross-linking efficiency between PFY and His decreased by 14% when pH was increased to 8.8 (FIG.3B). This decrease suggests that factors other than side chain deprotonation may have also contributed to final cross-linking extent, such as the stability of the linkage.

[0403] Example 5

[0404] Temperature impact on PFY reaction differs between His and Tyr

[0405] We discovered unexpectedly that the reaction products of PFY with His and Tyr showed divergent thermostability. We first incubated Afb(36PFY) with MBP-Z(6His) or MBP- Z(6Tyr) at 37 ^C for 16 hours to allow cross-linking. Aliquots of samples were treated with or without 95 ^C heating for 10 minutes followed with SDS-PAGE (FIGS.4A-4B). MBP-Z(6Tyr) showed robust cross-linking with Afb(36PFY) regardless the heating (FIG.4C), indicating that the resultant P(V)-O linkage was thermostable under the test condition. In contrast, MBP- Z(6His) showed robust cross-linking when the sample was not heat treated, while 95 ^C treatment resulted in drastic decrease of the cross-linked product (FIG.4D), indicating that the associated P(V)-N linkage was unstable at 95 ^C. We further tested incubation of Afb(36PFY) with MBP-Z(6His) or MBP-Z(6Tyr) at 4 ^C for 16 hours followed with SDS-PAGE analysis without heating the samples. At this low temperature, both His and Tyr could form crosslink with PFY, but His was more efficient than Tyr (FIGS.4C-4D). Therefore, His reacted with PFY efficiently at low temperature 4 ^C but its P(V)-N linkage was unstable at 95 ^C, while Tyr needed higher temperature 37 ^C to react with PFY efficiently and its P(V)-O linkage remained stable at 95 ^C. Using another protein pair which afforded robust cross-linking of PFY with Cys, we confirmed that the P(V)-S linkage was also stable at 95 ^C (FIGS.11A-11B).

[0406] Example 6

[0407] Effect of Na2SiO3 on PFY reaction varies with target residues.

[0408] We further explored potential reagents to catalyze the PFEx reaction in proteins. PFEx reactions among small molecules require a catalyst, for which 1,5,7-triazabicyclo[4.4.0]dec-5- ene (TBD) has been reported to be most efficient, and the presence of the silicon-containing additive hexamethyldisilazane. Sun et al, Chem 9, 2128–2143 (2023). Using Afb(36PFY) cross- linking with MBP-Z(6X) and MBP-Z(24PFY) cross-linking with Afb(7X), we found that TBD had no catalytic effect for PFY reactions in proteins, possibly because the solvent for proteins has to be aqueous while that for small molecules is mainly the aprotic acetonitrile. The additive hexamethyldisilazane is insoluble in water and thus cannot be used for proteins. We reasoned that the water soluble Na2SiO3 might function similarly as SiO2 is often used to scrub fluoride.

[0409] Indeed, when Na2SiO3was added to the reaction of MBP-Z(24PFY) with Afb(7X) (FIG.5A), we found that Na2SiO3 significantly increased the cross-linking between PFY and Cys (FIG.5B). The increase was dependent on the Na2SiO3concentration in the range of 1-5 mM; further increase to 10 mM then decreased the cross-linking. Similar increase was alsofound for the cross-linking between PFY and Tyr when Na2SiO3concentration was increased from 0 to 2 mM (FIG.5C). Interestingly, the cross-linking efficiency between PFY and His was decreased by the addition of Na2SiO3(FIG.5D). In this protein context, the effect of Na2SiO3on PFY reaction with Lys was not obvious possibly because these two sites were not optimal for their crosslink.

[0410] We then added Na2SiO3 to the reaction of Afb(36PFY) with MBP-Z(6X) (FIG.5E). In this protein context, PFY reaction with Lys or Cys was nonoptimal, and thus the effect of Na2SiO3 was similarly not obvious. However, for Tyr and His which yielded robust cross- linking with PFY, we found that the cross-linking efficiency between PFY and Tyr also increased when Na2SiO3 was added in the range of 1 to 5 mM and decreased at 10 mM (FIG. 5F); consistently, the cross-linking efficiency between PFY and His decreased upon Na2SiO3addition (FIG.5G). Together, these data indicate that Na2SiO3 boosts the PFY reaction with Tyr and Cys but diminishes the PFY reaction with His.

[0411] Example 7

[0412] Design, synthesis, and genetic incorporation of PFK with a flexible long side chain

[0413] PFY was designed as a Tyr analog with a rigid side chain and length close to those of canonical amino acids. To provide flexibility and longer reaction radius, we designed PFK by installing the phosphoramidofluoridate onto the Lys backbone, resulting in an extension of the final side chain length (FIG.6A). We synthesized PFK starting from dimethylphosphoramidic dichloride through fluoride-chloride halogen exchange, direct PFEx reaction with a protected 4- hydroxybenzoic acid, and conjugation to the protected Lys followed with deprotection.

[0414] With reference to FIG.7,compound 5 was synthesized. To a stirred solution of compound 1 (6.0 g, 37.0 mmol) in acetone (150 mL) was added KF (17.1 g, 296 mmol). The mixture was stirred at r.t. for 3 h and then filtered through Celite. The solvent was removed under reduced pressure to give phosphoramidic difluoride compound 2, which was directly used for the next step. To a stirred solution compound 3 (5.1 g, 22.4 mmol) and the freshly prepared compound 2 in ACN (100 mL) was added HMDS (6.0 g, 37.2 mmol) and BTMG (1.3 g, 7.4 mmol). The resulting reaction was stirred at room temperature for 1 h. Then 200 ml EtOAc was added to dilute the reaction mixture and the organic phase was washed sequentially with H2O (200 mL) and brine (200 mL). The organic phase was dried over anhydrous Na2SO4 and evaporated under reduced pressure to give the crude product, which was then purified by column chromatography (silica gel, Hex: EA=4:1) to give compound 4 as colorless oil (4.6 g, 60.9 %).

[0415] To a 50 mL round bottom flask was added 100 mg 10% Pd / C, followed by addition of 5 mL MeOH under inert atmosphere (Argon). Compound 4 (1.0 g, 3.0 mmol) in MeOH (5 mL)was added to the flask and the flask was sealed by a rubber stopper. The round bottom flask was vacuumed and charged with H2 gas by an H2 balloon. The mixture was stirred vigorously under H2for 1 h (monitored by LC-MS) and then filtered through Celite. The solvent was removed under reduced pressure to give compound 5 as white solid (627 mg, 84.6 %).1H NMR (MeOD): δ 8.09 (d, J = 8.4 Hz, 2H), 7.34 (d, J = 8.4 Hz, 2H), 2.85 (d, J = 2.0 Hz, 3H), 2.82 (d, J = 2.0 Hz, 3H).13C NMR (100 MHz, MeOD): δ 168.6, 154.6 (d, J = 6.2 Hz, C-F), 133.1, 129.8, 120.8 (d, J= 5.4 Hz, C-F), 36.6, 36.6. 31P NMR (162 MHz, MeOD) δ 0.0 (d, J = 977.6 Hz, P-F); 19F NMR(376 MHz, MeOD) δ -77.4 (d, J = 977.6 Hz, P-F). MS calcd for C9H12FNO4P [M+H]+248.0482, found: 248.0477.

[0416] Synthesis of Compound 7. To a stirred solution compound 5 (580 mg, 2.34 mmol) and 1-ethyl-3-(3-dimethylaminopropyl)carbodiimide (EDC HCl, 581 mg, 3.04 mmol) in anhydrous DCM (8 mL) was added compound 6 (benzenesulfonic acid salt form, 1.24 g, 2.34 mmol) and N,N-diisopropylethylamine (DIPEA, 301 mg, 2.34 mmol) in anhydrous DCM (5 mL). The resulting reaction was stirred at room temperature for 5 h. The solvent was removed under reduced pressure to give crude product, which was diluted with 50 ml EtOAc. The organic phase was washed sequentially with H2O (50 mL) and brine (50 mL), dried over anhydrous Na2SO4 and evaporated under reduced pressure to give the crude product, which was then purified by column chromatography (silica gel, DCM: MeOH = 25:1) to give compound 7 as white solid (948 mg, 67.6 %).

[0417] Synthesis of PFK. To a 50 mL round bottom flask was added 100 mg 10% Pd / C, followed by addition of 4 mL MeOH under inert atmosphere (Argon). Compound 7 (846 mg, 1.41 mmol) in MeOH (6 mL) was added to the flask and the flask was sealed by a rubber stopper. The round bottom flask was vacuumed and charged with H2 gas by an H2 balloon. The mixture was stirred vigorously under H2for 1 h (monitored by LC-MS) and then filtered through Celite. The solvent was removed under reduced pressure to PFK as white solid (513 mg, 97.2 %).1H NMR (MeOD): δ 7.90 (d, J = 8.4 Hz, 2H), 7.33 (d, J = 8.4 Hz, 2H), 3.54 (t, J = 6.2 Hz, 1H), 3.40 (t, J = 6.8 Hz, 2H), 2.84 (d, J = 2.0 Hz, 3H), 2.82 (d, J = 2.0 Hz, 3H), 1.95 -1.90 (m, 2H), 1.88 -1.80 (m, 2H), 1.71 -1.67 (m, 2H);13C NMR (100 MHz, MeOD):δ 174.4, 168.9, 153.5 (d, J = 6 Hz, C-F), 133.4, 130.6, 120.8 (d, J = 6 Hz, C-F), 56.1, 40.7, 36.6, 36.6, 32.0, 30.1, 23.7;31P NMR (162 MHz, MeOD) δ 0.1 (d, J = 940.0 Hz, P-F); 19F NMR (376 MHz, MeOD) δ -77.6(d, J = 940.0 Hz, P-F). MS calcd for C15H24FN3O5P [M+H]+376.1432, found: 376.142.

[0418] Example 8

[0419] To genetically encode PFK into proteins, we reasoned that FSKRS, the synthetase specific for a structurally similar Uaa FSK, should be able to incorporate PFK. Liu et al, J. Am.Chem. Soc.143, 10341–10351 (2021). When the tRNAPyl / FSKRS gene was co-transformed with a nanobody mNb6 gene containing an amber stop codon at site 54 [mNb6(54TAG)] in E. coli cells, full-length mNb6 protein was produced when 1 mM of PFK was added to the growth media (FIG.6B), indicating PFK incorporation. Schoo et al, Science 370, 1473–1479 (2020). To evaluate PFK incorporation in mammalian cells, we co-expressed the tRNAPyl / FSKRS gene and the EGFP(182TAG) gene in HEK-293T cells. Western blot analysis of the cell lysates showed that full-length EGFP was produced only when PFK was added to cell culture (FIG.6C). We further transfected the tRNAPyl / FSKRS gene into the HeLa-GFP(182TAG) reporter cell line that contains the genome-integrated GFP(182TAG) gene. Wang et al, Nat. Neurosci.10, 1063–1072 (2007). Flow cytometric analysis of these cells showed that 30% of cells became fluorescent when 1 mM PFK was added (FIG.12). These data indicate that the tRNAPyl / FSKRS pair was able to incorporate PFK into proteins in both E. coli and mammalian cells. We renamed FSKRS to PFKRS for clarity. The amino acid sequence of PFKRS is set forth as SEQ ID NO:9.

[0420] Example 9

[0421] PFK expands protein cross-linking unreachable by PFY in vitro and in cells

[0422] To confirm that PFK could react with target residue unreachable by PFY, we first incorporated PFK into nanobody mNb6 to evaluate the in vitro cross-linking of mNb6 with its binding target: the Spike protein of SARS-Cov-2. Based on the structure of mNb6-Spike complex (FIG.6D), we decided to incorporate PFK into mNb6 at sites 50-59 individually to target Tyr351 of the Spike protein’s receptor binding domain (RBD). Schoof et al, Science 370, 1473–1479 (2020). PFK-incorporated mNb6 mutant proteins were purified and incubated with the Spike RBD followed with Western blot analysis. Robust cross-linking of mNb6 with the Spike RBD was detected when PFK was incorporated at site 54 but no other sites (FIG.6E). Tandem MS analysis of the cross-linked proteins confirmed that PFK reacted with the target Tyr351 as expected. This site-specific cross-linking indicated that PFK-based PFEx reaction was also proximity-driven and non-random. In contrast, when the shorter PFY was incorporated at site 54, no cross-linking of mNb6 with the Spike RBD was detected (FIG.6F).

[0423] We next incorporated PFK into ecGST expressed in mammalian cells to verify its reactivity and compare with PFY in cells. At the dimer interface of ecGST, we incorporated PFK at site 103 to target Cys10 of the other monomer (FIG.6G). Nishida et al, J. Mol. Biol. 281, 135–147 (1998). The Cα-Cα distance between site 103 and site 10 is 4.2 Å longer than that between site 103 and site 106, which was used to test PFY cross-linking earlier. Nearby His106, Lys107 and Tyr157 were mutated to Ala so that Cys10 was the only target residue in proximity for PFK103. After ecGST(103PFK) was expressed in HEK-293T cells, the cell lysate wasanalyzed by Western blot, which showed robust dimeric ecGST cross-linking (FIG.6H). When we mutated Cys10 to His, strong dimeric ecGST cross-linking was also detected. Weak dimeric cross-linking of ecGST was detected when Cys10 was mutated to Tyr, and no dimeric cross- linking was detected when Cys10 was mutated to Ala in the negative control. These data confirmed that PFK was able to react with Cys, His, and Tyr placed in proximity in mammalian cells. In addition, when we replaced PFK with PFY at site 103, no ecGST dimeric cross-linking was detected for Cys, significantly weaker dimeric cross-linking was detected for His, while significantly stronger dimeric cross-linking was detected for Tyr (FIGS.6H-6I). This difference in cross-linking can be accounted for by the length difference in PFK and PFY side chain: within the fixed distance between site 103 and site 10, the longer PFK was able to reach the shorter side chains of His and Cys, while the shorter PFY matched the longer side chain of Tyr better than PFK. Taken together, these results indicate that PFK was able to react with residues unreachable by PFY in proteins.

[0424] Example 10

[0425] PFY crosslinks nucleophilic residues via proximity-enabled reactivity through the new click PFEx reaction. We discovered that Arg could also accelerate PFY-mediated protein crosslinking. The affibody-Z protein pair described in figure 5 was used here to show the effect. PFY was incorporated at site 36 of the affibody, and residue 32 was mutated to Arg.24 µM of the affibody protein was incubated with 6 µM of MBP-Z(N6Y) protein for different time followed with SDS-PAGE analysis under denatured conditions. As shown in figure 9, Affibody(36PFY) did not show apparent crosslinking with MBP-Z(N6Y) in 18 hours; in contrast, the Arg mutant Affibody(36PFY / 32R) could crosslink MBP-Z(N6Y) in 1 hour with crosslinking efficiency increasing with incubation time.

[0426] Materials and Methods for Examples

[0427] All primers were synthesized and purified by Integrated DNA Technologies (IDT), and plasmids were sequenced by Azenta Life Sciences. All molecular biology reagents were obtained from Vazyme. His-HRP antibody, GFP monoclonal antibodies, and GAPDH-HRP antibody were obtained from ProteinTech Group. pEvol-mFSYRS was used as previously described.1Primers pEvol-NPY-For and pEvol-NPY-Rev were used to construct plasmid pEvol- MmNpYRS.

[0428] All solvents were of reagent grade and were purchased from Fisher Scientific and Aldrich. Reagents were purchased from Aldrich, Enamine, and Asta Tech. The stationary phase of chromatographic purification is silica (230 × 400 mesh, Sorbtech). Silica gel TLC plate was purchased from Sorbtech.1H-NMR (400 MHz) and13C-NMR (100 MHz) spectra were recordedon a Bruker Avance 400 MHz NMR spectrometer. In recording the19F and31P NMR of new compounds, PhCF3 and PPh3 were used as internal standards, respectively.

[0429] SEQ ID NO:1 - Affibody (36TAG) MVDNFNKELSVAGREIVTLPNLNDPQKKAFIRSLWUDPSQSANLLAEAKKLNDAQAPK GSHHHHHH, where U is the amber codon TAG introduced at the 36thposition to encode Uaa incorporation.

[0430] MBP-Z (6X)

[0431] pBAD-MBP-Z (6A) was cloned with primers MBP-Z-N6A-For and MBP-Z-N6A-Rev. pBAD-MBP-Z (6H) was cloned with primers MBP-Z-N6H-For and MBP-Z-N6H-Rev. pBAD- MBP-Z (6Y) was cloned with primers MBP-Z-N6Y-For and MBP-Z-N6Y-Rev. pBAD-MBP-Z (6K) was cloned with primers MBP-Z-N6K-For and MBP-Z-N6K-Rev. pBAD-MBP-Z (6T) was cloned with primers MBP-Z-N6T-For and MBP-Z-N6T-Rev. pBAD-MBP-Z (6S) was cloned with primers MBP-Z-N6S-For and MBP-Z-N6S-Rev. pBAD-MBP-Z (6C) was cloned with primers MBP-Z-N6C-For and MBP-Z-N6C-Rev. pBAD-MBP-Z (6N) was cloned with primers MBP-Z-N6-For and MBP-Z-N6-Rev. pBAD-MBP-Z (6W) was cloned with primers MBP-Z- N6W-For and MBP-Z-N6W-Rev. pBAD-MBP-Z (6M) was cloned with primers MBP-Z-N6M- For and MBP-Z-N6M-Rev. pBAD-MBP-Z (6D) was cloned with primers MBP-Z-N6D-For and MBP-Z-N6D-Rev. pBAD-MBP-Z (6E) was cloned with primers MBP-Z-N6E-For and MBP-Z- N6E-Rev.

[0432] SEQ ID NO:2 MKIEEGKLVIWINGDKGYNGLAEVGKKFEKDTGIKVTVEHPDKLEEKFPQVAATGDGP DIIFWAHDRFGGYAQSGLLAEITPDKAFQDKLYPFTWDAVRYNGKLIAYPIAVEALSLIY NKDLLPNPPKTWEEIPALDKELKAKGKSALMFNLQEPYFTWPLIAADGGYAFKYENGK YDIKDVGVDNAGAKAGLTFLVDLIKNKHMNADTDYSIAEAAFNKGETAMTINGPWAW SNIDTSKVNYGVTVLPTFKGQPSKPFVGVLSAGINAASPNKELAKEFLENYLLTDEGLEA VNKDKPLGAVALKSYEEELAKDPRIAATMENAQKGEIMPNIPQMSAFWYAVRTAVINA ASGRQTVDEALKDAQTNSSSNNNNNNNNNNLGSSGLVPRGGVDNAFXAEQQNAFYEI LHLPNLNEEQRNAFIQSLKDDPSQSANLLAEAKKLNDAQAPKLEHHHHHH, where X is the position of the mutated residues, including A, H, Y, K, T, S, C, N, M, W, D, and E.

[0433] ecGST (103TAG106X107A). PCDNA3.1-ecGST (103TAG106Y107A) was cloned with primers pCDNA-ecGST-106Y107A-For and pCDNA-ecGST-106Y107A-Rev. PCDNA3.1-ecGST (103TAG106K107A) was cloned with primers pCDNA-ecGST-106K107A- For and pCDNA-ecGST-106K107A-Rev. PCDNA3.1-ecGST (103TAG106C107A) was cloned with primers pCDNA-ecGST-106C107A-For and pCDNA-ecGST-106C107A-Rev.

[0434] SEQ ID NO:3 MKLFYKPGACSLASHITLRESGKDFTLVSVDLMKKRLENGDDYFAVNPKGQVPALLLD DGTLLTEGVAIMQYLADSVPDRQLLAPVNSISRYKTIEWLNYIAUELXAGFTPLFRPDTP EEYKPTVRAQLEKKLQYVNEALKDEHWICGQRFTIADAYLFTVLRWAYAVKLNLEGL EHIAAFMQRMAERPEVQDALSAEGLKHHHHHH, where U is the amber codon TAG introduced at site 103 for Uaa incorporation and X is His106 mutated to A, Y, K, and C.

[0435] ecGST (103TAG106A107A157A10X). PCDNA3.1-ecGST (103TAG106A107A157A10A) was cloned with primers pCDNA-ecGST-10A-For and pCDNA- ecGST-10A-Rev. PCDNA3.1-ecGST (103TAG106A107A157A10H) was cloned with primers pCDNA-ecGST-10H-For and pCDNA-ecGST-10H-Rev. PCDNA3.1-ecGST (103TAG106A107A157A10K) was cloned with primers pCDNA-ecGST-10K-For and pCDNA- ecGST-10K-Rev. PCDNA3.1-ecGST (103TAG106A107A157A10Y) was cloned with primers pCDNA-ecGST-10Y-For and pCDNA-ecGST-10Y-Rev.

[0436] SEQ ID NO:4 MKLFYKPGAXSLASHITLRESGKDFTLVSVDLMKKRLENGDDYFAVNPKGQVPALLLD DGTLLTEGVAIMQYLADSVPDRQLLAPVNSISRYKTIEWLNYIAUELAAGFTPLFRPDTP EEYKPTVRAQLEKKLQYVNEALKDEHWICGQRFTIADAALFTVLRWAYAVKLNLEGL EHIAAFMQRMAERPEVQDALSAEGLKHHHHHH, where U is the amber codon TAG introduced at site 103 for Uaa incorporation, and X is Cys10 mutated to A, H, K, and Y.

[0437] SEQ ID NO:5 is ubiquitin MQIFVKTLTGKTITLEVEPSDTIENVKAKIQDKEGIPPDQQRLIFAGKQLEDGRTLSDYNI QKESTLHLVLRLRGGHHHHHH, where E51 and R54 is the amber codon TAG introduced at these sites separately for Uaa incorporation.

[0438] SEQ ID NO:6 is EGFP(182TAG) MVSKGEELFTGVVPILVELDGDVNGHKFSVSGEGEGDATYGKLTLKFICTTGKLPVPWP TLVTTLTYGVQCFSRYPDHMKQHDFFKSAMPEGYVQERTIFFKDDGNYKTRAEVKFEG DTLVNRIELKGIDFKEDGNILGHKLEYNYNSHNVYIMADKQKNGIKVNFKIRHNIEDGS VQLADHUQQNTPIGDGPVLLPDNHYLSTQSALSKDPNEKRDHMVLLEFVTAAGITLGM DELYKHHHHHH, where U is the amber codon TAG introduced at site 182 for Uaa incorporation.

[0439] SEQ ID NO:7 is SR4 QVQLVESGGGLVQAGGSLRLSCAASGFPVYSWNMWWYRQAPGKEREWVAAIESHGD STRYADSVKGRFTISRDNAKNTVYLQMNSLKPEDTAVYYCYVWVGHTYYGQGTQVT VSAGRAGEQKLISEEDLNSAVD, where H54 and S5 are the amber codon TAG introduced atthese sites separately for Uaa incorporation.

[0440] SEQ ID NO:8 is mNb6 QVQLVESGGGLVQAGGSLRLSCAASGYIFGRNAMGWYRQAPGKERELVAGITRRGSI TYYADSVKGRFTISRDNAKNTVYLQMNSLKPEDTAVYYCAADPASPAYGDYWGQGT QVTVSSHHHHHH, where GITRRGSITY (sites 50-59) is where the amber codon TAG was introduced at separately for Uaa incorporation.

[0441] SEQ ID NO:9 is PFKRS amino acid sequence MTVKYTDAQIQRLREYGNGTYEQKVFEDLASRDAAFSKEMSVASTDNEKKIKGMIANP SRHGLTQLMNDIADALVAEGFIEVRTPIFISKDALARMTITEDKPLFKQVFWIDEKRALR PMLAPNLGSVARDLRDHTDGPVKIFEMGSCFRKESHSGMHLEEFTMLNLFDMGPRGDA TEVLKNYISVVMKAAGLPDYDLVQEESDVYKETIDVEINGQEVCSAAVGPTPIDAAHD VHEPWSGAGFGLERLLTIREKYSTVKKGGASISYLNGAKIN

[0442] Incorporation of PFY into EGFP (182TAG) in E. coli

[0443] pBad-EGFP (182TAG) was co-transformed with pEVOL-mFSYRS or pEVOL- NpYRS into DH10β E. coli chemical competent cells and plated on LB agar plate supplemented with 100 µg / mL ampicillin and 34 µg / mL chloramphenicol. A single colony was picked and incubated into 2 mL 2xYT-Amp100Cm34 and grown at 37 °C. When OD600 reached 0.6-0.8, 500 µL cell culture was supplemented with 1 mM or 2 mM PFY and 0.2% arabinose, then incubation at 25 °C for 16 h. As a negative control, 500 µL cell culture was induced with 0.2% arabinose at 25 °C for 16 h. The green fluorescence intensity was measure with a plate reader (excitation at 485 nm, emission at 528 nm).

[0444] Protein expression and purification

[0445] For the incorporation of PFY into Affibody (36TAG), MBP-Z (24TAG), SR4 (54TAG), SR4 (57TAG), mNB6 (54TAG), ubiquitin (51TAG) or ubiquitin (54TAG), pBad- Affibody (36TAG), pBad-MBP-Z (24TAG), pBad- SR4 (54TAG), pBad-SR4 (57TAG), pBad- mNB6(54TAG), pBad-ubiquitin (51TAG) or pBad-ubiquitin (54TAG) was co-transformed with pEvol-PFYRS into DH10β E. coli competent cells. For the incorporation of PFK into mNB6 (54TAG), pBad-mNB6 (54TAG) was co-transformed with pEvol-PFKRS into DH10β. After transformation, a single colony was inoculated into 2 mL of 2xYT-Amp100Cm34 and left grown at 37 °C for 8 h. Then, 0.6 mL cell culture was diluted into 30 mL 2xYT medium and agitated vigorously at 37 °C. When OD600reached 0.8-1.0, the cell culture was added with 0.2% arabinose with or without 2 mM PFY or PFK, and the expression were carried out at 25 °C for 24 h.

[0446] The cell pellets were harvested and resuspended in 30 mL PBS lysis buffer (pH 7.4,137 mM NaCl, 2.7 mM KCl, 8 mM Na2HPO4, 2 mM KH2PO4, 20 mM imidazole, DNase 0.1 mg / mL and protease inhibitors) and lysed by ultrasonication. The soluble fractions were collected and incubated with His-Pur Ni-NTA Resin at 4 °C for 2 h with constant mechanical rotation. The slurry was loaded into column and washed with 150 mL PBS wash buffer (pH 7.4, 137 mM NaCl, 2.7 mM KCl, 8 mM Na2HPO4, 2 mM KH2PO4and 20 mM imidazole), and them eluted with 400 µL elution buffer (0.5x PBS, 500 mM imidazole) for 4 times. The eluates were concentrated and buffer exchanged into 200 µL of PBS buffer and stored at -80 °C.

[0447] Genetic incorporation of PFY into HEK-293T cells

[0448] HEK-293T cells were seeded with 4×105cells per well in a 6 well-cell culture dish containing 2 mL of DMEM media with 10% FBS and 1% P / S, and grown at 37 °C in a CO2 incubator overnight. The plasmid pMP-PFYRS (1.5 µg) was co-transfected with pcDNA 3.1- EGFP (182TAG) (1.5 µg) into target cells with 9 µL polyethyleneimine (1 mg / mL) in 2 mL DMEM media. Six hours post transfection, the cells were treated with or without 2 mM PFY. After incubation for 48 h, the cells were visualized on a Nikon EclipseTi confocal microscope. The cells were then rinsed with PBS twice and lysed with RIPA buffer. The samples were separated on SDS-PAGE and immunoblotted with 1:2500 anti-GFP monoclonal antibody (Proteintech #HRP-66005). Anti-GAPDH antibody (Proteintech #HRP-60004) was used as a reference protein.

[0449] Flow cytometric analysis of Uaa incorporation in mammalian cells

[0450] HEK-293T cells were seeded with 2×105cells per well in a 12 well-cell culture dish containing 1 mL of DMEM media with 10% FBS and 1% P / S, and grown at 37 °C in a CO2 incubator overnight. The plasmid pMP-PFYRS (0.75 µg) was co-transfected with pcDNA 3.1- EGFP (182TAG) (0.75 µg) into target cells with 4.5 µL polyethyleneimine (1 mg / mL) in 1 mL DMEM media. Six hours post transfection, the cells were treated with or without 2 mM PFY and cultured for 24-48 h. For PFK incorporation, plasmid pNEU-PFKRS (0.8 µg) was transfected into HeLa-GFP(182TAG) reporter cells using lipofectamine 3000 transfection reagent and treated with or without 1 mM PFK, incubation for additional 24 h. After incubation, the cells were washed with PBS and analyzed by BD LSRFortessaTMcell analyzer.

[0451] In vitro cross-linking of proteins

[0452] MBP-Z (4A6X7A) and Affibody (36PFY)

[0453] Purified 24 µM Affibody (36PFY) was incubation with 6 µM MBP-Z (4A7A6A / H / Y / K / T / S / C / N / M / W / D / E) in 6 µL PBS buffer at 37 °C for 16 h. The reaction mixture was added 10 µL 4x laemmli sample buffer containing 100 mM DTT and incubation at r.t for 1 h. The samples were separated on SDS-PAGE or Western blot with 1:5000 anti-Hismonoclonal antibody (Proteintech # HRP-66005).

[0454] To explore temperature impact on PFY reaction, 24 µM Affibody (36PFY) was incubation with 6 µM MBP-Z (4A7A6H / Y / K / C) in 6 µL PBS buffer at 37 °C or 4 °C for 16 h. The reaction mixture was added 10 µL 4x laemmli sample buffer containing 100 mM DTT. The samples incubated at 37 ^C were treated with or without 95 ^C heating for 10 minutes; the samples incubated at 4 ^C were maintained at 4 ^C. These samples were loaded into the same gel and gel electrophoresis was run in an ice-water bath to maintain low temperature. To explore potential reagents to catalyze the PFEx reaction in proteins, TBD (final conc.9 µM), Na2SiO3(final conc.0, 1, 1.5, 2, 5,10 mM), or both were added in the reaction buffer. The samples were separated on SDS-PAGE.

[0455] MBP-Z (24PFY) and Affibody (4A7X)

[0456] Purified 24 µM Affibody (4A7H / Y / K / C) was incubation with 6 µM MBP-Z (24PFY) in 6 µL PBS buffer, various concentrations of Na2SiO3(final conc.0, 1, 2, 5,10 mM) were added and incubation at 37 °C for 16 h. The reaction mixture was added 10 µL 4x laemmli sample buffer containing 100 mM DTT and incubation at r.t for 1 h. The samples were separated on SDS-PAGE.

[0457] SR4 (54PFY) or SR4 (57PFY) and Spike RBD

[0458] Purified 5 µM SR4 (54PFY) or SR4 (57PFY) was incubation with 2 µM Spike RBD protein in 10 µL PBS buffer and incubation at 37 °C for 16 h. The reaction mixture was added 10 µL 4x laemmli sample buffer containing 100 mM DTT and incubation at r.t for 1 h. The samples were separated on SDS-PAGE.

[0459] mNB6 (50-59PFK) and Spike RBD

[0460] Purified 5 µM mNB6 (PFK individually incorporated at each site for sites 50-59) was incubation with 0.5 µM Spike RBD protein in 10 µL PBS buffer and incubation at 37 °C for 16 h. The reaction mixture was added 10 µL 4x laemmli sample buffer containing 100 mM DTT. The samples were separated on SDS-PAGE and Western blot with 1:5000 anti-his monoclonal antibody.

[0461] Genetic incorporation of Uaa into ecGST mutants in HEK-293T cells

[0462] HEK-293T cells were seeded with 1×105cells per well in a 24 well-cell culture dish containing 0.5 mL of DMEM media with 10% FBS and 1% P / S, and grown at 37 °C in a CO2 incubator overnight. The plasmid pMP-PFYRS (0.5 µg) or pNEU-PFKRS (0.5 µg) was co- transfected with pcDNA 3.1-ecGST 103TAG / 107A / 106X, X = A, H, Y, K, or C) (0.5 µg) or pcDNA 3.1-ecGST (103TAG / 107A / 106A / 157A / 10X, X = A, H, Y, K, or C) (0.5 µg) intotarget cells with 3 µL polyethyleneimine (1 mg / mL) in 0.5 mL DMEM media. Six hours post transfection, the cells were treated with or without 2 mM PFY or PFK and cultured for 48 h. The cells were then rinsed with PBS and lysed with RIPA buffer. The dimerization of ecGST due to cross-linking was monitored by Western blot using 1:5000 anti-his monoclonal antibody. Anti- GAPDH antibody was used as a reference protein.

[0463] Mass spectrometric analysis. Electrospray ionization mass spectrometric analysis of intact protein was performed as previously described. Protein digestion and tandem mass spectrometry measurement were also performed as previously described. Specifically, the affibody (36PFY) / MBP-Z (4A6X7A) samples were digested with Glu-C and trypsin. The mNB6 (54PFK) / Spike RBD samples were digested with trypsin. Mass spectrometry raw data was searched by pLink 2 and analyzed by the pLabel software. See Klauser et al, Chem. Commun. 58, 6861–6864 (2022); Liu et al, J. Am. Chem. Soc.143, 10341–10351 (2021).

[0464] The examples and figures herein are described in Cao et al, Genetically enabling phosphorous fluoride exchange click chemistry in proteins, Chem, 10:1868-1884 (13 June 2024), the disclosure of which is incorporated by reference herein in its entirety.

[0465] It is understood that the examples described herein are for illustrative purposes only and that various modifications or changes in light thereof will be suggested to persons skilled in the art and are to be included within the spirit and scope of this application and claims. The section headings used herein are for organizational purposes only and are not to be construed as limiting the subject matter described. All documents, or portions of documents, cited in the application are expressly incorporated by reference herein in their entirety and for all purposes.

Claims

CLAIMS What is claimed is:

1. A compound of Formula (1) or a stereoisomer thereof or a compound of Formula (5) or a stereoisomer thereof: O L4P NR2R3wherein: L4is a, , x is an integer from 1 to 8; x1 is an integer from 0 to 5; L1is a bond, substituted or unsubstituted alkylene, or substituted or unsubstituted heteroalkylene; R1is hydrogen or an electron withdrawing group; and R2and R3are each independently substituted or unsubstituted C1-5alkyl.

2. The compound of claim 1, wherein L1is –NH-C(O)-, -C(O)-NH-, -C(O)-NH-(CH2)y-, -C(O)-O-NH-(CH2)y-, –NH-C(O)-(CH2)y- or –NH-C(O)-O-(CH2)y-, and y is an integer from 0 to 2.

3. The compound of claim 1, wherein R2and R3are each independently unsubstituted C1-4 alkyl.

4. The compound of claim 1, wherein: R1is hydrogen, halogen, -CX13, -CHX12, -CH2X1, -OCX13, -OCH2X1, -OCHX12, -CN, -SOn1R1A, -SOv1NR1AR1B, -NHC(O)NR1AR1B, -N(O)m1, -NR1AR1B, -C(O)R1A, -C(O)-OR1A, -C(O)NR1AR1B, -OR1A, -NR1ASO2R1B, -NR1AC(O)R1B, -NR1AC(O)OR1B, -NR1AOR1B, substituted or unsubstituted alkyl, or substituted or unsubstituted heteroalkyl; X1is independently –F, -Cl, -Br, or –I; R1Ais hydrogen, substituted or unsubstituted alkyl, or substituted or unsubstituted heteroalkyl;R1Bis hydrogen, substituted or unsubstituted alkyl, or substituted or unsubstituted heteroalkyl; n1 is an integer from 0 to 4; m1 is 1 or 2; and v1 is 1 or 2 5. The compound of claim 4, wherein R1is halogen or –O(CH2)1-4.

6. The compound of claim 1, wherein the compound of Formula (1) is: 2 . 7.(II) is: .

8. The1, or wherein x is 4 and L1is –C(O)NH-, wherein–C(O)- is adjacent the phenyl ring.

9. A protein comprising an unnatural amino acid, wherein the unnatural amino comprises a side chain of Formula (III) or Formula (IV): OR1O O P F ; wherein: L4is a bond, -x is an integer from 1 to 8; x1 is an integer from 0 to 5; L1is a bond, substituted or unsubstituted alkylene, or substituted or unsubstituted heteroalkylene; and R1is hydrogen or an electron withdrawing group; R2and R3are each independently substituted or unsubstituted C1-5 alkyl.

10. The protein of claim 9, wherein L1is –NH-C(O)-, -C(O)-NH-, -C(O)-NH-(CH2)y-, –NH-C(O)-(CH2)y- or –NH-C(O)-O-(CH2)y-, and y is an integer from 0 to 2.

11. The protein of claim 9, wherein R2and R3are each independently unsubstituted C1-4 alkyl.

12. The protein of claim 9, wherein R1is hydrogen, halogen, -CX13, -CHX12, -CH2X1, -OCX13, -OCH2X1, -OCHX12, -CN, -SOn1R1A, -SOv1NR1AR1B, -NHC(O)NR1AR1B, -N(O)m1, -NR1AR1B, -C(O)R1A, -C(O)-OR1A, -C(O)NR1AR1B, -OR1A, -NR1ASO2R1B, -NR1AC(O)R1B, -NR1AC(O)OR1B, -NR1AOR1B, substituted or unsubstituted alkyl, or substituted or unsubstituted heteroalkyl; X1is independently –F, -Cl, -Br, or –I; R1Ais hydrogen, substituted or unsubstituted alkyl, or substituted or unsubstituted heteroalkyl; R1Bis hydrogen, substituted or unsubstituted alkyl, or substituted or unsubstituted heteroalkyl; n1 is an integer from 0 to 4; m1 is 1 or 2; and v1 is 1 or 2 -C(O)NR1AR1B, -OR1A, -NR1ASO2R1B, -NR1AC(O)R1B, -NR1AC(O)OR1B, or -NR1AOR1B.

13. The protein of claim 12, wherein R1is halogen or –O(CH2)1-4.

14. The protein of claim 9, wherein the unnatural amino acid side chain is:.

15. The ide chain is: .

16. Theis 1 or wherein x is 4 and L1is –C(O)NH-, wherein the –C(O)- is adjacent the phenyl ring.

17. The protein of claim 9, wherein the protein is an antibody, a single-chain variable fragment, a single-domain antibody, an affibody, or an antigen-binding fragment.

18. The protein of claim 17, wherein the unnatural amino acid is within a CDR region of the protein.

19. The protein of claim 9, wherein the protein comprises: (i) an arginine proximal to the unnatural amino acid or (ii) a non-naturally occurring arginine proximal to the unnatural amino acid.

20. The protein of claim 9, wherein the protein comprises a non-naturally occurring arginine proximal to the unnatural amino acid.

21. The protein of claim 9, further comprising a detectable agent.

22. A nucleic acid encoding the protein of claim 9.

23. A vector comprising a nucleic acid, wherein the nucleic acid encodes the protein of claim 9.

24. A protein conjugate of Formula (4) or Formula (8):or ); wherein: R4and R5are L4is a bond, -(CH2)1-5-, -O-(CH2)1-5-, or –O-; x is an integer from 1 to 8; x1 is an integer from 0 to 5; L1is a bond, substituted or unsubstituted alkylene, or substituted or unsubstituted heteroalkylene; and R1is hydrogen or an electron withdrawing group; R2and R3are each independently substituted or unsubstituted C1-5 alkyl; L2is a bond, -NR2A-, -S-, -S(O)2-, -O-, -C(O)-, -C(O)O-, -OC(O)-, -N(R2A)C(O)-, -C(O)N(R2A)-, -NR2AC(O)NR2B-, -NR2AC(NH)NR2B-, -SO2N(R2A)-, -N(R2A)SO2-, -C(S)-, substituted or unsubstituted alkylene, substituted or unsubstituted heteroalkylene, substituted or unsubstituted cycloalkylene, substituted or unsubstituted heterocycloalkylene, substituted or unsubstituted arylene, substituted or unsubstituted heteroarylene, a combination of two thereof, or a combination of three thereof; L3is a bond, -N(R3A)-, -S-, -S(O)2-, -O-, -C(O)-, -C(O)O-, -OC(O)-, -N(R3A)C(O)-, -C(O)N(R3A)-, -NR3AC(O)NR3B-, -NR3AC(NH)NR3B-, -SO2N(R3A)-, -N(R3A)SO2-, -C(S)-, substituted or unsubstituted alkylene, substituted or unsubstituted heteroalkylene, substituted or unsubstituted cycloalkylene, substituted or unsubstituted heterocycloalkylene, substituted or unsubstituted arylene, substituted or unsubstituted heteroarylene, a combination of two thereof, or a combination of three thereof; and R2A, R2B, R3A, and R3Bare independently hydrogen, substituted or unsubstituted alkyl, substituted or unsubstituted heteroalkyl, substituted or unsubstituted cycloalkyl, substituted or unsubstituted heterocycloalkyl, substituted or unsubstituted aryl, substituted or unsubstituted heteroaryl, a combination of two thereof, or a combination of three thereof.

25. The protein conjugate of claim 24, wherein the protein conjugate of Formula (4) is:or 26. A c ysyl-tRNA synthetase.

27. The complex of claim 26, further comprising a tRNAPyl.

28. A cell comprising: (i) the compound of claim 1; (ii) the protein of claim 9; (iii) the nucleic acid of claim 22; (iv) the vector of claim 23; (v) the protein conjugate of claim 24; or (vi) the complex of claim 26.

29. A method of enhancing the bioreactivity or binding efficacy of the protein of claim 9, the method comprising mutating a naturally-occurring amino acid in the protein to arginine, thereby producing a non-naturally occurring arginine; wherein the non-naturally occurring arginine is proximal to the unnatural amino acid.

30. A method of enhancing the bioreactivity or binding efficacy of a protein, the method comprising: (i) mutating a first amino acid to an unnatural amino acid, and (ii) mutating a second amino acid proximal to the first amino acid to arginine; wherein the unnatural amino comprises a side chain of Formula (2) or Formula (6): OR1O O P F ; wherein: L4is a bond, -x is an integer from 1 to 8; x1 is an integer from 0 to 5; L1is a bond, substituted or unsubstituted alkylene, or substituted or unsubstituted heteroalkylene; and R1is hydrogen or an electron withdrawing group; R2and R3are each independently substituted or unsubstituted C1-5 alkyl.

Citation Information

Patent Citations

  • Site-specific generation of phosphorylated tyrosines in proteins

    US20190359965A1

  • Mutant Aminoacyl tRNA Synthetase

    US20220411778A1