Methods for n-terminal peptide sequencing and capture

The method addresses the sensitivity limitations of peptide sequencing by functionalizing peptides and proteins under mild conditions, enabling stable cleavage and capture of N-terminal amino acids, thus improving detection and analysis of low-abundance samples.

WO2026156355A1PCT designated stage Publication Date: 2026-07-23BOARD OF RGT THE UNIV OF TEXAS SYST +1
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
BOARD OF RGT THE UNIV OF TEXAS SYST
Filing Date
2026-01-20
Publication Date
2026-07-23

AI Technical Summary

Technical Problem

Current methods for peptide sequencing, such as mass spectrometry, lack sensitivity for low-abundance samples and are limited by the use of trifluoroacetic acid, which degrades acid-sensitive compounds like oligonucleotides, posing challenges for applications like peptide-DNA barcoding and fluorosequencing.

Method used

A method for functionalizing peptides and proteins under mild conditions (pH 6-8 and 37 °C) using compounds of Formula (I) and (II), followed by cleavage with diamine reagents to remove N-terminal amino acids, allowing for stable sequencing and capture.

Benefits of technology

Enables high-efficiency cleavage of N-terminal amino acids under mild conditions, preserving the stability of fluorescent tags and enhancing the sensitivity and scope of detection in peptide sequencing and capture.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure IMGF000002_0001
    Figure IMGF000002_0001
  • Figure IMGF000002_0002
    Figure IMGF000002_0002
  • Figure IMGF000003_0001
    Figure IMGF000003_0001
Patent Text Reader

Abstract

Methods for functionalizing peptides and proteins in mild conditions for use in, for example, N-terminal peptide sequencing and peptide and protein capture, and related compounds and compositions produced therein.
Need to check novelty before this filing date? Find Prior Art

Description

DESCRIPTIONMETHODS FOR N-TERMINAL PEPTIDE SEQUENCING AND CAPTURE

[0001] This application claims the benefit of priority to United States Provisional Application No. 63 / 747,084, filed on January 19, 2025, the entire contents of which are hereby incorporated by reference.STATEMENT AS TO FEDERALLY SPONSORED RESEARCH

[0002] This invention was made with government support under Grant No. R35 GM149308 awarded by the National Institute of Health. The government has certain rights in the invention.SEQUENCE LISTING INCORPORATION

[0003] This application contains a Sequence Listing XML, which has been submitted electronically and is hereby incorporated by reference in its entirety. Said Sequence Listing XML, created on January 20, 2026, is named UTFBP1390WO.xml, and is 20,067 bytes in size.BRIEF SUMMARY

[0004] The disclosure provides methods for functionalizing peptides and proteins in mild conditions for use in, for example, N-terminal peptide and protein sequencing and peptide and protein capture, and related compounds and compositions produced therein.BACKGROUND

[0005] Peptide sequencing plays an important role in determining the primary structure of proteins, which enhances our understanding of their functions supports the development of novel therapeutics and advances research in biotechnology and synthetic biology. However, the process to identify and quantify proteins often becomes challenging when it comes to reaction scale and sensitivity. Currently, mass spectrometry (MS) is the primary method for protein identification and quantification; however, the lack of sensitivity towards low-abundance samples or rare amino acid variants results in a decreased ability for analysis.

[0006] Highly parallel single-molecule protein sequencing method, known as fluorosequencing, is a viable an alternative to MS. Fluorosequencing allows for identification and quantification of millions of peptides at the same time and rivals MS due to its increased sensitivity and wider scope of detection. In this technique, proteins are enzymatically digested into peptide samples, then one or more specific amino acids are selectively labeled with an14928-1939-2137, v. 2identifier fluorophore, and the peptide samples are immobilized on a glass surface. Total internal reflection fluorescence (TIRF) microscopy is then used to track the reduction in fluorescence during consecutive rounds of Edman degradation which results in sequential removal of N-terminal amino acids. The resulting sparse fluorescent sequences are matched to their corresponding proteins using a reference database.

[0007] This method has been successful in identifying and quantifying protein samples. However, the use of trifluoroacetic acid (TFA) in Edman cycles limits the selection of fluorophores and acid-sensitive compounds, such as oligonucleotides, as both are prone to degradation under strong acidic or basic conditions. This presents significant challenges for applications like peptide-DNA barcoding and fluorosequencing, where the stability of fluorescent tags and other components is crucial.

[0008] To address this issue, new methods are needed which can perform high efficiency cleavage under mild aqueous conditions.SUMMARY

[0009] The disclosure provides methods for functionalizing peptides and proteins for use in, for example, N-terminal peptide and protein sequencing and peptide and protein capture. In particular, the methods herein can be performed under mild conditions (e.g., about pH 6-8 and about 37 °C). In particular, one aspect of the disclosure provides a method of functionalizing a peptide or protein having an N-terminus, the method comprising a coupling step comprising reacting the N-terminus of the peptide or protein with a compound of Formula (I):or a salt thereof, wherein:R1and R2are each, independently, an electron withdrawing group; orR1and R2are taken together with the carbon atom to which they are attached to form a 3-10 membered heterocyclyl or C3-10 carbocyclyl, wherein the 3-10 membered heterocyclyl or C3-10 carbocyclyl is substituted with one or more electron withdrawing groups and optionally further substituted, to form a compound of Formula (II):Peptide or Protein(II).24928-1939-2137, v. 2

[0010] In some embodiments, the method further comprises a step of washing to remove unreacted or excess amount of the compound of Formula (I) using an aqueous buffer, an organic solvent, or a mixture thereof.

[0011] In some embodiments, the method further comprises a step of removing unreacted or excess amount of the compound of Formula (I) by filtration.

[0012] In some embodiments, the method further comprises a step of cleaving the N-terminal amino acid residue of the peptide or protein to form a truncated peptide or protein.

[0013] In some embodiments, the step of cleaving comprises contacting a compound of Formula (II):Peptide or Protein(II),or a salt thereof,with a diamine reagent of Formula (III):H2N Linkeror a salt thereof,to form a compound of Formula (IV):R1R2Peptide or Protein(IV),or a salt thereof,wherein Linker is selected from the group consisting of a bond, optionally substituted Ci-6 alkylene, and optionally substituted Ci-6 heteroalkylene.

[0014] In some embodiments, the method further comprises intramolecularly reacting a compound of Formula (IV):R1R2Linker J Peptide or ProteinN^N(IV),or a salt thereof to form the truncated peptide or protein and a compound of Formula (V):34928-1939-2137, v. 2or a salt thereof; andR4is the side chain of the N-terminal amino acid residue of the peptide or protein.

[0015] In some embodiments, the method further comprises a step of detecting the presence of the truncated peptide or protein or the compound of Formula (V).

[0016] In another aspect, provided herein is a method of N-terminal peptide sequencing of a peptide or protein having an N-terminus, the method comprising the steps ofreacting the N-terminus of the peptide or protein with an EVA reagent:to form a compound of Formula (Il-a):cleaving the N-terminal amino acid residue of the peptide or protein to form a truncated peptide or protein by contacting a compound of Formula (Il-a) with a diamine reagent of Formula (III):Linker NH2(III),or a salt thereof,to form a compound of Formula (IV-a):44928-1939-2137, v. 2or a salt thereof, wherein the compound of Formula (IV-a) or salt thereof intramolecularly reacts to form the truncated peptide or protein and a compound of Formula (V-a):or a salt thereof, whereinLinker is a bond or -C2H4-; andR4is the side chain of the N-terminal amino acid residue of the peptide or protein.

[0017] In some embodiments, the method further comprises a step of washing to remove unreacted or excess amount of the EVA reagent using an aqueous buffer, an organic solvent, or a mixture thereof.

[0018] In some embodiments, the method further comprises a step of removing unreacted or excess amount of the EVA reagent by filtration.

[0019] In some embodiments, the method further comprises a step of detecting the presence of the truncated peptide or protein or the compound of Formula (V-a).

[0020] Also provided herein is a method of sequencing a plurality of peptide or proteins comprising sequencing a plurality of peptide or proteins functionalized by the method provided herein.

[0021] Another aspect of the disclosure provides a compound of Formula (IV):or a salt thereof, whereinR1and R2are each, independently, an electron withdrawing group; orR1and R2are taken together with the carbon atom to which they are attached to form a 3-10 membered heterocyclyl or C3-10 carbocyclyl, wherein the 3-10 membered heterocyclyl or C3-10 carbocyclyl is substituted with one or more electron withdrawing groups and optionally further substituted; and54928-1939-2137, v. 2Linker is selected from the group consisting of a bond, optionally substituted Ci-6 alkylene, and optionally substituted Ci-6 heteroalkylene.

[0022] Another aspect of the disclosure provides a compound of Formula (V):-Linker—or a salt thereof, whereinR1and R2are each, independently, an electron withdrawing group; orR1and R2are taken together with the carbon atom to which they are attached to form a 3-10 membered heterocyclyl or C3-10 carbocyclyl, wherein the 3-10 membered heterocyclyl or C3-10 carbocyclyl is substituted with one or more electron withdrawing groups and optionally further substituted;R4is the side chain of an amino acid residue; andLinker is selected from the group consisting of a bond, optionally substituted C1-6 alkylene, and optionally substituted C1-6 heteroalkylene.

[0023] The disclosure also provides compositions comprising one or more compounds (e.g., a compound of Formula (IV), a compound of Formula (V)) described herein.

[0024] The disclosure also provides kits comprising a compound of Formula (I) (e.g., the EVA reagent), or a salt thereof, for use in a method described herein.BRIEF DESCRIPTION OF THE DRAWINGS

[0025] The features of the present disclosure are set forth with particularity in the appended claims. A better understanding of the features and advantages of the present disclosure will be obtained by reference to the following detailed description that sets forth illustrative embodiments, in which the principles of the disclosure are utilized, and the accompanying drawings of which:

[0026] FIG. 1 shows a structure of 5-[Bis(methylthio)methylene]-2,2-dimethyl-l,3-dioxane-4, 6-dione. This compound is also referred to as the “EVA reagent” and its synthesis is described in Example 1.

[0027] FIG. 2 shows derivatives of the EVA reagent that can also be used in the processes and methods disclosed herein (e.g., in “Ava” chemistry for the removal of an N-terminal amino acid).64928-1939-2137, v. 2

[0028] FIG. 3 shows an exemplary mechanism of the EVA reagent using hydrazine. The EVA reagent shown in the box, upon reaction with the N-terminal amine of the peptide and upon action of hydrazine, removes the N-terminal amino acid under aqueous conditions.

[0029] FIG.4 shows an exemplary mechanism of the EVA reagent using ethylenediamine. The EVA reagent, upon reaction with the N-terminal amine of the model peptide, VYFWVY (SEQ ID NO: 1), and subsequent reaction with ethylenediamine, removes the N-terminal amino acid, through the formation of a stable 8-membered ring.

[0030] FIG.5 displays an exemplary LCMS trace of the product of the reaction shown in FIG.4, which demonstrates the products formed by the Ava chemistry using ethylenediamine. Peak (a) confirms the cleaved 8-membered ring containing the N-terminal amino acid and Peak (b) confirms the presence of truncated peptide, YFWVY (SEQ ID NO: 2). Unreacted peptide, VYFWVY (SEQ ID NO: 1), and unreacted EVA-peptide conjugate Xi YFWVY (SEQ ID NO: 3), wherein Xi is Vai conjugated to the EVA reagent, were also detected.

[0031] FIG. 6 shows a 280 nm LC trace of the EVA-conjugated model peptide intermediate. A model peptide, YGFWVY (SEQ ID NO: 4), was reacted with EVA reagent to form the EVA-conjugated model peptide X2GFWVY (SEQ ID NO: 5), wherein X2 is Tyr conjugated to the EVA reagent. Mass-spectrometry data with an annotated chemical structure is shown in the inset.

[0032] FIG. 7 shows a 280 nm LC trace of the cleavage products of EVA-tyrosine and the truncated peptide following the hydrazine treatment. The EVA-conjugated model peptide X2GFWVY (SEQ ID NO: 5), wherein X2 is Tyr conjugated to the EVA reagent, was cleaved with hydrazine to form EVA-tyrosine and the model truncated peptide, GFWVY (SEQ ID NO: 6). Mass-spectrometry of the LC peaks with annotated chemical formula is shown in the inset.

[0033] FIG. 8 provides a high-resolution mass-spectrometry analysis of the two products from FIG. 7 - truncated peptide, GFWVY (SEQ ID NO: 6), (top panel) and EVA-tyrosine adduct (bottom panel). The structures of the two compounds are annotated.

[0034] FIG. 9 shows (a) an LC-MS trace of an EVA-derivatized peptide having arginine at the N-terminus prior to sequencing and (b) after sequencing. The peptide RYFWVY (SEQ ID NO: 7), was reacted with EVA reagent to form the EVA-conjugated peptide X3 YFWVY (SEQ ID NO: 8), wherein X3 is Arg conjugated to the EVA reagent, which was subsequently cleaved with hydrazine to form EVA-arginine and the truncated peptide YFWVY (SEQ ID NO: 2).74928-1939-2137, v. 2

[0035] FIG. 10 shows (a) an LC-MS trace of an EVA-derivatized peptide having asparagine at the N-terminus prior to sequencing and (b) after sequencing. The peptide NYFWVY (SEQ ID NO: 9), was reacted with EVA reagent to form the EVA-conjugated peptide X4YFWVY (SEQ ID NO: 10), wherein X4 is Asn conjugated to the EVA reagent, which was subsequently cleaved with hydrazine to form EVA-asparagine and the truncated peptide YFWVY (SEQ ID NO: 2).

[0036] FIG. 11 shows (a) an LC-MS trace of an EVA-derivatized peptide having glutamine at the N-terminus prior to sequencing and (b) after sequencing. The peptide QYFWVY (SEQ ID NO: 11), was reacted with EVA reagent to form the EVA-conjugated peptide X5 YFWVY (SEQ ID NO: 12), wherein X5 is Gin conjugated to the EVA reagent, which was subsequently cleaved with hydrazine to form EVA-glutamine and the truncated peptide YFWVY (SEQ ID NO: 2).

[0037] FIG. 12 shows (a) an LC-MS trace of an EVA-derivatized peptide having histidine at the N-terminus prior to sequencing and (b) after sequencing. The peptide HYFWVY (SEQ ID NO: 13), was reacted with EVA reagent to form the EVA-conjugated peptide Xe YFWVY (SEQ ID NO: 14), wherein Xe is His conjugated to the EVA reagent, which was subsequently cleaved with hydrazine to form EVA-glutamine and the truncated peptide YFWVY (SEQ ID NO: 2).

[0038] FIG. 13 shows (a) an LC-MS trace of an EVA-derivatized peptide having phenylalanine at the N-terminus prior to sequencing and (b) after sequencing. The peptide FYFWVY (SEQ ID NO: 15), was reacted with EVA reagent to form the EVA-conjugated peptide X7 YFWVY (SEQ ID NO: 16), wherein X7 is Phe conjugated to the EVA reagent, which was subsequently cleaved with hydrazine to form EVA-glutamine and the truncated peptide YFWVY (SEQ ID NO: 2).

[0039] FIG. 14 shows an exemplary mechanism of the EVA reagent in a N-terminal capture reaction, including (a) the cyclized capture product and (b) the side product.DETAILED DESCRIPTION

[0040] The disclosure provides methods for functionalizing peptides and proteins for use in, for example, N-terminal peptide sequencing and peptide and protein capture. In particular, the methods herein can be performed under mild conditions (e.g., about pH 6-8 and about 37 °C). The practice of the present disclosure employs, unless otherwise indicated, conventional techniques of organic chemistry, pharmacology, molecular biology (including recombinant techniques), cell biology, biochemistry, and immunology. Such techniques are explained in the literature, such as in “Comprehensive Organic Synthesis” (B.M. Trost & I. Fleming, eds., 1991-84928-1939-2137, v. 21992); “Handbook of experimental immunology” (D.M. Weir & C.C. Blackwell, eds.); “Current protocols in molecular biology” (F.M. Ausubel et al., eds., 1987, and periodic updates); and “Current protocols in immunology” (J.E. Coligan et al., eds., 1991), each of which is herein incorporated by reference in its entirety.

[0041] Various aspects of the disclosure are set forth below in sections; however, aspects of the disclosure described in one particular section are not to be limited to any particular section. Further, when a variable is not accompanied by a definition, the previous definition of the variable controls.Definitions

[0042] The term "substantially" or "substantial" as used herein generally refers to at least about 60% or 60%, about 70% or 70%, or about or at 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% or higher relative to a reference such as, for example, the original composition or state of an entity. Thus, an agent that does not “substantially” couple to an internal amino acid indicates that at least about 60% or 60%, about 70% or 70%, or about or at 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or higher amounts of the agent have not reacted with the internal amino acid.

[0043] The term “selective” or “selectively”, as used herein, generally refers to a preference of at least about 50% or 50%, about 60% or 60%, about 70% or 70%, or about or at 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100% for one composition than another composition. For example, a reaction that is “selective” for a C-terminal amino acid of a peptide or protein has about a 50% or 50%, about 60% or 60%, about 70% or 70%, or about or at 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100% preference to react with the C-terminal amino acid than another group of the peptide or protein, such as, for example, an internal amino acid of the peptide or protein.

[0044] Where the use of the term “about” or “approximately” is before a quantitative value, the present disclosure also includes the specific quantitative value itself, unless specifically stated otherwise. As used herein, the term “about” or “approximately” refers to a ±10%, ±5%, ±3%, ±2%, or ±1% variation from the nominal value unless otherwise indicated or inferred from the context.

[0045] As used herein, the term "amino acid" in general refers to organic compounds that contain at least one amino group, — NH2, which may be present in its ionized form, — NHC, and one carboxyl group, — COOH, which may be present in its ionized form, — COO', where the carboxylic acids are deprotonated at neutral pH, having the formula of+NH3CHRCOO'. An94928-1939-2137, v. 2amino acid and thus a peptide has an N (amino)-terminal residue region and a C (carboxy)-terminal residue region. Types of amino acids may include at least 20 that are considered "natural" as they comprise the majority of biological proteins in mammals and include amino acid, such as, for example, lysine, cysteine, tyrosine, threonine, etc. Amino acids may also be grouped based upon their side chains, such as those with a carboxylic acid groups (at neutral pH), including aspartic acid or aspartate (Asp; D) and glutamic acid or glutamate (Glu; E); and basic amino acids (at neutral pH), including lysine (Lys; L), arginine (Arg; N), and histidine (His; H).

[0046] As used herein, the term "terminal" is referred to as singular terminus and plural termini. A “N-terminal amino acid residue” may refer to an amino acid residue at the end of a peptide or protein that has a free NH2 or NH3. A “C-terminal amino acid residue” may refer to an amino acid residue at the end of a peptide or protein that has a free COOH or COO'.

[0047] As used herein, the term "side chains," “residue,” or "R" refers to groups attached to the a-carbon (the carbon that couples the amine and carboxylic acid groups of an amino acid) that render each type of amino acid (e.g., natural amino acid). R groups have a variety of shapes, sizes, charges, and reactivities, such as, for example, charged polar side chains (e.g., positively or negatively charged, such as, for example, lysine (+), arginine (+), histidine (+), aspartate ( ), and glutamate ( )); amino acids can also be basic (e.g., lysine) or acidic (e.g., glutamic acid); uncharged polar side chains may comprise groups, such as, for example, hydroxyl, amide, or thiol groups (e.g., cysteine), which may be a chemically reactive side chain (e.g., a thiol group that can form bonds with another cysteine, serine (Ser) and threonine (Thr)); asparagine (Asn), glutamine (Gin), and tyrosine (Tyr); non-polar hydrophobic amino acid side chains (e.g., glycine, alanine, valine, leucine, and isoleucine) having aliphatic hydrocarbon side chains ranging in size from a methyl group (e.g., alanine) to isomeric butyl groups (e.g., leucine and isoleucine); methionine (Met) has a thiol ether side chain; proline (Pro) has a cyclic pyrrolidine side group. Phenylalanine (with its phenyl moiety) (Phe) and tryptophan (Trp) (with its indole group) contain aromatic side chains, which are characterized by bulk as well as lack of polarity.

[0048] Amino acids can be referred to by a name, 3-letter code, or 1 -letter code, for example, Cysteine, Cys, C; Lysine, Lys, K; Tryptophan, Trp, W, respectively.

[0049] " Unnatural" amino acids are those not naturally encoded or found in the genetic code nor produced via de novo metabolic pathways in mammals and plants. They can be synthesized by adding side chains not normally found or rarely found on amino acids in nature. Examples may include: P-amino acids (e.g., P-alanine), homo-amino acids (e.g., homoserine), proline derivatives (e.g., cis-4-Hydroxy-D-proline), 3-substituted alanine derivatives (e.g., 3,3-104928-1939-2137, v. 2diphenyl-D-alanine), glycine derivatives (e.g., sarcosine), ring-substituted phenylalanine and tyrosine derivatives (e.g., 4-chloro-L-phenylalanine and 3-chloro-L-tyrosine, respectively), linear core amino acids (e.g., 4-amino-3 -hydroxybutyric acid), andN-methyl amino acids (e.g., L-abrine).

[0050] As used herein, [3 amino acids, which have their amino group bonded to the [3 carbon rather than the a-carbon as in the 20 standard biological amino acids, are unnatural amino acids. The only common naturally occurring [3 amino acid is P-alanine.

[0051] As used herein, the terms "amino acid sequence," "peptide," "peptide sequence," "polypeptide," “oligopeptide,” "polypeptide sequence," and “oligopeptide sequence” as used herein refer to at least two amino acids or amino acid analogs that are covalently linked by a peptide (amide) bond or an analog of a peptide bond. The term peptide includes oligomers and polymers of amino acids or amino acid analogs. The term peptide also includes molecules that may be referred to as oligopeptides, which may contain from about two (2) to about twenty (20) amino acids. The term peptide may include molecules that are commonly referred to as polypeptides, which generally contain more than twenty (20) amino acids. The term peptide also includes molecules that are commonly referred to as proteins, which may contain at least about twenty (20) amino acids and a set of defined structural features (e.g., a set of secondary, tertiary, and quaternary structures). The amino acids of the peptide may be L-amino acids or D-amino acids. A peptide, polypeptide, or protein may be synthetic, recombinant, or naturally occurring. A synthetic peptide is a peptide that is produced by artificial means in vitro.

[0052] In some embodiments, the term “peptide or protein” encompasses a chemical structure of a chemically-modified peptide or protein (e.g., a peptide or protein without a terminal amine).

[0053] As used herein, the term “fluorescence” refers to the emission of visible light by a substance that has absorbed light of a different wavelength. Fluorescence may provide a nondestructive way of tracking and / or analyzing biological molecules based on the fluorescent emission at a specific wavelength. Proteins (including antibodies), peptides, nucleic acid, oligonucleotides (including single stranded and double stranded primers) may be “labeled” with a variety of extrinsic fluorescent molecules referred to as fluorophores.

[0054] As used herein, sequencing of peptides “at the single molecule level” refers to amino acid sequence information obtained from individual (i.e., single) peptide molecules, which can be in a mixture of diverse peptide molecules. It is not necessary that the present disclosure be limited to methods where the amino acid sequence information obtained from an individual peptide molecule is the complete or contiguous amino acid sequence of an individual peptide114928-1939-2137, v. 2molecule. It may be sufficient that only partial amino acid sequence information is obtained, allowing for identification of the peptide or protein. Partial amino acid sequence information, including for example, the pattern of a specific amino acid residue (i.e., lysine) within individual peptide molecules, may be sufficient to uniquely identify an individual peptide molecule. For example, a pattern of amino acids such as, for instance, X-X-X-Lys-X-X-X-X-Lys-X-Lys, which indicates the distribution of lysine molecules within an individual peptide molecule, may be searched against a known proteome of a given organism to identify the individual peptide molecule. It is not intended that sequencing of peptides at the single molecule level be limited to identifying the pattern of lysine residues in an individual peptide molecule; sequence information for any amino acid residue (including multiple amino acid residues) may be used to identify individual peptide molecules in a mixture of diverse peptide molecules.

[0055] As used herein, “single molecule sensitivity” refers to the ability to acquire data (including, for example, amino acid sequence information) from individual peptide molecules in a mixture of diverse peptide molecules. In one non-limiting example, the mixture of diverse peptide molecules may be immobilized on a solid surface (including, for example, a glass slide, or a glass slide whose surface has been chemically modified). This may include the ability to simultaneously record the fluorescent intensity of multiple individual (i.e., single) peptide molecules distributed across the glass surface. Optical devices are commercially available that can be applied in this manner. For example, a conventional microscope equipped with total internal reflection illumination and an intensified charge-couple device (CCD) detector is available (see Braslaysky etal., 2003). Imaging with a high sensitivity CCD camera allows the instrument to simultaneously record the fluorescent intensity of multiple individual (i.e., single) peptide molecules distributed across a surface. Image collection may be performed using an image splitter that directs light through two band pass filters (one suitable for each fluorescent molecule) to be recorded as two side-by-side images on the CCD surface. Using a motorized microscope stage with automated focus control to image multiple stage positions in the flow cell allows millions of individual single peptides (or more) to be sequenced in one experiment.

[0056] The term “single-cell proteomics,” as used herein, refers to the study of the proteome of a cell. The proteome may be of a single cell. The proteome may be of a cluster of cells. The cluster of cells may be at least two cells. The cluster of cells may be 2, 5, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, or more cells. The cluster of cells may be from 2 to 10 cells. In some embodiments, the proteome of a single cell comprises proteins, peptides, or a combination thereof. In some embodiments, studying the proteome comprises determining the amino acid124928-1939-2137, v. 2sequence for at least one peptide, protein, or combination thereof. In some embodiments, the amino acid sequence is determined by sequencing peptides, proteins, or a combination thereof. The cells may be eukaryotic, prokaryotic, or archaean.

[0057] The term “support,” as used herein, refers to as a solid or semi-solid support. In some embodiments, the support is a bead or a resin.

[0058] The term “barcode” or “barcode sequence” as used herein, refers to a molecule that can be identified to distinguish a probe, a peptide, a protein, or any combination thereof from another probe, peptide, protein, or any combination thereof. In general, a barcode or barcode sequence labels a molecule or provides a molecule with an identity. The barcode can be an artificial molecule or a naturally occurring molecule. In some embodiments, at least a portion of the barcodes in a population of barcodes comprise barcodes that are different from another barcode in the population of barcodes. In some embodiments, at least about 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 95%, 99%, or more of the barcodes are different. The diversity of different barcodes in a population of barcodes can be randomly generated or non-randomly generated.

[0059] The term “nucleic acid barcode sequence,” as used herein, refers to a molecule with a particular sequence of nucleic acid. Generally, a nucleic acid barcode sequence can include one or more nucleotide sequences that can be used to identify one or more particular nucleic acids. The nucleic acid barcode sequence can be an artificial sequence or can be a naturally occurring sequence. A nucleic acid barcode sequence can comprise at least about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20 or more consecutive nucleotides. In some embodiments, a nucleic acid barcode sequence comprises at least about 10, 20, 30, 40, 50, 60, 70, 80, 90, 100 or more consecutive nucleotides. In some embodiments, at least a portion of the nucleic acid barcode sequences in a population of nucleic acids comprising barcodes is different. In some embodiments, at least about 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 95%, 99%, or more of the nucleic acid barcode sequences are different. The diversity of different nucleic acid barcode sequences in a population of nucleic acids comprising nucleic acid barcode sequences can be randomly generated or non-randomly generated.

[0060] The term “nucleic acid” as used herein generally refers to a polymeric form of nucleotides of any length, either ribonucleotides (RNA), deoxyribonucleotides (DNA) or peptide nucleic acids (PNAs), that comprise purine and pyrimidine bases, or other natural, chemically or biochemically modified, non-natural, or derivatized nucleotide bases. The backbone of the polynucleotide can comprise sugars and phosphate groups, as may typically be found in RNA or DNA, or modified or substituted sugar or phosphate groups. A134928-1939-2137, v. 2polynucleotide may comprise modified nucleotides, such as methylated nucleotides and nucleotide analogs. The sequence of nucleotides may be interrupted by non-nucleotide components. Thus, the terms nucleoside, nucleotide, deoxynucleoside and deoxynucleotide generally include analogs such as those described herein. These analogs are those molecules having some structural features in common with a naturally occurring nucleoside or nucleotide such that when incorporated into a nucleic acid or oligonucleoside sequence, they allow hybridization with a naturally occurring nucleic acid sequence in solution. Typically, these analogs are derived from naturally occurring nucleosides and nucleotides by replacing and / or modifying the base, the ribose or the phosphodiester moiety. The changes can be tailor made to stabilize or destabilize hybrid formation or enhance the specificity of hybridization with a complementary nucleic acid sequence as desired. The nucleic acid molecule may be a DNA molecule. The nucleic acid molecule may be an RNA molecule.

[0061] The sequencing reactions may comprise, for example, capillary sequencing, next generation sequencing, Sanger sequencing, sequencing by synthesis, single molecule nanopore sequencing, sequencing by ligation, sequencing by hybridization, sequencing by nanopore current restriction, or a combination thereof. Sequencing by synthesis may comprise reversible terminator sequencing, processive single molecule sequencing, sequential nucleotide flow sequencing, or a combination thereof. The single molecule sequencing may provide single molecule resolution. Sequential nucleotide flow sequencing may comprise pyrosequencing, pH-mediated sequencing, semiconductor sequencing or a combination thereof. Conducting one or more sequencing reactions may comprise whole genome sequencing or exome sequencing. The hybridization reactions may comprise, for example, fluorescent in-situ hybridization (FISH), DNA paint, multi-barcode identification (e.g., MER-FISH).

[0062] The sequencing reactions or hybridization reactions may comprise one or more capture probes or libraries of capture probes. At least one of the one or more capture probe libraries may comprise one or more capture probes to 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15 or more genomic regions. The libraries of capture probes may be at least partially complementary. The libraries of capture probes may be fully complementary. The libraries of capture probes may be at least about 5%, 10%, 15%, 20%, %, 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 70%, 80%, 90%, 95%., 97%, or more complementary.

[0063] The methods and systems disclosed herein may further comprise conducting one or more sequencing reactions or hybridization reactions on one or more capture probe free nucleic acid molecules. The methods and systems disclosed herein may further comprise conducting144928-1939-2137, v. 2one or more sequencing reactions or hybridization reactions on one or more subsets on nucleic acid molecules comprising one or more capture probe free nucleic acid molecules.

[0064] The term “label” as used herein is the introduction of a chemical group to the molecule, which generates some form of measurable signal. Such a signal may include, but is not limited to, fluorescence, visible light, mass, radiation, or a nucleic acid sequence.

[0065] As used herein, Ci-Cxincludes C1-C2, C1-C3 . . . Ci-Cx. By way of example only, a group designated as “C1-C4” indicates that there are one to four carbon atoms in the moiety, i.e. groups containing 1 carbon atom, 2 carbon atoms, 3 carbon atoms or 4 carbon atoms. Thus, by way of example only, “C1-C4 alkyl” indicates that there are one to four carbon atoms in the alkyl group, i.e., the alkyl group is selected from among methyl, ethyl, propyl, Ao-propyl, n-butyl, Ao-butyl, .scc-butyl, and / -butyl.

[0066] An “alkyl” group refers to an aliphatic hydrocarbon group. The alkyl group is branched or straight chain. In some embodiments, the “alkyl” group has 1 to 10 carbon atoms, i.e. a Ci-Cioalkyl. Whenever it appears herein, a numerical range such as “1 to 10” refers to each integer in the given range; e.g., “1 to 10 carbon atoms” means that the alkyl group consist of 1 carbon atom, 2 carbon atoms, 3 carbon atoms, 4 carbon atoms, 5 carbon atoms, 6 carbon atoms, etc., up to and including 10 carbon atoms, although the present definition also covers the occurrence of the term “alkyl” where no numerical range is designated. In some embodiments, an alkyl is a Ci-Cealkyl. In one aspect the alkyl is methyl, ethyl, propyl, iso-propyl, n-butyl, iso-butyl, secbutyl, or t-butyl. Typical alkyl groups include, but are in no way limited to, methyl, ethyl, propyl, isopropyl, butyl, isobutyl, sec-butyl, tertiary butyl, pentyl, neopentyl, or hexyl.

[0067] An “alkylene” group refers to a divalent alkyl group. Any of the above-mentioned monovalent alkyl groups may be an alkylene by abstraction of a second hydrogen atom from the alkyl. In some embodiments, an alkylene is a Ci-Cealkylene. In other embodiments, an alkylene is a Ci-C4alkylene. In certain embodiments, an alkylene comprises one to four carbon atoms (e.g., C1-C4 alkylene). In other embodiments, an alkylene comprises one to three carbon atoms (e.g., C1-C3 alkylene). In other embodiments, an alkylene comprises one to two carbon atoms (e.g., C1-C2 alkylene). In other embodiments, an alkylene comprises one carbon atom (e.g., Ci alkylene). In other embodiments, an alkylene comprises two carbon atoms (e.g., C2 alkylene). In other embodiments, an alkylene comprises two to four carbon atoms (e.g., C2-C4 alkylene). Typical alkylene groups include, but are not limited to, -CH2-, -CH(CH3)-, -C(CH3)2-, -CH2CH2-, -CH2CH(CH3)-, -CH2C(CH3)2-, -CH2CH2CH2-, -CH2CH2CH2CH2-, and the like.

[0068] The term “alkenyl” refers to a type of alkyl group in which at least one carbon-carbon double bond is present. In one embodiment, an alkenyl group has the formula -C(R)=CR.2,154928-1939-2137, v. 2wherein R refers to the remaining portions of the alkenyl group, which may be the same or different. In some embodiments, R is H or an alkyl. In some embodiments, an alkenyl is selected from ethenyl (i.e., vinyl), propenyl (i.e., allyl), butenyl, pentenyl, pentadienyl, and the like. Non-limiting examples of an alkenyl group include -CH=CH2, -C(CH3)=CH2, -CH=CHCH3, -C(CH3)=CHCH3, and -CH2CH=CH2.

[0069] The term “alkynyl” refers to a type of alkyl group in which at least one carbon-carbon triple bond is present. In one embodiment, an alkenyl group has the formula -C=C-R, wherein R refers to the remaining portions of the alkynyl group. In some embodiments, R is H or an alkyl. In some embodiments, an alkynyl is selected from ethynyl, propynyl, butynyl, pentynyl, hexynyl, and the like. Non-limiting examples of an alkynyl group include -C=CH, -C=CCH3-C=CCH2CH3, -CH2C=CH.

[0070] An “alkoxy” group refers to a (alkyl)O- group, where alkyl is as defined herein.

[0071] The term “alkylamine” refers to the -N(alkyl)xHygroup, where x is 0 and y is 2, or where x is 1 and y is 1, or where x is 2 and y is 0.

[0072] The term “aromatic” refers to a planar ring having a delocalized 7i-electron system containing 4n+2 % electrons, where n is an integer. The term “aromatic” includes both carbocyclic aryl (“aryl”, e.g., phenyl) and heterocyclic aryl (or “heteroaryl” or “heteroaromatic”) groups (e.g., pyridine). The term includes monocyclic or fused-ring polycyclic (i.e., rings which share adjacent pairs of carbon or nitrogen atoms) groups.

[0073] The term “carbocyclic” or “carbocycle” refers to a ring or ring system where the atoms forming the backbone of the ring are all carbon atoms. The term thus distinguishes carbocyclic from “heterocyclic” rings or “heterocycles” in which the ring backbone contains at least one atom which is different from carbon. In some embodiments, at least one of the two rings of a bicyclic carbocycle is aromatic. In some embodiments, both rings of a bicyclic carbocycle are aromatic. Carbocycle includes cycloalkyl and aryl.

[0074] The term “oxo” refers to C=O.

[0075] As used herein, the term “aryl” refers to an aromatic ring wherein each of the atoms forming the ring is a carbon atom. In one aspect, aryl is phenyl or a naphthyl. In some embodiments, an aryl is a phenyl. In some embodiments, an aryl is a Ce-Cioaryl. Depending on the structure, an aryl group is a monoradical or a diradical (i.e., an arylene group).

[0076] The term “cycloalkyl” refers to a monocyclic or polycyclic aliphatic, non-aromatic group, wherein each of the atoms forming the ring (i.e. skeletal atoms) is a carbon atom. In some embodiments, cycloalkyls are spirocyclic or bridged compounds. In some embodiments,164928-1939-2137, v. 2cycloalkyls are optionally fused with an aromatic ring, and the point of attachment is at a carbon that is not an aromatic ring carbon atom. Cycloalkyl groups include groups having from 3 to 10 ring atoms. In some embodiments, cycloalkyl groups are selected from among cyclopropyl, cyclobutyl, cyclopentyl, cyclopentenyl, cyclohexyl, cyclohexenyl, cycloheptyl, cyclooctyl, spiro[2.2]pentyl, norbornyl and bicyclo[l.l.l]pentyl. In some embodiments, a cycloalkyl is a Cs-Cecycloalkyl. In some embodiments, a cycloalkyl is a monocyclic cycloalkyl. Monocyclic cycloalkyls include, but are not limited to, cyclopropyl, cyclobutyl, cyclopentyl, cyclohexyl, cycloheptyl, and cyclooctyl. Polycyclic cycloalkyls include, for example, adamantyl, norbornyl (i.e., bicyclo[2.2.1]heptanyl), norbomenyl, decalinyl, 7,7-dimethyl-bicyclo[2.2.1]heptanyl, and the like

[0077] The term “halo” or, alternatively, “halogen” or “halide” means fluoro, chloro, bromo or iodo. In some embodiments, halo is fluoro, chloro, or bromo.

[0078] The term “haloalkyl” refers to an alkyl in which one or more hydrogen atoms are replaced by a halogen atom. In one aspect, a fluoroalkyl is a Ci-Cefluoroalkyl.

[0079] The term “fluoroalkyl” refers to an alkyl in which one or more hydrogen atoms are replaced by a fluorine atom. In one aspect, a fluoroalkyl is a Ci-Cefluoroalkyl. In some embodiments, a fluoroalkyl is selected from trifluoromethyl, difluoromethyl, fluoromethyl, 2,2,2-trifluoroethyl, l-fluoromethyl-2-fluoroethyl, and the like.

[0080] The term “heteroalkyl” refers to an alkyl group in which one or more skeletal atoms of the alkyl are selected from an atom other than carbon, e.g., oxygen, nitrogen (e.g., -NH-, -N(alkyl)-, sulfur, or combinations thereof. A heteroalkyl is attached to the rest of the molecule at a carbon atom of the heteroalkyl. In one aspect, a heteroalkyl is a Ci-Ceheteroalkyl.

[0081] The term “heteroalkylene” refers to a divalent heteroalkyl group.

[0082] The term "heterocycle" or “heterocyclic” refers to heteroaromatic rings (also known as heteroaryls) and heterocycloalkyl rings (also known as heteroalicyclic groups) containing one to four heteroatoms in the ring(s), where each heteroatom in the ring(s) is selected from O, S and N, wherein each heterocyclic group has from 3 to 10 atoms in its ring system, and with the proviso that any ring does not contain two adjacent O or S atoms. In some embodiments, heterocycles are monocyclic, bicyclic, polycyclic, spirocyclic or bridged compounds. Nonaromatic heterocyclic groups (also known as heterocycloalkyls) include rings having 3 to 10 atoms in its ring system and aromatic heterocyclic groups include rings having 5 to 10 atoms in its ring system. The heterocyclic groups include benzo-fused ring systems. Examples of nonaromatic heterocyclic groups are pyrrolidinyl, tetrahydrofuranyl, dihydrofuranyl, tetrahydrothienyl, oxazolidinonyl, tetrahydropyranyl, dihydropyranyl, tetrahydrothiopyranyl,174928-1939-2137, v. 2piperidinyl, morpholinyl, thiomorpholinyl, thioxanyl, piperazinyl, aziridinyl, azetidinyl, oxetanyl, thietanyl, homopiperidinyl, oxepanyl, thiepanyl, oxazepinyl, diazepinyl, thiazepinyl, 1,2,3,6-tetrahydropyridinyl, pyrrolin-2-yl, pyrrolin-3-yl, indolinyl, 2H-pyranyl, 4H-pyranyl, dioxanyl, 1,3-dioxolanyl, pyrazolinyl, dithianyl, dithiolanyl, dihydropyranyl, dihydrothienyl, dihydrofuranyl, pyrazolidinyl, imidazolinyl, imidazolidinyl, 3-azabicyclo[3.1.0]hexanyl, 3-azabicyclo[4.1.0]heptanyl, 3H-indolyl, indolin-2-onyl, isoindolin-l-onyl, isoindoline-1, 3-dionyl, 3,4-dihydroisoquinolin-l(2H)-onyl, 3,4-dihydroquinolin-2(lH)-onyl, isoindoline-1, 3-dithionyl, benzo[d]oxazol-2(3H)-onyl, lH-benzo[d]imidazol-2(3H)-onyl, benzo[d]thiazol-2(3H)-onyl, and quinolizinyl. Examples of aromatic heterocyclic groups are pyridinyl, imidazolyl, pyrimidinyl, pyrazolyl, triazolyl, pyrazinyl, tetrazolyl, furyl, thienyl, isoxazolyl, thiazolyl, oxazolyl, isothiazolyl, pyrrolyl, quinolinyl, isoquinolinyl, indolyl, benzimidazolyl, benzofuranyl, cinnolinyl, indazolyl, indolizinyl, phthalazinyl, pyridazinyl, triazinyl, isoindolyl, pteridinyl, purinyl, oxadiazolyl, thiadiazolyl, furazanyl, benzofurazanyl, benzothiophenyl, benzothiazolyl, benzoxazolyl, quinazolinyl, quinoxalinyl, naphthyridinyl, and furopyridinyl. The foregoing groups are either C-attached (or C-linked) or TV-attached where such is possible. For instance, a group derived from pyrrole includes both pyrrol-l-yl (TV-attached) or pyrrol-3-yl (C-attached). Further, a group derived from imidazole includes imidazol-l-yl or imidazol-3-yl (both TV-attached) or imidazol-2-yl, imidazol-4-yl or imidazol-5-yl (all C-attached). The heterocyclic groups include benzo-fused ring systems. Non-aromatic heterocycles are optionally substituted with one or two oxo (=0) moieties, such as pyrrolidin-2-one. In some embodiments, at least one of the two rings of a bicyclic heterocycle is aromatic. In some embodiments, both rings of a bicyclic heterocycle are aromatic.

[0083] The terms “heteroaryl” or, alternatively, “heteroaromatic” refers to an aryl group that includes one or more ring heteroatoms selected from nitrogen, oxygen and sulfur. Illustrative examples of heteroaryl groups include monocyclic heteroaryls and bicyclic heteroaryls. Monocyclic heteroaryls include pyridinyl, imidazolyl, pyrimidinyl, pyrazolyl, triazolyl, pyrazinyl, tetrazolyl, furyl, thienyl, isoxazolyl, thiazolyl, oxazolyl, isothiazolyl, pyrrolyl, pyridazinyl, triazinyl, oxadiazolyl, thiadiazolyl, and furazanyl. Bicyclic heteroaryls include indolizine, indole, benzofuran, benzothiophene, indazole, benzimidazole, purine, quinolizine, quinoline, isoquinoline, cinnoline, phthalazine, quinazoline, quinoxaline, 1,8-naphthyridine, and pteridine. In some embodiments, a heteroaryl contains 0-4 N atoms in the ring. In some embodiments, a heteroaryl contains 1-4 N atoms in the ring. In some embodiments, a heteroaryl contains 0-4 N atoms, 0-1 0 atoms, and 0-1 S atoms in the ring. In some embodiments, a heteroaryl contains 1-4 N atoms, 0-1 O atoms, and 0-1 S atoms in the ring. In some184928-1939-2137, v. 2embodiments, heteroaryl is a Ci-Cgheteroaryl. In some embodiments, monocyclic heteroaryl is a Ci-Csheteroaryl. In some embodiments, monocyclic heteroaryl is a 5-membered or 6-membered heteroaryl. In some embodiments, bicyclic heteroaryl is a Ce-Cgheteroaryl.

[0084] A “heterocycloalkyl” or “heteroalicyclic” group refers to a cycloalkyl group that includes at least one heteroatom selected from nitrogen, oxygen and sulfur. In some embodiments, a heterocycloalkyl is fused with an aryl or heteroaryl. In some embodiments, the heterocycloalkyl is oxazolidinonyl, pyrrolidinyl, tetrahydrofuranyl, tetrahydrothienyl, tetrahydropyranyl, tetrahydrothiopyranyl, piperidinyl, morpholinyl, thiomorpholinyl, piperazinyl, piperidin-2-onyl, pyrrolidine-2, 5-dithionyl, pyrrolidine-2, 5-dionyl, pyrrolidinonyl, imidazolidinyl, imidazolidin-2-onyl, or thiazolidin-2-onyl. The term heteroalicyclic also includes all ring forms of the carbohydrates, including but not limited to the monosaccharides, the disaccharides and the oligosaccharides. In one aspect, a heterocycloalkyl is a C2-Cioheterocycloalkyl. In another aspect, a heterocycloalkyl is a C4-Cioheterocycloalkyl. In some embodiments, a heterocycloalkyl contains 0-2 N atoms in the ring. In some embodiments, a heterocycloalkyl contains 0-2 N atoms, 0-2 O atoms and 0-1 S atoms in the ring.

[0085] The term “bond” or “single bond” refers to a chemical bond between two atoms, or two moi eties when the atoms joined by the bond are considered to be part of larger substructure. In one aspect, when a group described herein is a bond, the referenced group is absent thereby allowing a bond to be formed between the remaining identified groups.

[0086] The term “moiety” refers to a specific segment or functional group of a molecule. Chemical moieties are often recognized chemical entities embedded in or appended to a molecule.

[0087] As used herein, the term “derivative” generally refers to a chemical species that is derivatized by substitution by one or more substituents. Exemplary substituents include, but are not limited to D, halogen, -CN, -NH2, -NH(alkyl), -N(alkyl)2, -OH, -CO2H, -CO2alkyl, -C(=O)NH2, -C(=O)NH(alkyl), -C(=O)N(alkyl)2, -S(=O)2NH2, -S(=O)2NH(alkyl), -S(=O)2N(alkyl)2, -CH2CO2H, -CH2CO2alkyl, -CH2C(=O)NH2, -CH2C(=O)NH(alkyl), -CH2C(=O)N(alkyl)2, -CH2S(=O)2NH2, - CH2S(=O)2NH(alkyl), - CH2S(=O)2N(alkyl)2, alkyl, alkenyl, alkynyl, cycloalkyl, fluoroalkyl, heteroalkyl, alkoxy, fluoroalkoxy, heterocycloalkyl, aryl, heteroaryl, aryloxy, alkylthio, arylthio, alkylsulfoxide, arylsulfoxide, alkylsulfone, and arylsulfone.

[0088] The term “optionally substituted” or “substituted” means that the referenced group is optionally substituted with one or more additional group(s). In some other embodiments,194928-1939-2137, v. 2optional substituents are individually and independently selected from D, halogen, -CN, -NH2, -NH(alkyl), -N(alkyl)2, -OH, -CO2H, -CO2alkyl, -C(=O)NH2, -C(=O)NH(alkyl), -C(=O)N(alkyl)2, -S(=O)2NH2, -S(=O)2NH(alkyl), -S(=O)2N(alkyl)2, -CH2CO2H, -CH2CO2alkyl, -CH2C(=O)NH2, -CH2C(=O)NH(alkyl), -CH2C(=O)N(alkyl)2, CH2S(=O)2NH2, - CH2S(=O)2NH(alkyl), - CH2S(=O)2N(alkyl)2, alkyl, alkenyl, alkynyl, cycloalkyl, fluoroalkyl, heteroalkyl, alkoxy, fluoroalkoxy, heterocycloalkyl, aryl, heteroaryl, aryloxy, alkylthio, arylthio, alkylsulfoxide, arylsulfoxide, alkylsulfone, and arylsulfone. The term “optionally substituted” or “substituted” means that the referenced group is optionally substituted with one or more additional group(s) individually and independently selected from D, halogen, -CN, -NH2, -NH(alkyl), -N(alkyl)2, -OH, -CO2H, -CO2alkyl, -C(=O)NH2, -C(=O)NH(alkyl), -C(=O)N(alkyl)2, -S(=O)2NH2, -S(=O)2NH(alkyl), -S(=O)2N(alkyl)2, alkyl, cycloalkyl, fluoroalkyl, heteroalkyl, alkoxy, fluoroalkoxy, heterocycloalkyl, aryl, heteroaryl, aryloxy, alkylthio, arylthio, alkylsulfoxide, arylsulfoxide, alkylsulfone, and arylsulfone. In some other embodiments, optional substituents are independently selected from D, halogen, -CN, -NH2, -NH(CH3), -N(CH3)2, -OH, -CO2H, -CO2(Ci-C4alkyl), -C(=O)NH2, -C(=O)NH(Ci-C4alkyl), -C(=O)N(Ci-C4alkyl)2, -S(=O)2NH2, -S(=O)2NH(Ci-C4alkyl), -S(=O)2N(CI-C4alkyl)2, Ci-C4alkyl, C3-C6cycloalkyl, Ci-C4fluoroalkyl, Ci-C4heteroalkyl, Ci-C4alkoxy, Ci-C4fluoroalkoxy, -SCi-C4alkyl, -S(=O)Ci-C4alkyl, and -S(=O)2Ci-C4alkyl. In some embodiments, optional substituents are independently selected from D, halogen, -CN, -NH2, -OH, -NH(CH3), -N(CH3)2, -CH3, -CH2CH3, -CF3, -OCH3, and -OCF3. In some embodiments, substituted groups are substituted with one or two of the preceding groups. In some embodiments, substituted groups are substituted with one of the preceding groups. In some embodiments, an optional substituent on an aliphatic carbon atom (acyclic or cyclic) includes oxo (=0).

[0089] As described herein, the term “handle” refers to a molecule that can couple to the C-terminal carboxylic acid of a protein or peptide. A handle may comprise a backbone (e.g., alkylene, polyethylene glycol, and amide groups), a nucleophile (e.g., amine or thiol), an electrophile (e.g., Michael acceptor), a detection unit (e.g., fluorophore, nucleic acid oligomer, or charged group), a functionalization unit (e.g., biotin, azide, alkyne, thiol, alkene, carboxylic acid, or amine), or any combination thereof. A handle may comprise at least one linker.

[0090] A “linker”, as described herein, couples at least two molecules. In some embodiments, a linker couples at least two molecules directly or indirectly. A linker may be a bifunctional molecule for labeling amino acid side chains. One end of the molecule may comprise an amino acid specific functional group (e.g., iodoacetamide for labeling thiol residues on cysteines) and204928-1939-2137, v. 2the other end may be a different functional group amenable for labeling. If no reporter is required to be attached, then the functional group may be an inert group (e.g., alkane). The reporter end of the tag molecule may be a fluorophore. A tag may comprise at least one charged molecule that can produce a distinct signal (e.g., fluorescent or electric).

[0091] The term “reporter” or “tag” refers to a molecule that produces an identifiable signal. Examples of a reporter include fluorophores (e.g., a cluster of fluorophores), DNA molecules that can be hybridized, or molecules that produces a distinct electrical signal state.

[0092] The term “reactive agent,” as used herein, generally refers to a chemical or biological agent that reacts with a peptide or protein. The “reactive agent” may react selectively with the C -terminal amino acid of a peptide or protein.

[0093] The term, “internal amino acid residue,” as used herein, generally refers to an amino acid residue between a C-terminal amino acid residue or an N-terminal amino acid residue of a peptide or protein.

[0094] The term, “nucleophile,” as used herein, generally refers to a chemical species (e.g., a first atom) that donates an electron pair to form a chemical bond with another chemical species (e.g., a second atom). Examples of atoms that can act as nucleophiles are halogens (e.g., fluoride, chloride, bromine, iodine), oxygen, sulfur, nitrogen, and carbon. Examples of nucleophiles include, but are not limited to, electron rich chemical species, negatively charged chemical species, amines, alcohols, thiols, sulfides, alkynes, alkenes, carboxylic acids, nitriles, water, azides, nitrites, hydroxylamines, hydrazines, and carbazides. The term, “electrophile,” as used herein, generally refers to a chemical species (e.g., a first atom) that accepts an electron pair to form a chemical bond with another chemical species (e.g., a second atom). Examples of atoms that can act as electrophiles are hydrogen, halogens, sulfur, and carbon. Examples of electrophiles include, but are not limited to, electron poor chemical species, positively charged chemical species, alkenes, dienes, acylates, acrylamides, cyanates, carboxylic acids, amides, esters, sulfones, aldehydes, and conjugated systems (e.g., a Michael acceptor or a conjugated aromatic system). For example, a nucleophile can react with an electrophile to form a chemical bond between the nucleophile and the electrophile.

[0095] As used herein, the term “electron withdrawing group” or “EWG,” which, as will be appreciated by those of ordinary skill in the art, is a group that has greater electronegativity than the carbon atom to which it is bound. In general, this means that the electron withdrawing group is more electron-withdrawing than a hydrogen atom. Accordingly, representative electron-withdrawing substituents include, without limitation, oxo, halo, cyano, Ci-C2ohaloalkyl (preferably Ci-Cnhaloalkyl, more preferably Ci-Ce haloalkyl), Ci-214928-1939-2137, v. 2C20 alkyl sulfonyl (preferably C1-C12 alkyl sulfonyl, more preferably Ci-Ce alkyl sulfonyl), C5-C24 aryl sulfonyl (preferably C5-C14 aryl sulfonyl), carboxyl, C5-C24 aryloxy (preferably C5-C14 aryloxy), C1-C20 alkoxy (preferably C1-C12 alkoxy, more preferably Ci-Ce alkoxy), Ce-C24 aryloxycarbonyl (preferably O,-C 14 aryl oxy carbonyl), C1-C20 alkoxycarbonyl (preferably C1-C12 alkoxycarbonyl, more preferably Ci-Ce alkoxycarbonyl), C6-C24 arylcarbonyl (preferably O,-C 14 aryl carbonyl), C1-C20 alkylcarbonyl (preferably C1-C12 alkylcarbonyl, more preferably Ci-Ce alkylcarbonyl), formyl, nitro, quaternary amino (including alkyl- and arylsubstituted amino), sulfhydryl, C1-C20 alkylthio (preferably C1-C12 alkylthio, more preferably Ci-Ce alkylthio), C5-C24 arylthio (preferably C5-C14 arylthio), hydroxyl, C5-C24 aryl (preferably Cs-Cwaryl), C1-C20 alkyl (preferably C1-C12 alkyl, more preferably Ci-Ce alkyl), any of which may be substituted and / or heteroatom-containing. Cyano groups are of particular interest herein.

[0096] As used herein, the term “Meldrum’s acid” refers to 2, 2-dimethyl- 1,3 -dioxane-4, 6-O^Odi one which has the structure:

[0097] As used herein, the term “EVA reagent” refers to 5-[Bis(methylthio)methylene]-2,2-dimethyl- 1,3 -dioxane-4, 6-dione which has the structure:I. Peptide Sequencing Overview

[0098] Peptide sequencing is essential for understanding protein structures, developing therapeutics, and advancing biotechnology. While Edman degradation has been the gold standard, its limitations with acid-sensitive samples, such as those containing oligonucleotides, restrict its application in contexts like fluorosequencing. This study introduces a mild method for N-terminal peptide sequencing that avoids acids or strong bases. By conjugating the N-terminal to a reagent and treating it with hydrazine, the N-terminal amino acid is selectively cleaved with high yield in under an hour. The cleaved derivative is UV-active, enabling precise identification via LC-MS.Edman Degradation Chemistry

[0099] Edman degradation is currently the “gold standard” method for N-terminal peptide sequencing. Edman degradation involves the use of a phenylisothiocyanate reagent, which224928-1939-2137, v. 2couples to N-terminal amino acids and / or lysine residues under basic conditions. This coupling forms a 5-membered intermediate ring. The Edman degradation method uses organic solvents which can often be harsh including bases such as pyridine, triethylamine, N, N-diisopropylethylamine, and combinations thereof and acids such as trifluoroacetic acid (which is the most common) and methanolic hydrochloric acid. The reaction steps of Edman degradation are as follows:

[0100] Step 1 (Base): The peptide or protein for sequencing is incubated with base (such as, pyridine, triethylamine, N, N-diisopropylethylamine, and combinations thereof) and an Edman reagent (i.e., a phenylisothiocyanate reagent). The Edman reagent forms a 5-membered intermediate with the peptide or protein to be sequenced (“Edman Intermediate I”).

[0101] Step 2 (First Wash): After base incubation, the Edman Intermediate I is washed with an inert solvent (e.g., methanol, heptane, etc.)

[0102] Step 3 (Acid): After the first wash, the Edman Intermediate is incubated with an acid (e.g., trifluoroacetic acid or methanolic hydrochloric acid) to form Edman Intermediate II and retained peptide or protein.

[0103] Step 4 (Second Wash or Release): After acid incubation, the Edman Intermediate II and retained peptide or protein are subjected to a second wash with an inert solvent (e.g., methanol, heptane, etc.). One can capture the released N-terminal amino acid (Edman Intermediate II) to analyze by HPLC (standard method for pure protein sequencing). One can also image the retained peptide for methods such as fluorosequencing.

[0104] Exemplary incubation conditions for each of the aforementioned steps are listed below:

[0105] Step 1 (Base): 40-60 °C for 20 mins

[0106] Step 2 (First Wash): 5 mins of solvents; Performed 2x

[0107] Step 3 (Acid): 40-60 °C for 5 mins

[0108] Step 4 (Second Wash or Release): 5 mins of solvents; Performed 2x“Ava” Chemistry

[0109] One aspect of this disclosure provides a new method for N-terminal peptide sequencing. This method, unlike Edman degradation, is suitable for N-terminal peptide or protein sequencing in mild conditions. As used herein, “mild conditions” refer to conditions which utilize a water-based solution, wherein the pH is in the neutral range (about pH 6-8) and the temperature is between about °25-°40 C.EVA reagent and derivatives234928-1939-2137, v. 2

[0110] Ava chemistry involves the use of an reagent with the general structure shown in Compound (1) in FIG.2. One such reagent is the EVA reagent (5-[Bis(methylthio)methylene]-2, 2-dimethyl- 1,3 -dioxane-4, 6-dione, see FIG. 1). The EVA reagent reacts with nucleophiles such as amines (e.g., N-terminal amines, lysine resides) and thiols (e.g., cysteine residues). The synthesis of the EVA reagent is described in Example 1.[OHl] Derivatives of the EVA reagent are also suitable for use in the method of the present disclosure. Exemplary EVA reagent derivatives are disclosed in FIG. 2. Exemplary syntheses of EVA reagent derivatives are disclosed in Example 3.Process

[0112] The EVA reagent and derivatives thereof can be used in the methods described herein for N-terminal peptide or protein sequencing in mild conditions. A variety of peptides and proteins are envisioned to be sequenced by this method, including but not limited to natural peptides, unnatural peptides, peptide derivatives containing amide bonds and alpha-amino acids, or combinations thereof. The EVA reagent or derivative thereof may be coupled to a peptide or protein and subsequently cleaved from the peptide or protein in mild conditions. The mild conditions of coupling and cleavage occur between about pH 6-8, in an aqueous buffer or a mix of aqueous and polar solvents (e.g., acetonitrile, methanol, ethanol, dimethylformamide, etc.) and at moderate temperatures, between about 25°-40 °C. The cleavage is facilitated by a diamine (e.g., hydrazine and hydrazine derivatives).Peptide processing

[0113] Nucleophilic side-chains on peptides such as cysteine and / or lysine may be protected prior to the N-terminal peptide sequencing method. Cysteine protection can be performed by alkylation using reagents such as by use of iodoacetamide, chloroacetamide, etc. Lysine protection can be performed by succinimidyl-ester. In some embodiments, the lysine side chain is not protected prior to the N-terminal peptide sequencing method. Either of both of the cysteine residue and lysine residue can be functionalized with a fluorophore. In some embodiments, one or more of the amino acids on the peptide or protein are functionalized. The C -terminus of the peptide can be functionalized with an alkyne moiety. In some embodiments, the peptides are immobilized on a solid support. In some embodiments, the support may be a glass slide. In some embodiments, the peptide or protein may be immobilized through its C-terminus.EVA reagent coupling

[0114] In some embodiments, the EVA reagent or derivative reacts with the N-terminal amine on the peptide. The reaction is carrier out in an aqueous buffer or in an aqueous buffer / organic244928-1939-2137, v. 2buffer. In some embodiments, the aqueous buffer is a sodium phosphate buffer at about pH 6-8 and 0.1 M. In some embodiments, the pH is about 7.6. In some embodiments, the organic buffer is acetonitrile or dimethylformamide (DMF). In some embodiments, the aqueous buffer / organic buffer is prepared in a 1:1 mixture.

[0115] In some embodiments, the reaction is sparged with nitrogen. In some embodiments, the reaction is sparged with nitrogen for about 5 minutes. In some embodiments, the reaction does not require nitrogen. In some embodiments, the reaction can be open to ambient air. In some embodiments, the reaction is done in a reaction vial and the reaction vial is opened to ambient air. In some embodiments, opening the reaction vial to ambient air allows the release of methyl mercaptan.

[0116] In some embodiments, the EVA reagent or derivative thereof is mixed with the one or more peptides or proteins with at least about 1:1 equivalents. In some embodiments, more equivalents of the EVA reagent or derivative are used. In some embodiments, 2x, 5x, lOx, or lOOx the EVA reagent or derivative is used. In some embodiments, the one or more peptides or proteins are the same. In some embodiments, the one or more peptides or proteins are different from each other. In preferred embodiments, the reaction conditions are 15 minutes at 37 °C.

[0117] In some embodiments, there is a washing step. In some embodiments, there is not a washing step. If peptides are on solid-support, coupling EVA reagent can be removed through use of repeated aqueous buffer, organic solvents or combinations of these solutions.

[0118] In some embodiments, there is a separation step. In some embodiments, there is not a separation step. The excess EVA reagent can be separated from the coupled peptide through filtration methods, such as by using C-18 tips.

[0119] The EVA-coupled peptide is subjected to cleavage conditions that remove the 1st N-terminal amino acid. In some embodiments, the cleavage reagent used is a diamine reagent. Exemplary reagents include but are not limited to hydrazine, hydrazine variants, hydroxylamines, diamine conjugated reagents, and ethylene diamine. The cleavage reagent can be added at 5x, lOx, and 12x equivalencies compared to the peptide or protein. In some embodiments, 12x equivalency is used. The cleavage reaction is carried out in aqueous buffer or aqueous buffer / organic solvents. In some embodiments, the aqueous buffer is a sodium phosphate buffer at about pH 6-8 and about 0.1 M. In some embodiments, the pH is about 7.5. In some embodiments, the organic solvent is acetonitrile. In some embodiments, the ratio of aqueous buffer / organic solvent is 3:1. In preferred embodiments, the cleavage conditions are 30 minutes at 37 °C.Utility254928-1939-2137, v. 2

[0120] Ava chemistry can be used as a substitute for Edman degradation methods used for peptide and protein characterization. It may be used in place of traditional Edman degradation process, where the liberated residue from a pure protein or peptide is analyzed by a LC-MS. It may also be used in modern next-generation protein sequencing platforms, where peptides and proteins are analyzed through degradative chemistries. Exemplary examples of next-generation protein sequencing include but are not limited to fluorosequencing and methods using DNA (e.g., ProtSeq, Glyphic patents, reverse translation etc).

[0121] EVA reagents or derivatives thereof (e.g., a compound in FIG. 1 or FIG. 2) may be used as a peptide and protein capture reagent. In some embodiments, the EVA reagent or derivative thereof may be coupled to a solid support (e.g., magnetic beads or resins). Peptides or proteins are conjugated through covalent bonds from solution. In some embodiments, the EVA reagent or derivative thereof is first coupled to an amine (e.g., anN-terminal amine, lysine residue) or thiol (e.g., cysteine residue) on a peptide or protein. Then, the EVA-peptide or protein conjugate is conjugated to the solid support. The EVA-peptide or protein conjugate can be released via a diamine (e.g., a hydrazine, hydrazine variants, hydroxylamines, diamine conjugated reagents, and ethylene diamine). In some embodiments, the cleavage is efficient. As used herein, “efficient cleavage” is cleavage which is greater than or equal to 50%.II. N-Terminal Sequencing Methods

[0122] One aspect of the disclosure provides methods of functionalizing peptides or proteins. The methods may be useful, for example, in N-terminal sequencing methods or methods comprising capture of peptides or proteins to a solid support. Exemplary methods are described in the following section.

[0123] One aspect of the disclosure provides a method of functionalizing a peptide or protein having an N-terminus, the method comprising a coupling step comprising reacting the N-terminus of the peptide or protein with a compound of Formula (I):or a salt thereof, wherein:R1and R2are each, independently, an electron withdrawing group; orR1and R2are taken together with the carbon atom to which they are attached to form a 3-10 membered heterocyclyl or C3-10 carbocyclyl, wherein the 3-10 membered heterocyclyl or264928-1939-2137, v. 2C3-10 carbocyclyl is substituted with one or more electron withdrawing groups and optionally further substituted,to form a compound of Formula (II):Peptide or Protein(II).

[0124] In some embodiments, the terminus is the N-terminus of the peptide or protein. In some embodiments, the compound of Formula (I) is conjugated to a solid support. In some embodiments, the compound of Formula (II) is conjugated to a solid support.

[0125] In some embodiments, R1and R2are each, independently, an electron withdrawing group.

[0126] In some embodiments, R1and R2are identical.

[0127] In some embodiments, the electron withdrawing group is cyano or -CORa, wherein Rais selected from the group consisting of-Ci-6 alkyl, -Ci-6 alkoxy, and optionally substituted -Ce-io aryl.

[0128] In some embodiments, Rais selected from the group consisting of methyl, methoxy, and phenyl.

[0129] In some embodiments, R1and R2are selected from the group consisting of cyano, -C(O)C6H5, -C(O)CH3, and -C(O)(OCH3).

[0130] In some embodiments, R1and R2are taken together with the carbon atom to which they are attached to form a 3-10 membered heterocyclyl or C3-io carbocyclyl, wherein the 3-10 membered heterocyclyl or C3-io carbocyclyl is substituted with one or more electron withdrawing groups and optionally further substituted.

[0131] In some embodiments, the 3-10 membered heterocyclyl or C3-io carbocyclyl is further substituted with one or more R3, wherein R3is selected from the group consisting of-Ci-6 alkyl, -C2-6 alkenyl, and -C3-io carbocyclyl. In some embodiments, R3is -C1-6 alkyl.

[0132] In some embodiments, the one or more electron withdrawing groups is oxo.

[0133] In some embodiments, the 3-10 membered heterocyclyl or C3-io carbocyclyl is derived from Meldrum’s acid.274928-1939-2137, v. 2

[0134] In some embodiments, the compound of Formula (I) is selected from the groupIn some embodiments, the derivative is substituted with one or more Rb, wherein Rbis selected from the group consisting of halo, cyano, -Ci-6 alkyl, -C2-6 alkenyl, and -C3-10 carbocyclyl.

[0135] In some embodiments, the compound is selected from the group consisting ofor a derivative thereof. In some embodiments, the derivative is substituted with one or more Rb, wherein Rbis selected from the group consisting of halo, cyano, -C1-6 alkyl, -C2-6 alkenyl, and -C3-10 carbocyclyl.

[0137] In some embodiments, the compound of Formula (I) is an EVA reagent:

[0138] In some embodiments, the coupling step occurs in an aqueous buffer or in a mixture of an aqueous buffer and an organic solvent.

[0139] In some embodiments, the aqueous buffer is a sodium phosphate buffer.284928-1939-2137, v. 2

[0140] In some embodiments, the pH of the sodium phosphate buffer is between about 6.0 about 8.0 and the concentration of the sodium phosphate buffer is about 0.1 M.

[0141] In some embodiments, the pH is about 7.6.

[0142] In some embodiments, the organic solvent is acetonitrile or dimethylformamide.

[0143] In some embodiments, the ratio of the mixture of the aqueous buffer and the organic solvent is 1:1.

[0144] In some embodiments, the reaction occurs in a reaction vessel that is sparged with an inert gas.

[0145] In some embodiments, the inert gas is nitrogen.

[0146] In some embodiments, the ratio of the compound of Formula (I) and the peptide or protein is about 1:1.

[0147] In some embodiments, the compound of Formula (I) is present in excess of the peptide or protein.

[0148] In some embodiments, the ratio of the compound of Formula (I) and the peptide or protein is selected from the group consisting of about 2:1, about 5:1, about 10:1, and about 100:1.

[0149] In some embodiments, the reaction time of the coupling step is less than about an hour.

[0150] In some embodiments, the reaction time of the coupling step is about 15 minutes.

[0151] In some embodiments, the coupling step is performed at about 37 °C.

[0152] In some embodiments, the method further comprises a step of washing to remove unreacted or excess amount of the compound of Formula (I) using an aqueous buffer, an organic solvent, or a mixture thereof.

[0153] In some embodiments, the method further comprises a step of removing unreacted or excess amount of the compound of Formula (I) by filtration.

[0154] In some embodiments, the filtration comprises using C-18 tips.

[0155] In some embodiments, the method further comprises a step of cleaving the N-terminal amino acid residue of the peptide or protein to form a truncated peptide or protein.

[0156] In some embodiments, the step of cleaving comprises contacting a compound of Formula (II):R1R2Peptide or ProteinH(II),or a salt thereof,with a diamine reagent of Formula (III):294928-1939-2137, v. 2H2N Linker NH2(III),or a salt thereof,to form a compound of Formula (IV):or a salt thereof,wherein Linker is selected from the group consisting of a bond, optionally substituted Ci-6 alkylene, and optionally substituted Ci-6 heteroalkylene.

[0157] In some embodiments, Formula (IV) is represented by the structure:, whereinR4is the side chain of the N-terminal amino acid residue of the peptide or protein; R5is the side chain of the amino acid residue directly adjacent to the N-terminal amino acid residue of the peptide or protein;AAis an amino acid residue; andn is an integer number representing the total number of amino acid residues in the peptide or protein minus the first two N-terminal amino acid residues. For example, if the total number of amino acids in a peptide or protein is 5, n is 3. In another example, if the total number of amino acids in a peptide or protein is 100, n is 98.

[0158] In some embodiments, the method further comprises intramolecularly reacting a compound of Formula (IV):or a salt thereof to form the truncated peptide or protein and a compound of Formula (V):304928-1939-2137, v. 2or a salt thereof; andR4is the side chain of the N-terminal amino acid residue of the peptide or protein.

[0159] In some embodiments, the diamine reagent is selected from the group consisting of hydrazine and ethylenediamine. In some embodiments, the diamine reagent is hydrazine. In some embodiments, the diamine reagent is ethylenediamine.

[0160] In some embodiments, the step of cleaving occurs in an aqueous buffer or in a mixture of an aqueous buffer and an organic solvent.

[0161] In some embodiments, the pH of the sodium phosphate buffer is between about 6.0 and 8.0 and the concentration of the sodium phosphate buffer is about 0.1 M.

[0162] In some embodiments, the pH is about 7.5.

[0163] In some embodiments, the organic solvent is acetonitrile.

[0164] In some embodiments, the ratio of the mixture of the aqueous buffer and the organic solvent is 1:3.

[0165] In some embodiments, the ratio of the diamine reagent and the compound of Formula (II) is about 1:1.

[0166] In some embodiments, the diamine reagent is present in excess of the compound of Formula (II).

[0167] In some embodiments, the ratio of diamine reagent and the compound of Formula (II) is selected from the group consisting of about 5:1, about 10:1, and about 12:1.

[0168] In some embodiments, the ratio is about 12:1.

[0169] In some embodiments, the reaction time of the step of cleaving is less than about an hour.

[0170] In some embodiments, the reaction time of the step of cleaving is about 30 minutes.

[0171] In some embodiments, the step of cleaving is performed at about 37 °C.

[0172] In some embodiments, the peptide or protein comprises natural amino acids or derivatives thereof.

[0173] In some embodiments, the peptide or protein comprises unnatural amino acids, alphaamino acids or combinations or derivatives thereof.

[0174] In some embodiments, the method further comprises a step of sequencing the peptide or protein.

[0175] In some embodiments, the sequencing is N-terminal sequencing.

[0176] In some embodiments, the method further comprises a step of detecting the presence of the truncated peptide or protein or the compound of Formula (V).314928-1939-2137, v. 2

[0177] In some embodiments, the step of detecting comprises detecting the compound of Formula (V).

[0178] In some embodiments, the step of detecting comprises performing mass spectrometry.

[0179] In some embodiments, the mass spectrometry comprises LC-MS.

[0180] In some embodiments, the step of detecting comprises detecting the truncated peptide or protein.

[0181] In some embodiments, the step of detecting comprises fluorosequencing.

[0182] In some embodiments, the truncated peptide or protein is attached to an oligonucleotide barcode, and the step of detecting comprises detecting the oligonucleotide barcode. In some embodiments, the truncated peptide or protein is attached to an oligonucleotide barcode, and the step of detecting comprises sequencing the oligonucleotide barcode. In some embodiments, the truncated peptide or protein is attached to an oligonucleotide barcode, and the step of detecting comprises detecting the oligonucleotide barcode. In some embodiments, the truncated peptide or protein is attached to an oligonucleotide barcode, and the step of detecting comprises detecting or sequencing the oligonucleotide barcode. In some embodiments, the oligonucleotide is DNA. In some embodiments, the oligonucleotide is RNA.

[0183] Also provided herein is a method of sequencing a plurality of peptide or proteins comprising sequencing a plurality of peptide or proteins functionalized by the method provided herein. In some embodiments, the peptides or proteins are unique from each other.

[0184] In some embodiments, the solid support is a magnetic bead or resin.

[0185] In another aspect, provided herein is a method of N-terminal peptide sequencing of a peptide or protein having an N-terminus, the method comprising the steps ofreacting the N-terminus of the peptide or protein with an EVA reagent:salt thereof,to form a compound of Formula (Il-a):324928-1939-2137, v. 2cleaving the N-terminal amino acid residue of the peptide or protein to form a truncated peptide or protein by contacting a compound of Formula (Il-a) with a diamine reagent of Formula (III):H2N LinkerNH2(III),or a salt thereof,to form a compound of Formula (IV-a):or a salt thereof, wherein the compound of Formula (IV-a) or salt thereof intramolecularly reacts to form the truncated peptide or protein and a compound of Formula (V-a):or a salt thereof, whereinLinker is a bond or -C2H4-; andR4is the side chain of the N-terminal amino acid residue of the peptide or protein.

[0186] In some embodiments, Formula (IV-a) is represented by the structure:R4is the side chain of the N-terminal amino acid residue of the peptide or protein; R5is the side chain of the amino acid residue directly adjacent to the N-terminal amino acid residue of the peptide or protein;AAis an amino acid residue; and334928-1939-2137, v. 2n is an integer number representing the total number of amino acid residues in the peptide or protein minus the first two N-terminal amino acid residues. For example, if the total number of amino acids in a peptide or protein is 5, n is 3. In another example, if the total number of amino acids in a peptide or protein is 100, n is 98.

[0187] In some embodiments, the diamine reagent is hydrazine. In some embodiments, the diamine reagent is ethylenediamine.

[0188] In some embodiments, the method further comprises a step of washing to remove unreacted or excess amount of the EVA reagent using an aqueous buffer, an organic solvent, or a mixture thereof.

[0189] In some embodiments, the method further comprises a step of removing unreacted or excess amount of the EVA reagent by filtration.

[0190] In some embodiments, the method further comprises a step of detecting the presence of the truncated peptide or protein or the compound of Formula (V-a).

[0191] In some embodiments, the step of detecting comprises detecting the compound of Formula (V-a).

[0192] In some embodiments, the step of detecting comprises performing mass spectrometry.

[0193] In some embodiments, the mass spectrometry comprises LC-MS.

[0194] In some embodiments, the step of detecting comprises detecting the truncated peptide or protein.

[0195] In some embodiments, the step of detecting comprises fluorosequencing.

[0196] In some embodiments, the truncated peptide or protein is attached to an oligonucleotide barcode, and the step of detecting comprises detecting the oligonucleotide barcode. In some embodiments, the truncated peptide or protein is attached to an oligonucleotide barcode, and the step of detecting comprises sequencing the oligonucleotide barcode. In some embodiments, the truncated peptide or protein is attached to an oligonucleotide barcode, and the step of detecting comprises detecting the oligonucleotide barcode. In some embodiments, the truncated peptide or protein is attached to an oligonucleotide barcode, and the step of detecting comprises detecting or sequencing the oligonucleotide barcode. In some embodiments, the oligonucleotide is DNA. In some embodiments, the oligonucleotide is RNA.

[0197] In another aspect, the disclosure provides a kit comprising a compound of Formula (I), or a salt thereof, for use in a method described herein.

[0198] In another aspect, the disclosure provides a kit comprising the EVA reagent, or a salt thereof, for use in a method described herein.344928-1939-2137, v. 2III. Compounds and Compositions

[0199] Another aspect of the disclosure provides compounds and compositions for use in or generated by the methods described herein.

[0200] One aspect of the disclosure provides a compound of Formula (IV):R1R2Peptide or Protein(IV),or a salt thereof, whereinR1and R2are each, independently, an electron withdrawing group; orR1and R2are taken together with the carbon atom to which they are attached to form a 3-10 membered heterocyclyl or C3-10 carbocyclyl, wherein the 3-10 membered heterocyclyl or C3-10 carbocyclyl is substituted with one or more electron withdrawing groups and optionally further substituted; andLinker is selected from the group consisting of a bond, optionally substituted C1-6 alkylene, and optionally substituted C1-6 heteroalkylene.

[0201] In some embodiments, Formula (IV) is represented by the structure:, whereinR4is the side chain of the N-terminal amino acid residue of the peptide or protein; R5is the side chain of the amino acid residue directly adjacent to the N-terminal amino acid residue of the peptide or protein;AAis an amino acid residue; andn is an integer number representing the total number of amino acid residues in the peptide or protein minus the first two N-terminal amino acid residues. For example, if the total number of amino acids in a peptide or protein is 5, n is 3. In another example, if the total number of amino acids in a peptide or protein is 100, n is 98.

[0202] In some embodiments, R1and R2are each, independently, an electron withdrawing group.

[0203] In some embodiments, R1and R2are identical.

[0204] In some embodiments, the electron withdrawing group is cyano or -CORa, wherein Rais selected from the group consisting of-Ci-6 alkyl, -Ci-6 alkoxy, and -Ce-io aryl.4928-1939-2137, v. 2

[0205] In some embodiments, Rais selected from the group consisting of methyl, methoxy, and phenyl.

[0206] In some embodiments, R1and R2are selected from the group consisting of cyano, -C(O)C6H5, -C(O)CH3, and -C(O)(OCH3).

[0207] In some embodiments, R1and R2are taken together with the carbon atom to which they are attached to form a 3-10 membered heterocyclyl or C3-io carbocyclyl, wherein the 3-10 membered heterocyclyl or C3-io carbocyclyl is substituted with one or more electron withdrawing groups and optionally further substituted.

[0208] In some embodiments, the 3-10 membered heterocyclyl or C3-io carbocyclyl is further substituted with one or more R3, wherein R3is selected from the group consisting of-Ci-6 alkyl, -C2-6 alkenyl, and -C3-io carbocyclyl.

[0209] In some embodiments, R3is -C1-6 alkyl.

[0210] In some embodiments, the one or more electron withdrawing groups is oxo.

[0211] In some embodiments, the 3-10 membered heterocyclyl or C3-io carbocyclyl is derived from Meldrum’s acid.

[0212] In some embodiments, Linker is a bond.

[0213] In some embodiments, Linker is -C2H4-.

[0214] In some embodiments, the compound is represented by Formula (IV-a):O'Linker .Peptide or ProteinH H (IV-a)or a salt thereof.

[0215] In some embodiments, Formula (IV-a) is represented by the structure:R4is the side chain of the N-terminal amino acid residue of the peptide or protein; R5is the side chain of the amino acid residue directly adjacent to the N-terminal amino acid residue of the peptide or protein;AAis an amino acid residue; and364928-1939-2137, v. 2n is an integer number representing the total number of amino acid residues in the peptide or protein minus the first two N-terminal amino acid residues. For example, if the total number of amino acids in a peptide or protein is 5, n is 3. In another example, if the total number of amino acids in a peptide or protein is 100, n is 98.

[0216] Another aspect of the disclosure provides a compound of Formula (V):R1R2° (V),or a salt thereof, whereinR1and R2are each, independently, an electron withdrawing group; orR1and R2are taken together with the carbon atom to which they are attached to form a 3-10 membered heterocyclyl or C3-10 carbocyclyl, wherein the 3-10 membered heterocyclyl or C3-10 carbocyclyl is substituted with one or more electron withdrawing groups and optionally further substituted;R4is the side chain of an amino acid residue; andLinker is selected from the group consisting of a bond, optionally substituted C1-6 alkylene, and optionally substituted C1-6 heteroalkylene.

[0217] In some embodiments, R1and R2are each, independently, an electron withdrawing group.

[0218] In some embodiments, R1and R2are identical.

[0219] In some embodiments, the electron withdrawing group is cyano or -CORa, wherein Rais selected from the group consisting of-Ci-6 alkyl, -C1-6 alkoxy, and -Ce-io aryl.

[0220] In some embodiments, Rais selected from the group consisting of methyl, methoxy, and phenyl.

[0221] In some embodiments, R1and R2are selected from the group consisting of cyano, -C(O)C6H5, -C(O)CH3, and -C(O)(OCH3).

[0222] In some embodiments, R1and R2are taken together with the carbon atom to which they are attached to form a 3-10 membered heterocyclyl or C3-10 carbocyclyl, wherein the 3-10 membered heterocyclyl or C3-10 carbocyclyl is substituted with one or more electron withdrawing groups and optionally further substituted.4928-1939-2137, v. 2

[0223] In some embodiments, the 3-10 membered heterocyclyl or C3-10 carbocyclyl is further substituted with one or more R3, wherein R3is selected from the group consisting of-Ci-6 alkyl, -C2-6 alkenyl, and -C3-10 carbocyclyl.

[0224] In some embodiments, R3is -C1-6 alkyl.

[0225] In some embodiments, the one or more electron withdrawing groups is oxo.

[0226] In some embodiments, the 3-10 membered heterocyclyl or C3-10 carbocyclyl is derived from Meldrum’s acid.

[0227] In some embodiments, R4is the side chain of an amino acid, wherein the amino acid had been truncated from the N-terminus of a peptide or protein.

[0228] In some embodiments, Linker is a bond.

[0229] In some embodiments, Linker is-C2H4-

[0230] In some embodiments, the compound is represented by Formula (V-a):or a salt thereof.

[0231] Another aspect of the disclosure provides a composition comprising one or more compounds disclosed herein.EXAMPLES

[0232] The present disclosure now being generally described, will be more readily understood by reference to the following examples, which are included merely for purposes of illustration of certain aspects and embodiments of the present disclosure, and is not intended to be limiting.General Methods Used

[0233] LC-MS: Liquid Chromatography / Mass spectra (LCMS) and flow injection analysismass spectrometry (FIA-MS) data were collected by an Agilent 6120 Single Quadrupole or Agilent 6125B Single Quadrupole LC / MS System (UT Austin Mass Spectrometry Facility).

[0234] High-resolution mass-spectrometry: High-resolution mass spectrometry (HRMS) data were collected using an Agilent 6530 Accurate-Mass Q-TOF LC / MS or Agilent 6546 Q-TOF LC / MS (UT Austin Mass Spectrometry Facility).4928-1939-2137, v. 2

[0235] Peptide synthesis: Peptide synthesis was carried out using an automated microwave-assisted solid-phase peptide synthesizer (Liberty Blue, CEM Corporation) with Fmoc chemistry. The DIC / Oxyma method was employed for coupling reactions in a 1:1:1 ratio with amino acids, conducted at 90 °C for 120 seconds. Deprotection steps were performed using 20% piperidine in DMF at 90 °C for 60 seconds. Peptides were cleaved from the resin with a mixture of trifluoroacetic acid (TFA), triisopropylsilane (TIS), and water (95:2.5:2.5) for 2.5 hours, after which the cleavage mixture was concentrated under nitrogen. The resulting peptides were precipitated using ice-cold diethyl ether. Reverse-phase column chromatography and HPLC purifications were conducted using a Shimadzu Prominence HPLC system equipped with a Zorbax SB-C18 preparative column (21.2 x 250 mm) containing 7.0 pm packing material. Analytical HPLC traces were obtained with a Zorbax SB-C18 analytical column (4.6 x 250 mm) featuring 5.0 pm packing material. A 5-95% gradient elution of acetonitrile / water with 0.1% (by volume) formic acid was applied.Example 1. Preparation of EVA reagent: 5-[Bis(methylthio)methylene]-2,2-dimethyl-l,3-dioxane-4, 6-dione

[0236] The method for synthesis of the EVA reagent (FIG. 1) is adapted from previously described method (Ben Cheikh, Abdelhamid, et al, The Journal of Organic Chemistry 56.3 (1991): 970-975). To a solution of Meldrum’s acid (10 g, 0.07 mol) in DMSO (30 mL) was successively added triethylamine (14.16 g, 0.14 mol) and carbon disulfide (5.30 g; 0.07 mol). After the mixture was stirred for 1 h at room temperature under Ar, methyl iodide (19.87 g; 0.14 mol) was added dropwise. After the mixture was stirred for 14 h, ice was added, and the yellow precipitate was collected and recrystallized from methanol, giving 9.5 g (55%); mp 119-121 °C (lit. mp 119-121 °C); IR (CDCL) 3000, 1728, 1680, 1410, 1310, 1280, 1040, 955 cm1; 'H NMR (CDCL) 8 1.75 (s, 6 H), 2.65 (s, 6 H);13C NMR (CDCL) 621.3 (q), 26.6 (q), 102.9 (s), 103.0 (s), 159.7 (s), 192.3 (s); MS m / z 248 (17), 191 (15), 190 (17), 172 (35), 146 (29), 118 (29), 100 (27), 99 (98), 85 (21). Anal. Calcd for C9H12O4S2: C, 43.53; H, 4.87. Found: C, 43.55; H, 4.80.Example 2. Preparation of EVA regent derivatives

[0237] The core structure of EVA reagent comprises two electron withdrawing groups (EWG) and a dithioalkyldiene moeity (shown in (1) in FIG. 2). The other structures shown in FIG. 2 are exemplary compounds which are capable of performing the methods herein.394928-1939-2137, v. 2Example 3. Synthetic methodology for synthesis of EVA reagent derivatives

[0238] Certain EVA reagent derivatives (e.g., compounds of Formula (I)) can be prepared via the below synthetic methodology shown in Scheme 1. Afunctional handle (alkyne) is installed for further derivatization.Scheme 1.

[0239] Other EVA reagent derivative of Formula (I) can be synthesized by derivatizing Compound 5 from FIG.2. Two reaction schemes are described below, Scheme 2 and Scheme 3, beginning with two distinct starting materials.Scheme 2.>Example 4. Mechanism of EVA assisted N-terminal degradation (termed “Ava chemistry”) using hydrazine404928-1939-2137, v. 2

[0240] Working out the reaction mechanisms (as shown in FIG.3), it is noted that the reagent reacts with the N-terminal amine to release methyl mercaptan. Substitution of the intermediate with an aliphatic diamine (hydrazine), generates a more effective nucleophile that attacks the carbonyl group of the next N-terminal residue, forming a six-membered ring, which rearranges and cleaves the terminal amino acid. The mechanism suggests that, this method does not require a strong base for two reasons: (i) the nucleophilicity of the amine adjacent to an alkene provides enhanced reactivity; and (ii) the conformation of the EVA-derivatized peptide with the diamine aligns all bonds from the amine through the N-terminus in the same plane.Example 5. Mechanism of EVA assisted N-terminal degradation (termed “Ava chemistry”) using ethylenediamine

[0241] To further support the hypothesis regarding the favorable stereochemistry and its effect on the reaction, the cleavage reaction was conducted using ethylenediamine (depicted in FIG.4) on EVA adduct of model peptide (NH2-VYFWVY-COOH) (SEQ ID NO: 1). The results (shown in FIG. 5) confirmed the stability of the eight-membered ring product, demonstrating that the rigid structure of this EVA derivative effectively overcomes transannular strain. This stability arises from the unique geometry of the eight-membered ring, where the planar arrangement of bonds minimizes steric hindrance and effectively overcomes the transannular strain.Example 6. Derivatization of Model Peptide with an EVA Reagent under mild-conditions

[0242] A model peptide (NH2-YGFWVY-COOH) (SEQ ID NO: 4) was dissolved in a 1:1 mixture of acetonitrile and 50 mM phosphate buffer (pH 7.6). EVA reagent (1 equivalent relative to peptide) was added, and the reaction mixture was maintained at 37 °C with continuous nitrogen sparging to facilitate the removal of methyl mercaptan. After 20 minutes, reaction completion was assessed by LC-MS and showed essentially complete conversion to the EVA-derivatized peptide (see LC-MS traces in FIG. 6). This example demonstrates that compounds of Formula (I) (e.g., the EVA reagent) rapidly and selectively modify the N-terminus of peptides under mild, near-neutral conditions.Example 7. Cleavage of the N-terminal amino acid - Tyrosine with EVA and hydrazine414928-1939-2137, v. 2

[0243] The EVA-derivatized model peptide X2GFWVY (SEQ ID NO: 5), wherein X2 is Tyr conjugated to the EVA reagent, from Example 6 (1 equivalent) was dissolved in a 3:2 mixture of acetonitrile and 50 mM phosphate buffer (pH 7.6). Similar results are also observed with higher % of acetonitrile (5:1 v:v). Hydrazine (12 equivalents) was added, and the mixture was incubated at 37 °C for 30 minutes, again under continuous nitrogen sparging.

[0244] LC-MS analysis confirmed formation of the truncated peptide together with the corresponding EVA-derived amino-acid adduct (FIG. 7).

[0245] Conversion efficiency was determined by LC-MS, by measuring the area under the curve of the N-terminal truncated peptide to all the peaks (reactant + side-products) observed in the 280 nm trace. For the model peptide, the conversion efficiency was 88%. High-resolution mass-spectrometry analysis (HRMS) of the product confirms the cleaved derivative of tyrosine and truncated peptide (FIG. 8)Example 8. EVA-Mediated N-Terminal Sequencing Across Multiple Amino Acids

[0246] The procedure of Examples 6 and 7 was applied to peptides containing different N-terminal amino acids, including tyrosine, glycine, arginine, aspartic acid, asparagine, glutamine, and tryptophan. For each peptide: Derivatization with EVA reagent proceeded to high conversion within ~20 minutes. Subsequent hydrazine treatment produced the truncated peptide and the corresponding EVA-bound amino-acid byproduct. Representative data from N-terminal arginine, asparagine, glutamine, histidine and phenylalanine are shown in FIG. 9-FIG. 13, respectively.

[0247] Conversion values determined by LC-MS are summarized in Table 1 and Table 2 with cleavage yields all ranging from 66% to 95%, depending on the identity of the N-terminal residue and the pH of the phosphate buffer (pH 7.5 or pH 6.5, the results of which are listed in Table 1 or Table 2 respectively). This example demonstrates that EVA derivatization followed by hydrazine addition enables efficient, mild N-terminal removal across multiple amino acids.Table 1. N-terminal cleavage under mild conditions (at about pH 7.5): % conversion calculated using LC-MS traces424928-1939-2137, v. 2Table 2. N-terminal cleavage under mild conditions (at about pH 6.5): % conversion calculated using LC-MS traces>>>>>>Example 9. EVA reagent as a peptide-capture reaction

[0248] Peptide capture using EVA was performed by dissolving peptide (1 equivalent) in a mixture of phosphate buffer (50-100 mM; pH 6-8) and DMF, followed by addition of EVA (1 equivalent). The mixture was stirred at 37 °C for 20 hours under nitrogen sparging. LC-MS analysis revealed formation of a cyclized capture product and minor side products (see, FIG.14).

[0249] This example shows that EVA reagent also functions as a reversible N-terminal capture reagent, forming stable capture adducts that can later participate in release chemistry.

[0250] A use case is envisioned of the functionalizing EVA reagent on a solid-support (magnetic beads, resin) - either through a single step or two step chemistry. Numerous methods are available for chemical derivatization of the solid-support and introducing a complementary reactive group on the EVA reagent. For example, the resin can be an azide resin and the EVA reagent can be modified to contain an alkyne handle. A standard Copper-assisted click reaction434928-1939-2137, v. 2can couple the peptides covalently on the resin, either through the N-terminal amine or the lysine side-chain.

[0251] Downstream biomolecular analysis of peptides or proteins on solid-support can be performed using methods such as antibody binding, sorting of beads based on fluorescent markers etc.

[0252] In the capture reaction, the N-terminal of the peptide undergoes conjugation with EVA (click reaction), leaving a methyl mercaptan. Subsequently, the second amine of the peptide's backbone initiates a nucleophilic attack, leading to the departure of a second methyl mercaptan and yielding a five-membered ring cyclized product. We anticipate that by introducing different types of amines, the cyclized product will undergo a declick reaction, resulting in the release of the peptide sample and the EVA derivative as the final products.Example 10. Subtractive N-terminal chemistry used in next-generation peptide and protein sequencing

[0253] Recent efforts to create protein analogs of DNA sequencing have produced several single-molecule peptide and protein sequencing technologies capable of identifying and quantifying proteins in complex mixtures with exceptional sensitivity. These approaches include optical fluorescence-based readouts, strategies that transduce peptide information into DNA, and nanopore-derived methods.

[0254] We developed one such platform, fluorosequencing, in which peptides are site-selectively fluorophore-labeled, immobilized for TIRF imaging, and sequenced by monitoring fluorescence loss during Edman degradation. However, we and others have found that the acidic conditions required for Edman chemistry, particularly TFA, degrade many useful fluorophores and also damage DNA. This limitation also affects multiple sequencing modalities, including methods that use N-terminal binders or antibody -based recognition units, using DNA molecules where depurination under Edman conditions necessitated either alternative reagents or chemically modified oligonucleotides. Consequently, overcoming the acid requirement of Edman degradation would broadly expand the compatibility and robustness of emerging single-molecule protein sequencing platforms.

[0255] Apart from the use cases described in fluorosequencing and other DNA based techniques, it can also be used for sequencing for Cyclesequencing (Schilling, Kevin, et al. "A Universal-Binder, De Novo Peptide Sequencing Platform with a Fluorescence Lifetime Fingerprint." (2025)) and Nanopore methods such as those described by Glyphic (see., e.g., Estandian, Daniel Masao, et al. "Single-molecule protein and peptide sequencing." U.S. Patent444928-1939-2137, v. 2No. 11,499,979. 15 Nov. 2022). Both these methods use Edman degradation as a method for cleavage of the N-terminal amino acid and can be substituted for a milder cleavage method.INCORPORATION BY REFERENCE

[0256] The entire disclosure of each of the patent documents and scientific articles referred to herein is incorporated by reference for all purposes.EQUIVALENTS

[0257] The invention may be embodied in other specific forms without departing from the spirit or essential characteristics thereof. The foregoing embodiments are therefore to be considered in all respects illustrative rather than limiting the invention described herein. Scope of the invention is thus indicated by the appended claims rather than by the foregoing description, and all changes that come within the meaning and range of equivalency of the claims are intended to be embraced therein.454928-1939-2137, v. 2

Claims

CLAIMSWhat is claimed is:

1. A method of functionalizing a peptide or protein having an N-terminus, the method comprising a coupling step comprising reacting the N-terminus of the peptide or protein with a compound of Formula (I):"or a salt thereof, wherein:R1and R2are each, independently, an electron withdrawing group; orR1and R2are taken together with the carbon atom to which they are attached to form a 3-10 membered heterocyclyl or C3-10 carbocyclyl, wherein the 3-10 membered heterocyclyl or C3-10 carbocyclyl is substituted with one or more electron withdrawing groups and optionally further substituted.to form a compound of Formula (II):

2. The method of claim 1, wherein R1and R2are each, independently, an electron withdrawing group.

3. The method of claim 2, wherein R1and R2are identical.

4. The method of claim 2 or 3, wherein the electron withdrawing group is cyano or - CORa, wherein Rais selected from the group consisting of-Ci-6 alkyl, -Ci-6 alkoxy, and optionally substituted -Ce-io aryl.

5. The method of claim 4, wherein Rais selected from the group consisting of methyl, methoxy, and phenyl.464928-1939-2137, v.

26. The method of claim 2 or 3, wherein R1and R2are selected from the group consisting of cyano, -C(O)C6H5, -C(O)CH3, and -C(O)(OCH3).

7. The method of claim 1, wherein R1and R2are taken together with the carbon atom to which they are attached to form a 3-10 membered heterocyclyl or C3-10 carbocyclyl, wherein the 3-10 membered heterocyclyl or C3-10 carbocyclyl is substituted with one or more electron withdrawing groups and optionally further substituted.

8. The method of claim 7, wherein the 3-10 membered heterocyclyl or C3-10 carbocyclyl is further substituted with one or more R3, wherein R3is selected from the group consisting of -C1-6 alkyl, -C2-6 alkenyl, and -C3-10 carbocyclyl.

9. The method of claim 8, wherein R3is -C1-6 alkyl.

10. The method of any one of claims 7-9, wherein the one or more electron withdrawing groups is oxo.

11. The method of any one of claims 7-10, wherein the 3-10 membered heterocyclyl or C3-10 carbocyclyl is derived from Meldrum’s acid.

12. The method of claim 1, wherein the compound of Formula (I) is selected from theand a derivative thereof.

13. The method of claim 12, wherein the derivative is substituted with one or more Rb, wherein Rbis selected from the group consisting of halo, cyano, -C1-6 alkyl, -C2-6 alkenyl, and -C3-10 carbocyclyl.474928-1939-2137, v.

214. The method of claim 1, wherein the compound is selected from the group consisting15. The method of claim 1, wherein the compound of Formula (I) is an EVA reagent:derivative thereof.

16. The method of claim 15, wherein the derivative is substituted with one or more Rb, wherein Rbis selected from the group consisting of halo, cyano, -Ci-6 alkyl, -C2-6 alkenyl, and -C3-10 carbocyclyl.

17. The method of claim 1, wherein the compound of Formula (I) is an EVA reagent:

18. The method of any one of claims 1-17, wherein the coupling step occurs in an aqueous buffer or in a mixture of an aqueous buffer and an organic solvent.

19. The method of claim 18, wherein the aqueous buffer is a sodium phosphate buffer.

20. The method of claim 19, wherein the pH of the sodium phosphate buffer is between about 6.0 about 8.0 and the concentration of the sodium phosphate buffer is about 0.1 M.

21. The method of claim 20, wherein the pH is about 7.6.484928-1939-2137, v.

222. The method of claim 18, wherein the organic solvent is acetonitrile or dimethylformamide.

23. The method of claim 18 or 22, wherein the ratio of the mixture of the aqueous buffer and the organic solvent is 1:1.

24. The method of any one of claims 1-23, wherein the compound of Formula (I) is conjugated to a solid support.

25. The method of any one of claims 1-23, wherein the compound of Formula (II) is conjugated to a solid support.

26. The method of any one of claims 1-25, wherein the ratio of the compound of Formula (I) and the peptide or protein is about 1:1.

27. The method of any one of claims 1-25, wherein the compound of Formula (I) is present in excess of the peptide or protein.

28. The method of claim 27, wherein the ratio of the compound of Formula (I) and the peptide or protein is selected from the group consisting of about 2:1, about 5:1, about 10:1, and about 100:1.

29. The method of any one of claims 1-28, wherein the reaction time of the coupling step is less than about an hour.

30. The method of any one of claims 1-29, wherein the reaction time of the coupling step is about 15 minutes.

31. The method of any one of claims 1-30, wherein the coupling step is performed at about 37 °C.

32. The method of any one of claims 1-31, wherein the method further comprises a step of washing to remove unreacted or excess amount of the compound of Formula (I) using an aqueous buffer, an organic solvent, or a mixture thereof.

33. The method of any one of claims 1-32, wherein the method further comprises a step of removing unreacted or excess amount of the compound of Formula (I) by filtration.

34. The method of claim 33, wherein the filtration comprises using C-18 tips.494928-1939-2137, v.

235. The method of any one of claims 1-34, wherein the method further comprises a step of cleaving the N-terminal amino acid residue of the peptide or protein to form a truncated peptide or protein.

36. The method of claim 35, wherein the step of cleaving comprises contacting a compound of Formula (II):or a salt thereof,with a diamine reagent of Formula (III):H2N LinkerNH2(III),or a salt thereof,to form a compound of Formula (IV):R1R2iH_i2MNJ Linker I JI^N N Peptide or ProteinH H (IV),or a salt thereof,wherein Linker is selected from the group consisting of a bond, optionally substituted Ci-6 alkylene, and optionally substituted Ci-6 heteroalkylene.

37. The method of claim 36, wherein the method further comprises intramolecularly reacting a compound of Formula (IV):or a salt thereof to form the truncated peptide or protein and a compound of Formula (V):504928-1939-2137, v. 2or a salt thereof; andR4is the side chain of the N-terminal amino acid residue of the peptide or protein.

38. The method of any one of claims 35-37, wherein the diamine reagent is selected from the group consisting of hydrazine and ethylenediamine.

39. The method of any one of claims 35-38, wherein the step of cleaving occurs in an aqueous buffer or in a mixture of an aqueous buffer and an organic solvent.

40. The method of claim 39, wherein the pH of the sodium phosphate buffer is between about 6.0 and 8.0 and the concentration of the sodium phosphate buffer is about 0.1 M.

41. The method of claim 40, wherein the pH is about 7.5.

42. The method of claim 39, wherein the organic solvent is acetonitrile.

43. The method of claim 39 or 42, wherein the ratio of the mixture of the aqueous buffer and the organic solvent is 1:3.

44. The method of any one of claims 36-43, wherein the ratio of the diamine reagent and the compound of Formula (II) is about 1:1.

45. The method of any one of claims 36-44, wherein the diamine reagent is present in excess of the compound of Formula (II).

46. The method of claim 45, wherein the ratio of diamine reagent and the compound of Formula (II) is selected from the group consisting of about 5:1, about 10:1, and about 12:1.

47. The method of claim 46, wherein the ratio is about 12:1.514928-1939-2137, v.

248. The method of any one of claims 35-47, wherein the reaction time of the step of cleaving is less than about an hour.

49. The method of any one of claims 35-48, wherein the reaction time of the step of cleaving is about 30 minutes.

50. The method of any one of claims 35-49, wherein the step of cleaving is performed at about 37 °C.

51. The method of any one of claims 1-50, wherein the peptide or protein comprises natural amino acids or derivatives thereof.

52. The method of any one of claims 1-50, wherein the peptide or protein comprises unnatural amino acids, alpha-amino acids or combinations or derivatives thereof.

53. The method of any one of claims 1-52, wherein the method further comprises a step of sequencing the peptide or protein.

54. The method of claim 53, wherein the sequencing is N-terminal sequencing.

55. The method of any one of claims 37-54, wherein the method further comprises a step of detecting the presence of the truncated peptide or protein or the compound of Formula (V).

56. The method of claim 55, wherein the step of detecting comprises detecting the compound of Formula (V).

57. The method of claim 56, wherein the step of detecting comprises performing mass spectrometry.

58. The method of claim 57, wherein the mass spectrometry comprises LC-MS.

59. The method of claim 55, wherein the step of detecting comprises detecting the truncated peptide or protein.

60. The method of claim 59, wherein the step of detecting comprises fluorosequencing.524928-1939-2137, v.

261. The method of claim 59, wherein the truncated peptide or protein is attached to an oligonucleotide barcode, and the step of detecting comprises detecting or sequencing the oligonucleotide barcode.

62. A method of sequencing a plurality of peptide or proteins comprising sequencing a plurality of peptide or proteins functionalized by the method of any one of claims 1- 61.

63. The method of claim 62, wherein the peptides or proteins are unique from each other.

64. The method of claim 1, wherein the solid support is a magnetic bead or resin.

65. A compound of Formula (IV):R1R2Peptide or ProteinH H (IV),or a salt thereof, whereinR1and R2are each, independently, an electron withdrawing group; orR1and R2are taken together with the carbon atom to which they are attached to form a 3-10 membered heterocyclyl or C3-10 carbocyclyl, wherein the 3-10 membered heterocyclyl or C3-10 carbocyclyl is substituted with one or more electron withdrawing groups and optionally further substituted; andLinker is selected from the group consisting of a bond, optionally substituted C1-6 alkylene, and optionally substituted C1-6 heteroalkylene.

66. The compound of claim 65, wherein R1and R2are each, independently, an electron withdrawing group.

67. The compound of claim 66, wherein R1and R2are identical.

68. The compound of claim 66 or 67, wherein the electron withdrawing group is cyano or -CORa, wherein Rais selected from the group consisting of-Ci-6 alkyl, -C1-6 alkoxy, and -Ce- 10 aryl.534928-1939-2137, v.

269. The compound of claim 68, wherein Rais selected from the group consisting of methyl, methoxy, and phenyl.

70. The compound of any one of claims 66-69, wherein R1and R2are selected from the group consisting of cyano, -C(O)CeH5, -C(O)CH3, and -C(O)(OCH3).

71. The compound of claim 65, wherein R1and R2are taken together with the carbon atom to which they are attached to form a 3-10 membered heterocyclyl or C3-10 carbocyclyl, wherein the 3-10 membered heterocyclyl or C3-10 carbocyclyl is substituted with one or more electron withdrawing groups and optionally further substituted.

72. The compound of claim 71, wherein the 3-10 membered heterocyclyl or C3-10 carbocyclyl is further substituted with one or more R3, wherein R3is selected from the group consisting of-Ci-6 alkyl, -C2-6 alkenyl, and -C3-10 carbocyclyl.

73. The compound of claim 72, wherein R3is -C1-6 alkyl.

74. The compound of any one of claims 71-73, wherein the one or more electron withdrawing groups is oxo.

75. The compound of any one of claims 69-73, wherein the 3-10 membered heterocyclyl or C3-10 carbocyclyl is derived from Meldrum’s acid.

76. The compound of any one of claims 65-72, wherein Linker is a bond.

77. The compound of any one of claims 65-72, wherein Linker is -C2H4-.

78. The compound of claim 65, wherein the compound is represented by Formula (IV-a):or a salt thereof.

79. A compound of Formula (V):544928-1939-2137, v. 2or a salt thereof, whereinR1and R2are each, independently, an electron withdrawing group; orR1and R2are taken together with the carbon atom to which they are attached to form a 3-10 membered heterocyclyl or C3-10 carbocyclyl, wherein the 3-10 membered heterocyclyl or C3-10 carbocyclyl is substituted with one or more electron withdrawing groups and optionally further substituted;R4is the side chain of an amino acid residue; andLinker is selected from the group consisting of a bond, optionally substituted C1-6 alkylene, and optionally substituted C1-6 heteroalkylene.

80. The compound of claim 79, wherein R1and R2are each, independently, an electron withdrawing group.

81. The compound of claim 80, wherein R1and R2are identical.

82. The compound of claim 80 or 81, wherein the electron withdrawing group is cyano or -CORa, wherein Rais selected from the group consisting of-Ci-6 alkyl, -C1-6 alkoxy, and -Ce- 10 aryl.

83. The compound of claim 82, wherein Rais selected from the group consisting of methyl, methoxy, and phenyl.

84. The compound of any one of claims 80-83, wherein R1and R2are selected from the group consisting of cyano, -C(O)CeH5, -C(O)CH3, and -C(O)(OCH3).

85. The compound of claim 79, wherein R1and R2are taken together with the carbon atom to which they are attached to form a 3-10 membered heterocyclyl or C3-10 carbocyclyl, wherein the 3-10 membered heterocyclyl or C3-10 carbocyclyl is554928-1939-2137, v. 2substituted with one or more electron withdrawing groups and optionally further substituted.

86. The compound of claim 85, wherein the 3-10 membered heterocyclyl or C3-10 carbocyclyl is further substituted with one or more R3, wherein R3is selected from the group consisting of-Ci-6 alkyl, -C2-6 alkenyl, and -C3-10 carbocyclyl.

87. The compound of claim 86, wherein R3is -C1-6 alkyl.

88. The compound of any one of claims 85-87, wherein the one or more electron withdrawing groups is oxo.

89. The compound of any one of claims 85-88, wherein the 3-10 membered heterocyclyl or C3-10 carbocyclyl is derived from Meldrum’s acid.

90. The compound of any one of claims 79-89, wherein R4is the side chain of an amino acid, wherein the amino acid had been truncated from the N-terminus of a peptide or protein.

91. The compound of any one of claims 79-90, wherein Linker is a bond.

92. The compound of any one of claims 79-90, wherein Linker is-C2H4-93. The compound of claim 79, wherein the compound is represented by Formula (V-a):(V-a)or a salt thereof.

94. A kit comprising a compound of Formula (I), or a salt thereof, for use in a method of any one of claims 1-64.

95. A method of N-terminal peptide sequencing of a peptide or protein having an N- terminus, the method comprising the steps of564928-1939-2137, v. 2reacting the N-terminus of the peptide or protein with an EVA reagent:salt thereof,to form a compound of Formula (Il-a):cleaving the N-terminal amino acid residue of the peptide or protein to form a truncated peptide or protein by contacting a compound of Formula (Il-a) with a diamine reagent of Formula (III):or a salt thereof,to form a compound of Formula (IV-a):or a salt thereof, wherein the compound of Formula (IV-a) or salt thereof intramolecularly reacts to form thetruncated peptide or protein and a compound of Formula (V-a):574928-1939-2137, v. 2or a salt thereof, whereinLinker is a bond or -C2H4-; andR4is the side chain of the N-terminal amino acid residue of the peptide or protein.

96. The method of claim 95, wherein the diamine reagent is hydrazine.

97. The method of claim 95, wherein the diamine reagent is ethylenediamine.

98. The method of any one of claims 95-97, wherein the method further comprises a step of washing to remove unreacted or excess amount of the EVA reagent using an aqueous buffer, an organic solvent, or a mixture thereof.

99. The method of any one of claims 95-98, wherein the method further comprises a step of removing unreacted or excess amount of the EVA reagent by filtration.

100. The method of any one of claims 95-99, wherein the method further comprises a step of detecting the presence of the truncated peptide or protein or the compound of Formula (V-a).

101. The method of claim 100, wherein the step of detecting comprises detecting the compound of Formula (V-a).

102. The method of claim 101, wherein the step of detecting comprises performing mass spectrometry.

103. The method of claim 102, wherein the mass spectrometry comprises LC-MS.

104. The method of claim 100, wherein the step of detecting comprises detecting the truncated peptide or protein.584928-1939-2137, v. 2105. The method of claim 104, wherein the step of detecting comprises fluorosequencing.

106. The method claim 104, wherein the truncated peptide or protein is attached to an oligonucleotide barcode, and the step of detecting comprises detecting the oligonucleotide barcode.

107. A kit comprising the EVA reagent, or a salt thereof, for use in a method of any one of claims 95-106.594928-1939-2137, v. 2