Sequencing reagents using guanidinylating agents

Sequencing reagents with guanidinylating agents address the limitations of current protein sequencing methods by enabling single-molecule detection and spatial localization of proteins, enhancing the identification of low copy-number proteins.

WO2026030423A1PCT designated stage Publication Date: 2026-02-05GLYPHIC BIOTECHNOLOGIES INC
View PDF 14 Cites 0 Cited by

Patent Information

Application Number
PCT/US2025/039826
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-07-31
Filing Date
2025-07-30
Publication Date
2026-02-05

AI Technical Summary

Technical Problem

Current protein sequencing methods, such as mass spectrometry and Edman degradation, lack single-molecule sensitivity and spatial information, while immunohistochemistry provides limited scalability and sequence information, making it difficult to quantify and identify low copy-number proteins effectively.

Method used

The use of sequencing reagents comprising guanidinylating agents with reactive groups, linkers, and substrates allows for the sequencing of polymeric analytes like polypeptides, enabling single-molecule detection and spatial localization of proteins through methods involving binding, cleavage, and detection using click chemistry and nanopores.

Benefits of technology

The proposed sequencing reagents and methods enhance the ability to sequence proteins with single-molecule sensitivity and provide spatial information, overcoming the limitations of existing techniques by improving detection and localization of low copy-number proteins.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2025039826_05022026_PF_FP_ABST
    Figure US2025039826_05022026_PF_FP_ABST
Patent Text Reader

Abstract

The present disclosure provides reagents and methods useful for single-molecule sequencing of proteins through use of a sequencing reagents of Formula (Ia), (Ib), (I-Aa), (I-Ab), (IIa), (IIb), (II-Aa), (II-Ab), (IIIa), (IIIb), (III-Aa), (III-Ab), (IVa), (IVb), (IV-Aa), (IV-Ab), (IV-Ba), (IV-Bb), (V), (V-A), (V-B), (VI), (VI-A), (VII), or (VII-A). The reagents and methods described herein provide for high-throughput and high efficiency protein and peptide sequencing in mild conditions allowing for high resolution investigation of complex biological systems.
Need to check novelty before this filing date? Find Prior Art

Description

SEQUENCING REAGENTS USING GUANIDINYLATING AGENTSCROSS REFERENCE

[0001] This application claims the benefit of U.S. Provisional Patent Application No.63 / 677,787, filed July 31, 2024, which is incorporated by reference herein in its entirety.BACKGROUND

[0002] Proteins serve a critical role at a cellular level, carrying out a variety of integral functions. Having the technology required to quantify and identify proteins is crucial to understanding their contributions to biological function. Advancements in proteomics have lagged behind while DNA sequencing has rapidly advanced the study of genomics primarily due to technologies that allow for high-throughput sequencing. Current methodologies available for studying proteins include mass spectrometry, Edman sequencing, and immunohistochemistry.

[0003] Mass spectrometry (MS) has enabled protein identification and quantification based on the mass / charge ratio of peptide fragments, which can be bioinformatically mapped back to a genomic database. However, MS has yet to quantify a complete set of proteins from a biological system, despite significant advancements. MS exhibits attomole detection for whole proteins and subattomole sensitives after fractionation. Yet, functionally-important, low copy-number proteins that make up about 10% of mammalian protein expression remain undetected.

[0004] Edman degradation allows for sequential and selective removal of single N-terminal amino acids which are subsequently identified via HPLC (High-Performance Liquid Chromatography). Edman protein sequencing removes the first N-terminal amino acid for identification using phenyl isothiocyanate (PITC) to conjugate to the N-terminal amino acid, then upon acid and heat treatment, the PITC-labeled N-terminal amino acid is removed. Although Edman sequencing can have 98% efficiency, a major drawback is that it is inherently low throughput, requiring a single highly purified protein, and the inapplicability to systems- wide biology. Moreover, Edman degradation presents a number of other drawbacks including harsh reaction conditions such as heat and acidic conditions which are not amenable to using or analyzing nucleic acids, resultant chiral residues which can hinder detection via stereoisomerspecific detection agents, and modification of lysine residues, which can prevent further functionalization of the lysine side chains. Both Edman degradation and mass spectrometry can sequence proteins but lack single-molecule sensitivity and do not provide spatial information of proteins in the context of cells.10005] In regard to spatial information, immunohistochemistry is a protein identification method that allows visualization of cellular localization of proteins but does not provide sequence information. Immunohistochemistry involves the identification of proteins via recognition with fluorophore-conjugated antibodies. This approach excludes protein sequence information but can identify proteins and their respective localizations. A major limitation is the scalability, since even the perfect construction of specific antibodies for every protein in the proteome would require around 25,000 antibodies and, -6250 rounds of four-color imaging.SUMMARY

[0006] Considering the present need for improved methods of single molecule protein sequencing, presented herein are formulas, compounds, and methods for addressing the abovementioned need. In some embodiments, sequencing reagents are provided herein comprising a reactive group, a substrate-tethering moiety, and a linker. In some instances, these sequencing reagents allow for sequencing of polymeric analytes, such as polypeptides. Provided herein are sequencing reagents comprising reactive groups comprising guanidinylating agents. Also provided herein are methods of sequencing polymeric analytes using the sequencing reagents provided herein. A brief summary of various exemplary embodiments is presented. Some simplifications and omissions may be made in the following summary, which is intended to highlight and introduce aspects of certain embodiments disclosed herein, but not to limit the scope of the disclosure. Detailed descriptions of various embodiments adequate to allow those of ordinary skill in the art to make and use the concepts disclosed herein will follow in later sections.

[0007] One aspect of the present disclosure provides a sequencing reagent of Formula (la) or Formula (lb), or a stereoisomer, tautomer, or salt thereof.Formula (la) Formula (lb).

[0008] In some embodiments, LG is a leaving group which is optionally substituted with one or more electron-withdrawing group (e.g., electron-withdrawing protecting group). In some embodiments, L1is a linker, a bond, or absent. In some embodiments, R1, R2, and R3are each independently hydrogen, a protecting group, an amino-protecting group, an electron-withdrawing group, an amino-electron-withdrawing group, or an electron- withdraw! ng protecting group. In some embodiments, B is a reactive moiety or polymer.

[0009] In some embodiments, the compound of Formula (la) or Formula (lb) is represented by the compound of Formula (I-Aa) or Formula (I-Ab), wherein Ring A is C4-C12 heteroaryl or C4-C12 heterocycle optionally substituted with one or more electron-withdrawing groups (e.g., electron-withdrawing protecting groups).Formula (I-Aa) Formula (I-Ab).

[0010] One aspect of the present disclosure provides a sequencing reagent of Formula (Ila) or Formula (lib), or a stereoisomer, tautomer, or salt thereof.Formula (Ila) Formula (lib).

[0011] In some embodiments, LG is a leaving group which is optionally substituted with one or more electron-withdrawing groups (e.g., electron-withdrawing protecting groups). In some embodiments, R1is hydrogen, a protecting group, an amino-protecting group, an electronwithdrawing group, an amino-electron-withdrawing group, or an electron-withdrawing protecting group. In some embodiments, R2is hydrogen, a protecting group, an amino- protecting group, an electron-withdrawing group, an amino-electron-withdrawing group, or an electron-withdrawing protecting group. In some embodiments, L1is a linker, a bond, or absent. In some embodiments, B is a reactive moiety or a polymer. In some embodiments, n is an integer from 0 to 3.

[0012] In some embodiments, the compound of Formula (Ila) or Formula (lib) is represented by the compound of Formula (II-Aa) or Formula (Il-Ab), wherein Ring A is C4-C12 heteroaryl or C4-C12 heterocycle optionally substituted with one or more electron-withdrawing groups (e.g., electron-withdrawing protecting groups).F ormul a (II- Aa) F ormul a (II- Ab ) .

[0013] One aspect of the present disclosure provides a sequencing reagent of Formula (Illa), Formula (Illb), Formula (IVa), or Formula (IVb), or a stereoisomer, tautomer, or salt thereof.Formula (Illa), Formula (Illb), Formula (IVa), Formula (IVb).

[0014] In some embodiments, LG is a leaving group, optionally substituted with one or more electron-withdrawing groups (e.g., electron-withdrawing protecting groups). In some embodiments, L1is a linker, a bond, or absent. In some embodiments, L2is a hydrogen or a linker. In some embodiments, L3and L4are independently linkers, taken together to form heterocycle. In some embodiments, B is a reactive moiety or a polymer. In some embodiments, R1is hydrogen, a protecting group, an amino-protecting group, an electron-withdrawing group, an amino-electron-withdrawing group, or an electron-withdrawing protecting group. In some embodiments, R2is hydrogen, a protecting group, an amino-protecting group, an electronwithdrawing group, an amino-electron-withdrawing group, or an electron-withdrawing protecting group.

[0015] In some embodiments, the compound of Formula (Illa), Formula (Illb), Formula (IVa), or Formula (IVb) is represented by the compound of Formula (III-Aa), Formula (III-Ab), Formula (IV-Aa), or Formula (IV-Ab), wherein Ring A is C4-C12 heteroaryl or C4-C12 heterocycle optionally substituted with one or more electron-withdrawing groups (e.g., electronwithdrawing protecting groups). In some embodiments, L3and L4are taken together to form C5- Cs heterocycle.Formula (III-Aa), Formula (III-Ab), Formula (IV-Aa), Formula (IV-Ab).

[0016] In some embodiments, the compound of Formula (IVa), Formula (IVb), Formula (IV-Aa), or Formula (IV-Ab) is represented by the compound of Formula (IV-Ba) or Formula (IV-Bb), wherein n is an integer from 0 to 3.Formula (IV-Ba) Formula (IV-Bb).

[0017] One aspect of the present disclosure provides a sequencing reagent of Formula (V) or a stereoisomer, tautomer, or salt thereof.Formula (V).

[0018] In some embodiments, the compound of Formula (V) is represented by the compound of Formula (V-A) or Formula (V-B), wherein each Ring A is independently C4-C12 heteroaryl or C4-C12 heterocycle, optionally substituted with one or more electron-withdrawing groups.Formula (V-A) Formula (V-B).

[0019] One aspect of the present disclosure provides a sequencing reagent of Formula (VI) or Formula (VII), or a stereoisomer, tautomer, or salt thereof. In some embodiments, each LG is independently a leaving group, optionally substituted with one or more electron-withdrawing groups. In some embodiments, L1is a linker, a bond, or absent. In some embodiments, L2is a hydrogen or a linker. In some embodiments, L3and L4are independently linkers, taken together to form heterocycle. In some embodiments, B is a reactive moiety or a polymer.Formula (VI) Formula (VII).10020] In some embodiments, the compound of Formula (VI) or Formula (VII) is represented by the compound of Formula (VI-A) or Formula (VILA), wherein each Ring A is independently C4-C12 heteroaryl or C4-C12 heterocycle, optionally substituted with one or more electron-withdrawing groups.Formula (VI- A) Formula (VII- A).

[0021] In some embodiments, the leaving group is electron-withdrawing. In some embodiments, the leaving group is C4-C12 heteroaryl or C4-C12 heterocycle each optionally substituted with one or more electron-withdrawing groups (e.g., electron-withdrawing protecting groups).

[0022] In some embodiments, the leaving group or Ring A is C4-C6 heteroaryl, optionally substituted with one or more electron-withdrawing groups (e.g., electron-withdrawing protecting groups). In some embodiments, the leaving group or Ring A is an azole or azine, optionally substituted with one or more electron-withdrawing groups (e.g., electron-withdrawing protecting groups). In some embodiments, the leaving group or Ring A is diazole, triazole, or tetrazole, optionally substituted with one or more electron-withdrawing groups (e.g., electron-withdrawing protecting groups).

[0023] In some embodiments, the leaving group or Ring A is diazole, triazole, or tetrazole, each optionally fused with aryl or heteroaryl.

[0024] In some embodiments, the leaving group or Ring A is triazole fused with aryl or heteroaryl.

[0025] In some embodiments, the leaving group or Ring A is:0, wherein each of X1-X5 is independently selected from N, CH, or C-EWG; each of Y1-Y4 is independently selected from N, CH, or C-EWG; and EWG is an electron-withdrawing group (e.g., electron-withdrawing protecting group); wherein at least one of X1-X5 is N.

[0026] Another aspect of the present disclosure provides a non-transitory computer readable medium comprising machine executable code that, upon execution by one or more computer processors, implements any of the methods above or elsewhere herein.

[0027] In some embodiments, the leaving group or Ring, optionally substituted with one or more electron-withdrawing groups (e.g., electronwithdrawing protecting groups). In some embodiments, the leaving group or Ring A is C4-C6 heteroaryl fused with C4-C6 aryl, optionally substituted with one or more electron-withdrawing groups (e.g., electron-withdrawing protecting groups). In some embodiments, the leaving group or Ring A is diazole, triazole, or tetrazole, fused with aryl, optionally substituted with one or more electron-withdrawing groups (e.g., electron-withdrawing protecting groups). In somemore electron-withdrawing groups (e.g., electron-withdrawing protecting groups).

[0028] In some embodiments, the electron-withdrawing group is each independently selected from halogen, haloalkyl, nitro, sulfonate, amino, alkylamino, cyano, or a carbonyl. In some embodiments, the carbonyl comprises an acetyl group or a derivative thereof, a carboxylic acid, or an aldehyde. In some embodiments, at least one electron- withdraw! ng group comprises Ci-Ce haloalkyl. In some embodiments, at least one electron-withdrawing group comprises CF3. In some embodiments, at least one electron-withdrawing group comprises nitro. In some embodiments, at least one electron-withdrawing group comprises halogen. In some embodiments, the halogen is fluoro or chloro.

[0029] In some embodiments, the leaving group or Ring

[0030] In some embodiments, the linker is configured to enable a nucleic acid attraction. In some embodiments, the nucleic acid attraction is ionic. In some embodiments, the nucleic acid attraction is non-ionic. In some embodiments, the linker is configured to provide steric space for click chemistry via B. In some embodiments, the linker is electron-withdrawing. In some embodiments, a first end of the linker is electron-withdrawing. In some embodiments, both a first end and a second end of the linker are electron-withdrawing. In some embodiments, each linker is independently selected from optionally substituted alkyl, optionally substituted heteroalkyl, an oligonucleotide, DNA, RNA, or a peptide. In some embodiments, the linker is or comprises, wherein at least one of x is an integer greater than 0. In some embodiments, at least one linker is -CH2-. In some embodiments, L1and L2are taken together to form Cs-Cs heterocycle.

[0031] In some embodiments, at least one of R1, R2, and R3is an amine-protecting group. In some embodiments, R1, R2, and R3are each independently hydrogen, a protecting group, or an amino-protecting group, wherein the protecting group or amino-protecting group comprises carbobenzyl oxy (CBz) (e.g., benzyl carbamate), acetamide (Ac), trifluoroacetyl (TFAc), phthalimide, benzyl, benzylamine (Bn), benzoyl (Bz), triphenylmethyl (Tr) (e.g., triphenylmethylamine), benzylideneneamine, tosyl (Ts) (e.g., / ?-toluenesulfonamide), Methylsulfonylethoxycarbonyl (Msc), tert-butyloxycarbonyl (Boc) (e.g., / -butyl carbamate), or fluorenylmethyloxycarbonyl (Fmoc) (e.g., 9-fluorenylmethyl carbamate), 2,7-disulfo-9- fluorenylmethoxy carbonyl (Smoc), or acetyl.

[0032] In some embodiments, the reactive moiety B comprises a click chemistry moiety. In some embodiments, the click chemistry moiety B is an azide, a sulfonyl azide, an alkyne, a thio acid, a tetrazine, bicyclo[6.1.0]nonyne (BCN), a diarylcyclooctyne (e.g., DBCO),transcyclooctene (TCO), cyclopropane, norbornene, or a spiroalkene. In some embodiments, the reactive moiety B comprises a phosphite, phosphine, pentafluorophenyl ester (PFP), tetrafluorophenyl ester (TFP), 4-sulfo-2,3,5,6-tetrafluorophenyl ester (STP), or thio-phthalimide.

[0033] In some embodiments, the compound is represented by

[0034] In some embodiments, the compound is represented by

[0035] In some embodiments, the compound is represented by the structure

[0036] In some embodiments, the compound is represented by the structure

[0037] In some embodiments, the compound is represented by the structure

[0038] In some embodiments, the compound is represented by the structure

[0039] In some embodiments, the reactive moiety B comprises a polymer. In some embodiments, the polymer comprises a polynucleotide, a polypeptide, a polymeric material, or a combination thereof. In some embodiments, the polynucleotide comprises a deoxyribonucleic acid (DNA), a ribonucleic acid (RNA), a locked nucleic acid (LNA), a parallel-stranded DNA(p-DNA), a non-constrained nucleic acid (NNA), or a modified version thereof. In some embodiments, the polynucleotide is a deoxyribonucleic acid (DNA).

[0040] In some embodiments, the compound is a sequencing reagent.

[0041] One aspect of the present disclosure provides a method of using the sequencing reagent. The method comprising: (a) providing a polymeric analyte; (b) contacting the polymeric analyte with the sequencing reagent, wherein the sequencing reagent binds to a monomer of the polymeric analyte, thereby generating a sequencing reagent-monomer complex; and (c) cleaving the sequencing reagent-monomer complex from the polymeric analyte, thereby providing a cleaved sequencing reagent-monomer complex.

[0042] In some embodiments, the method further comprises coupling the sequencing reagent or the sequencing reagent-monomer complex to a capture moiety.

[0043] In some embodiments, the capture moiety comprises a DNA molecule. In some embodiments, the capture moiety comprises a modified DNA molecule. In some embodiments, the sequencing reagent comprises a modified DNA molecule. In some embodiments, the modified DNA molecule comprises a click chemistry moiety-modified DNA molecule.

[0044] In some embodiments, the sequencing reagent is coupled to the modified DNA molecule via click chemistry. In some embodiments, the polymeric analyte comprises a polypeptide. In some embodiments, the monomer comprises a terminal amino acid residue. In some embodiments, the method further comprises, detecting the cleaved sequencing reagent- monomer complex or derivative thereof, wherein detecting comprises contacting the sequencing reagent-monomer complex or derivative thereof with a binding agent. In some embodiments, the binding agent comprises an antibody, nanobody, single chain variable fragment (scFv), or aptamer. In some embodiments, the binding agent comprises a polymerizable molecule. In some embodiments, polymerizable molecule comprises a nucleic acid molecule. In some embodiments, the method further comprises deprotecting the protecting groups (e.g., R1, R2, and / or R3) of the sequencing reagent. In some embodiments, the deprotecting is performed before cleaving the sequence reagent-monomer complex from the polymeric analyte. In some embodiments, the deprotecting is performed before coupling the sequencing reagent to the capture moiety. In some embodiments, the deprotecting is performed in the presence of a base.

[0045] In some embodiments, the method further comprises (d) contacting the polymeric analyte with an additional sequencing reagent, wherein the additional sequencing reagent binds to an additional monomer of the polymeric analyte, thereby generating an additional sequencing reagent-monomer complex; (e) coupling the additional sequencing reagent to the cleaved sequencing reagent-monomer complex; and (f) cleaving the additional sequencing reagent-monomer complex from the polymeric analyte, thereby providing a stacked sequencing reagent- monomer complex.

[0046] In some embodiments, wherein in (a), the polymeric analyte is coupled to a substrate. In some embodiments, cleaving the sequencing reagent-monomer complex from the polymeric analyte is performed chemically or enzymatically. In some embodiments, cleaving the sequencing reagent-monomer complex from the polymeric analyte is performed chemically in the presence of a base. In some embodiments, the method further comprises detecting the cleaved sequencing reagent-monomer complex or derivative thereof. In some embodiments, the detecting is performed using a nanopore. In some embodiments, the detecting is performed by measuring a signal from the nanopore as the cleaved sequencing reagent-monomer complex or derivative thereof translocates through the nanopore, thereby generating a measured signal, and using the measured signal to identify the monomer.

[0047] Another aspect of the present disclosure provides a system comprising one or more computer processors and computer memory coupled thereto. The computer memory comprises machine executable code that, upon execution by the one or more computer processors, implements any of the methods above or elsewhere herein.

[0048] Additional aspects and advantages of the present disclosure will become readily apparent to those skilled in this art from the following detailed description, wherein only illustrative embodiments of the present disclosure are shown and described. As will be realized, the present disclosure is capable of other and different embodiments, and its several details are capable of modifications in various obvious respects, all without departing from the disclosure. Accordingly, the drawings and description are to be regarded as illustrative in nature, and not as restrictive.INCORPORATION BY REFERENCE

[0049] All publications, patents, and patent applications mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent, or patent application was specifically and individually indicated to be incorporated by reference. To the extent publications and patents or patent applications incorporated by reference contradict the disclosure contained in the specification, the specification is intended to supersede and / or take precedence over any such contradictory material.BRIEF DESCRIPTION OF THE DRAWINGS

[0050] The novel features of the disclosure are set forth with particularity in the appended claims. A better understanding of the features and advantages of the present disclosure will be obtained by reference to the following detailed description that sets forth illustrative embodiments, in which the principles of the disclosure are utilized, and the accompanying drawings (also “Figure” and “FIG.” herein), of which:

[0051] FIG. 1A schematically shows an example workflow for processing polymeric analytes molecules (e.g., peptides) described herein. FIG. IB schematically shows another example workflow for processing polymeric analytes in solution or on a substrate. FIG. 1C schematically shows another example workflow for processing polymeric analytes in solution or on a substrate. FIG. ID schematically shows yet another example workflow for processing polymeric analytes in solution or on a substrate.

[0052] FIG. 2A schematically shows an example workflow for processing polymeric analytes molecules (e.g., peptides) described herein. FIG. 2B schematically shows another example workflow for processing polymeric analytes in solution. FIG. 2C schematically shows another example workflow for processing polymeric analytes and detection. FIG. 2D schematically shows another example workflow for processing polymeric analytes and detection using a nanopore sequencing system.

[0053] FIG. 3 schematically shows an exemplary linker for coupling polymerizable molecules to polymeric analytes.

[0054] FIG. 4 schematically shows a computer system that is programmed or otherwise configured to implement methods provided herein.

[0055] FIG. 5 shows a scheme for attachment of a compound (e.g., sequencing reagent) provided herein to a peptide and modified DNA.

[0056] FIG. 6 shows a general reaction mechanism for using a guanidinylating reagent.

[0057] FIG. 7A shows a scheme for preparation of an intermediate compound en route to a compound provided herein. FIG. 7B shows a scheme for preparation of another intermediate compound en route to a compound provided herein.

[0058] FIG. 8A shows a scheme for preparation of a compound provided herein. FIG. 8B shows a scheme for preparation of a compound provided herein. FIG. 8C shows a scheme for preparation of a compound provided herein. FIG. 8D shows a scheme for preparation of a compound provided herein. FIG. 8E shows a scheme for preparation of a compound provided herein. FIG. 8F shows a scheme for preparation of a compound provided herein. FIG. 8G shows a scheme for preparation of a compound provided herein.

[0059] FIG. 9A shows schemes for preparation of two compounds provided herein. FIG. 9B shows alternative schemes for preparation of two compounds provided herein.

[0060] FIG. 10A shows a scheme for preparation of a compound provided herein. FIG. 10B shows a scheme for preparation of a compound provided herein.

[0061] FIG. 11A shows a scheme for preparation of a compound provided herein. FIG. 11B shows a scheme for preparation of a compound provided herein. FIG. 11C shows a scheme for preparation of a compound provided herein. FIG. 11D shows a scheme for preparation of a compound provided herein.

[0062] FIG. 12A shows an example of using a compound (e.g., a guanidinylating reagent) for the conjugation and cleavage of a peptide. FIG. 12B shows the extracted ion chromatogram (EIC) analysis of the starting material tripeptide AGF. FIG. 12C shows the EIC analysis of the Boc protected conjugate (Compound 1-K)-AGF after conjugating the guanidinylating reagent Compound 1-K with the tripeptide AGF. FIG. 12D shows the EIC analysis of the deprotected conjugate (Compound 1-K)-AGF after removing both Boc protecting groups from the Boc protected conjugate (Compound 1-K)-AGF. FIG. 12E shows the EIC analysis of the dipeptide GF after cleaving off the N-terminal amino acid by the guanidinylating reagent Compound 1-K.DETAILED DESCRIPTION

[0063] While various embodiments of the disclosure have been shown and described herein, it will be obvious to those skilled in the art that such embodiments are provided by way of example only. Numerous variations, changes, and substitutions may occur to those skilled in the art without departing from the disclosure. It should be understood that various alternatives to the embodiments of the disclosure described herein may be employed.Definitions

[0064] Whenever the term “at least,” “greater than,” or “greater than or equal to” precedes the first numerical value in a series of two or more numerical values, the term “at least,” “greater than” or “greater than or equal to” applies to each of the numerical values in that series of numerical values. For example, greater than or equal to 1, 2, or 3 is equivalent to greater than or equal to 1, greater than or equal to 2, or greater than or equal to 3.

[0065] Whenever the term “no more than,” “less than,” or “less than or equal to” precedes the first numerical value in a series of two or more numerical values, the term “no more than,” “less than,” or “less than or equal to” applies to each of the numerical values in that series ofnumerical values. For example, less than or equal to 3, 2, or 1 is equivalent to less than or equal to 3, less than or equal to 2, or less than or equal to 1.

[0066] References to “one embodiment,” “an embodiment,” “example embodiment,” “some embodiments,” “certain embodiments,” “various embodiments,” etc., indicate that the embodiment s) of the disclosed technology so described may include a particular feature, structure, or characteristic, but not every embodiment necessarily includes the particular feature, structure, or characteristic. Further, repeated use of the phrase “in one embodiment” does not necessarily refer to the same embodiment, although it may.

[0067] Ranges may be expressed herein as from “about” or “approximately” or “substantially” one particular value and / or to “about” or “approximately” or “substantially” another particular value. When such a range is expressed, other exemplary embodiments include from the one particular value and / or to the other particular value. Further, the term “about” means within an acceptable error range for the particular value as determined by one of ordinary skill in the art, which will depend in part on how the value is measured or determined, i.e., the limitations of the measurement system. For example, “about” can mean within an acceptable standard deviation, per the practice in the art. Alternatively, “about” can mean a range of up to ±20%, preferably up to ±10%, more preferably up to ±5%, and more preferably still up to ±1% of a given value. Alternatively, particularly with respect to biological systems or processes, the term can mean within an order of magnitude, preferably within 2-fold, of a value. Where particular values are described in the application and claims, unless otherwise stated, the term “about” is implicit and in this context means within an acceptable error range for the particular value.

[0068] By “comprising” or “containing” or “including” is meant that at least the named compound, element, particle, or method step is present in the composition or article or method, but does not exclude the presence of other compounds, materials, particles, method steps, even if the other such compounds, material, particles, method steps have the same function as what is named.

[0069] Throughout this description, various components may be identified having specific values or parameters, however, these items are provided as exemplary embodiments. Indeed, the exemplary embodiments do not limit the various aspects and concepts of the present disclosure as many comparable parameters, sizes, ranges, and / or values may be implemented. The terms “first,” “second,” and the like, “primary,” “secondary,” and the like, do not denote any order, quantity, or importance, but rather are used to distinguish one element from another.

[0070] As used herein, the term “protein” generally refers to a molecule comprising two or more amino acids joined by a peptide bond. A protein may also be referred to as a “polypeptide”, “oligopeptide”, or “peptide”. A protein can be a naturally occurring molecule, or a synthetic molecule (e.g., an artificial protein, peptide, enzyme). A protein may include one or more non-natural amino acids, modified amino acids, or non-amino acid linkers. A protein may contain D-amino acid enantiomers, L- amino acid enantiomers or both. Amino acids of a protein may be modified naturally or synthetically, such as by post-translational modifications or by chemical modification. In some circumstances, different proteins may be distinguished from each other based on different genes from which they are expressed in an organism, different primary sequence length or different primary sequence composition. Proteins expressed from the same gene may nonetheless be different proteoforms, for example, being distinguished based on non-identical length, non-identical amino acid sequence or non-identical post-translational modifications. Different proteins can be distinguished based on one or both of gene of origin and proteoform state.

[0071] As used herein, the term “peptide” may refer to any short, single peptide chain. A peptide may be no more than about 100, 95, 90, 85, 80, 75, 70, 65, 60, 55, 50, 45, 40, 35, 30, 25, 20, 15, 10, 5, or less than about 5 amino acids in length. A peptide may have a known or unknown biological function or activity. Peptides can include natural, synthetic, modified, or degraded proteins or peptides, or a combination thereof. Peptides can include proteinogenic, natural, synthetic, or modified amino acids or amino acid residues, or a combination thereof.

[0072] As used herein, the term “single analyte” may refer to an analyte that is individually manipulated or distinguished from other analytes. A single analyte may comprise a biomolecule or a synthetic molecule. A single analyte may comprise a small molecule. A single analyte can be a single molecule (e.g., a single biomolecule such as a single protein, nucleic acid molecule, affinity reagent, lipid, carbohydrate, metabolite, hapten, small molecule, pharmaceutical compound, nanoparticle, amino acid derivative, synthetic amino acid, etc.), a single complex of two or more molecules (e.g., a multimeric protein having two or more separable subunits, a single protein attached to a nucleic acid molecule or a single protein attached to an affinity reagent), a single particle, or the like. Reference herein to a “single analyte” in the context of a composition, system or method herein does not necessarily exclude application of the composition, system or method to multiple single analytes that are manipulated or distinguished individually, unless indicated contextually or explicitly to the contrary.

[0073] As used herein, “polypeptide” refers to two or more amino acids linked together by a peptide bond. The term “polypeptide” includes proteins that have a C-terminal end and an N-terminal end as generally known in the art and may be synthetic in origin or naturally occurring. As used herein “at least a portion of the polypeptide” refers to 2 or more amino acids of the polypeptide. A polypeptide may comprise one or more peptides. Optionally, a portion of the polypeptide includes at least: 1, 5, 10, 20, 30 or 50 amino acids, either consecutive or with gaps, of the complete amino acid sequence of the polypeptide, or the full amino acid sequence of the polypeptide.

[0074] As used herein, “affixed” refers to a connection between two molecules that are held in physical proximity. The term “affixed” encompasses both an indirect or direct connection and may be reversible or irreversible, for example the connection is optionally a covalent bond or a non-covalent bond.

[0075] As used herein, the term “sample” refers to a collected substance or material that comprises or is suspected to comprise one or more analytes of interest (e.g., biomolecules, e.g., polypeptides). A sample may be modified for purposes such as storage or stability. A sample may be naturally occurring or synthetic. A sample may be processed to separate or remove unwanted fractions or impurities from the analyte(s) of interest. A sample may be enriched or purified. For example, a sample may comprise a fraction of a separation process (e.g., chromatography, fractionation, electrophoresis, etc.). Alternatively, a sample may not be subjected to processing that separates or removes any unwanted fractions or impurities from the analyte(s) of interest. A sample may be obtained from any suitable source or location, including from organisms, cells, tissues, cell preparations, cell-free compositions, the environment (e.g., air, water, dirt, soil, agriculture, soil, dust, sewage). A sample may be obtained from an organism or part of an organism, such as from a fluid, tissue, or cell. A sample may include biological and / or non-biological components. As used herein, the terms “biological sample” or “biological source” refer to a sample that is derived from a predominantly biological system or organism, such as one or more viral particles, cells (e.g. individualized cells), organelles (e.g. individualized organelles), tissues, organs, bodily fluids, bone, cartilage, and exoskeleton. A biological sample may comprise a prokaryotic cell (e.g., bacteria) or eukaryotic cell (e.g., fungus, protist, algae, plant, animal). A biological sample may comprise a majority of biological material on a mass basis, excluding the weight of fluid within the sample. Biological samples may comprise one or more proteins, referred to herein as protein samples. Biological samples can be acquired from various sources, e.g., from a clinical patient sample, such as blood, serum, plasma, Cerebral Spinal Fluid (CSF), saliva, mucosal secretions, sputum, urine, lymph, perspiration, vaginal fluid, semen, fecal matter, amniotic fluid, perspiration, synovial fluid, fine needle aspirates, a tissue biopsy, a tumor biopsy, etc. A biological sample may be processed topurify and retain one or more biomolecules (e.g., proteins, nucleic acids, carbohydrates, lipids, glycoproteins, lipoproteins, metabolites, etc.) from the biological sample. A biological sample (e.g., a protein sample) may be derived from cultured cells, which may be treated or untreated. A biological sample (e.g., a protein sample) can also result from tissue specimens, such as biopsy samples, which may optionally be processed to liberate biomolecules (e.g., proteins) contained therein. Tissue samples may also be derived from in vivo specimens, including fresh, frozen, acute, and fixed tissues. A sample or a biological sample may comprise non-biological molecules, including but not limited to nanoparticles, polymers, haptens, small molecules, chemicals, fluorescent reagents, inert materials, pharmaceuticals, food additives, environmental contaminants, solvents, industrial chemicals, nanomaterials, radioisotopes, and / or by-products from non-biological molecules.

[0076] As used herein, the terms “antibody” and “immunoglobulin” may generally refer to proteins that can recognize and bind to a specific antigen. An antibody or immunoglobulin may refer to an antibody isotype, fragments of antibodies including, but not limited to, Fab, Fv, scFv, vHH, and Fd fragments, chimeric antibodies, humanized antibodies, single-chain antibodies, and fusion proteins including an antigen-binding portion of an antibody and a non-antibody protein. The antibodies may be detectably labeled, e.g., with a fluorophore, radioisotope, enzyme (e.g., a peroxidase), epitope tag, which generates a detectable product, fluorescent protein, nucleic acid barcode sequence, and the like. The antibodies may be further conjugated to other moieties, such as members of specific binding pairs, e.g., biotin (member of biotin-avidin specific binding pair), and the like. Also encompassed by the terms are nanobodies, Fab', Fv, F(ab')2, scFv, and other antibody fragments that retain specific binding to antigen. Antibodies may exist in a variety of other forms including, for example, Fv, Fab, and (Fab)2, diabodies, monobodies, single domain antibodies (sdAb), as well as bi-functional (i.e., bi-specific, e.g., bi-specific T-cell engager) hybrid antibodies (e.g., Lanzavecchia et al., Eur. J. Immunol. 17, 105 (1987)) and in single chains (e.g., Huston et al., Proc. Natl. Acad. Sci. U.S.A., 85, 5879-5883 (1988) and Bird et al., Science, 242, 423-426 (1988), which are incorporated herein by reference). (See, generally, Hood et al., Immunology, Benjamin, N.Y., 2nd ed. (1984), and Hunkapiller and Hood, Nature, 323, 15-16 (1986), which are herein incorporated by reference). Naturally occurring immunoglobulins or antibody types include immunoglobulin A, immunoglobulin G, immunoglobulin D, immunoglobulin E, immunoglobulin M, or other immunoreactive components.

[0077] “Binding” or “coupling” as used herein generally refers to a covalent or non-covalent interaction between two molecules (referred to herein as “binding partners”, e.g., a substrate andan enzyme or an antibody and an epitope). Binding between binding partners may be specific or non-specific. Binding between binding partners may involve one or more additional molecules (e.g., biomolecules) or enhancer molecules or substrates.

[0078] As used herein, “specifically binds” or “binds specifically” generally refers to an interaction between binding partners (e.g., a binding partner and a cognate molecule) such that the binding partners bind to one another, but do not bind to other molecules that may be present in the environment (e.g., in a biological sample, in tissue, in an in vitro assay) under a set of conditions. A specific binding interaction may entail a binding partner that binds to a cognate molecule. The specific binding interaction may entail the binding of the binding partner to its cognate molecule at a significantly or substantially higher level or with greater affinity as compared to the binding of the binding partner to a non-cognate molecule. A specific binding interaction may entail a first binding partner that has greater selectivity of binding to the cognate molecule as compared to a non-cognate molecule.

[0079] The terms “nucleic acid”, “nucleic acid molecule”, “oligonucleotide” and “polynucleotide” may be used interchangeably herein and generally refer to a polymeric form of naturally occurring or synthetic nucleotides, or analogs thereof, of any length. A nucleic acid molecule may comprise one or more deoxyribonucleotides, deoxynucleotide triphosphates, dideoxynucleotide triphosphates, deoxynucleotide hexaphosphates, dideoxynucleotide hexaphosphates, ribonucleotides, hexitol nucleotides, cyclohexane nucleotides, or analogs or combinations thereof. A nucleic acid molecule may comprise, e.g., DNA, RNA, HNA, CeNA, and modified forms thereof. A nucleic acid molecule may comprise nucleotides that are linked by phosphodiester bonds. A nucleic acid molecule may have any two- or three-dimensional structure, and may perform any function, known or unknown. A nucleic acid molecule may be single stranded, double stranded, or partially double stranded. Non-limiting examples of polynucleotides include a gene, a gene fragment, exons, introns, messenger RNA (mRNA), transfer RNA, ribosomal RNA, noncoding RNA, small interfering RNA, short hairpin RNA, micro RNA, scaRNA, ribozymes, riboswitches, viral RNA, complementary DNA (cDNA), cosmid DNA, mitochondrial DNA, chromosomal or genomic DNA, viral DNA, recombinant polynucleotides, branched polynucleotides, plasmids, vectors, isolated DNA of any sequence, control regions, isolated RNA of any sequence, nucleic acid probes, nucleic acid adapters, and primers. The nucleic acid molecule may be linear, circular, or any other geometry. Examples of polynucleotide analogs include but are not limited to xeno nucleic acid (XNA), bridged nucleic acid (BNA), glycol nucleic acid (GNA), hexitol nucleic acid (HNA), cyclohexane nucleic acid (CeNA), 2’-F-Arabinonucleic acids (2’-F-ANA), peptide nucleic acids (PNAs), yPNAs,morpholino polynucleotides, locked nucleic acids (LNAs), threose nucleic acid (TNA), 2'-O- Methyl polynucleotides, 2'-O-alkyl ribosyl substituted polynucleotides, phosphorothioate polynucleotides, and boronophosphate polynucleotides. A polynucleotide analog may possess purine or pyrimidine analogs, including for example, 7-deaza purine analogs, 8-halopurine analogs, 5-halopyrimidine analogs, inverted base, or universal base analogs that can pair with any base, including hypoxanthine, nitroazoles, isocarbostyril analogues, azole carboxamides, and aromatic triazole analogues, or base analogs with additional functionality, such as a biotin moiety for affinity binding. In addition, a nucleic acid molecule may exist permanently or transitionally in single-stranded or double-stranded form, including homoduplex, heteroduplex, and hybrid states.

[0080] As used herein, the term “amino acid” generally refers to an organic compound that combines to form a protein or peptide. An amino acid generally comprises an amine group, a carboxylic acid group, and a side-chain specific to each amino acid, which serve as a monomeric subunit of a peptide. An amino acid may include the 20 standard, naturally occurring or canonical amino acids as well as non-standard or non-canonical amino acids. The standard, naturally-occurring or canonical amino acids include Alanine (A or Ala), Cysteine (C or Cys), Aspartic Acid (D or Asp), Glutamic Acid (E or Glu), Phenylalanine (F or Phe), Glycine (G or Gly), Histidine (H or His), Isoleucine (I or He), Lysine (K or Lys), Leucine (L or Leu), Methionine (M or Met), Asparagine (N or Asn), Proline (P or Pro), Glutamine (Q or Gin), Arginine (R or Arg), Serine (S or Ser), Threonine (T or Thr), Valine (V or Vai), Tryptophan (W or Trp), and Tyrosine (Y or Tyr). An amino acid may be an L-amino acid or a D-amino acid. Non-standard amino acids may be modified amino acids, amino acid analogs, amino acid mimetics, non-standard proteinogenic amino acids, or non-proteinogenic amino acids that occur naturally or are chemically synthesized. Examples of non-standard amino acids include, but are not limited to, selenocysteine, pyrrolysine, and N-formylmethionine, (3 -amino acids, Homoamino acids, Proline and Pyruvic acid derivatives, 3-substituted alanine derivatives, glycine derivatives, ring-substituted phenylalanine and tyrosine derivatives, linear core amino acids, and N-methyl amino acids.

[0081] As used herein, the term “amino acid type” generally refers to one of the standard, naturally-occurring or canonical amino acids, e.g., one member of the group consisting of Alanine (A or Ala), Cysteine (C or Cys), Aspartic Acid (D or Asp), Glutamic Acid (E or Glu), Phenylalanine (F or Phe), Glycine (G or Gly), Histidine (H or His), Isoleucine (I or He), Lysine (K or Lys), Leucine (L or Leu), Methionine (M or Met), Asparagine (N or Asn), Proline (P or Pro), Glutamine (Q or Gin), Arginine (R or Arg), Serine (S or Ser), Threonine (T or Thr), Valine(V or Vai), Tryptophan (W or Trp), Tyrosine (Y or Tyr), derivatives thereof, and modified forms of any of the aforementioned amino acids. The term “amino acid type” may be used herein to distinguish a plurality of amino acids that comprise different side chain groups, rather than a plurality of amino acids that are identical (e.g., different positional amino acids of a single peptide that have the same side chain). An amino acid type may comprise a modified version of one of the standard, naturally-occurring or canonical amino acids, e.g., post translational modifications, an epigenetic modification, or chemical or enzymatic modifications. In some instances, an amino acid type can include non-canonical amino acids.

[0082] As used herein, the term “post-translational modification” refers to modifications that occur on a peptide subsequent to translation. A post-translational modification may be a covalent modification or enzymatic modification. Examples of post-translation modifications include, but are not limited to, acylation, acetylation, alkylation (including methylation), benzoylation, biotinylation, butyrylation, carbamylation, carbonylation, carboxylation, crotonylation, deamidation, deiminiation, dimethylation, diphthamide formation, disulfide bridge formation, eliminylation, flavin attachment, formylation, gamma-carboxylation, glutamyl ati on, glutarylation, glycylation, glycosylation, glypiation, heme C attachment, hydroxylation, hypusine formation, iodination, isoprenylation, lipidation, lipoylation, malonylation, methylation, myristolylation, nitration, oxidation, transglutamination, palmitoylation, pegylation, phosphopantetheinylation, phosphorylation, prenylation, propionylation, pyroglutamate formation, retinylidene Schiff base formation, S- glutathionylation, S-nitrosylation, S-sulfenylation, S-adenosylation, sulfation, selenation, stearoylation, succinylation, sulfination, trimethylation, ubiquitination, and C-terminal amidation. A post-translational modification includes modifications of the amino terminus and / or the carboxyl terminus of a peptide. Modifications (both naturally occurring and synthetic) of the terminal amino group include, but are not limited to, des-amino, N-lower alkyl, N-di- lower alkyl, and N-acyl modifications, N-terminal cyclization, deamination, oxidation, ubiquitination, SUMOylation, Neddylation, ISGylation, pupylation, eliminylation, biotinylation, lipidation, N-terminal methylation, N-terminal acetylation, N-terminal propionylation, N- terminal butyrylation, N-terminal crotonylation, N-terminal myristoylation, N-terminal palmitoylation, N-terminal stearoylation, and N-terminal benzoylation. Modifications of the terminal carboxy group include, but are not limited to, amide, lower alkyl amide, dialkyl amide, and lower alkyl ester modifications (e.g., wherein lower alkyl is C1-C4 alkyl). A post- translational modification also includes modifications, such as but not limited to those described above, of amino acids falling between the amino and carboxy termini. The term post-translational modification can also include peptide modifications that include one or more detectable labels. A post-translational modification may be naturally occurring or synthetic.

[0083] As used herein, the term “binding agent” refers to a molecule, e.g., a nucleic acid molecule, a peptide, a polypeptide, a protein, carbohydrate, a synthetic molecule, or a small molecule that binds to, associates with, unites with, recognizes, or combines with another molecule. The binding agent may bind to a macromolecule or a component or feature of a macromolecule. A binding agent may form a covalent association or non-covalent association with a molecule, a macromolecule, or a component or feature of a macromolecule. A binding agent may also be a chimeric binding agent, composed of two or more types of molecules, such as a nucleic acid molecule-peptide chimeric binding agent, a carbohydrate-peptide chimeric binding agent, or a lipid-peptide chimeric binding agent. A binding agent may be a naturally occurring, synthetically produced, or recombinantly expressed molecule. A binding agent may bind to a single monomer or subunit of a polymeric analyte, such as a macromolecule (e.g., a single amino acid of a peptide) or bind to a plurality of linked subunits of a macromolecule (e.g., a dipeptide, tripeptide, or higher order peptide of a longer peptide, polypeptide, or protein molecule). A binding agent may bind to a linear molecule or a molecule having a three- dimensional structure (also referred to as conformation). For example, an antibody binding agent may bind to linear peptide, polypeptide, or protein, or bind to a conformational peptide, polypeptide, or protein. A binding agent may bind to an N-terminal peptide, a C-terminal peptide, or an intervening peptide of a peptide, polypeptide, or protein molecule. A binding agent may bind to an N-terminal amino acid, C-terminal amino acid, or an intervening amino acid of a peptide molecule. A binding agent may preferably bind to a chemically modified or labeled amino acid over a non-modified or unlabeled amino acid. For example, a binding agent may preferably bind to an amino acid that has been modified with an acetyl moiety, guanyl moiety, dansyl moiety, PTC moiety, DNP moiety, SNP moiety, etc., over an amino acid that does not possess such a moiety. A binding agent may bind to a post-translational modification, either naturally occurring or synthetic, of a peptide molecule. A binding agent may exhibit selective binding to a component or feature of a macromolecule (e.g., a binding agent may selectively bind to one of the 20 possible natural amino acid residues and bind with very low affinity or not at all to the other 19 natural amino acid residues). A binding agent may exhibit less selective binding, where the binding agent is capable of binding a plurality of components or features of a macromolecule (e.g., a binding agent may bind with similar affinity to two or more different amino acid residues). A binding agent may comprise a tag, which may be coupled to the binding agent via a linker.

[0084] As used herein, the term “linker” generally refers to a molecule or moiety that is involved in joining two or more molecules. A linker may facilitate a covalent or noncovalent interaction of two or more molecules. A linker may be a crosslinker. The linker can be unifunctional, bifunctional, trifunctional, quadrifunctional, or polyfunctional. A linker can be chiral or achiral, or may contain one or more chiral centers, to influence enzymatic or chemical reactivity, conformational dynamics, steric interactions, or electronic properties of the linked molecules. A linker can be or comprise a nucleotide, a nucleotide analog, an amino acid, a peptide, a polypeptide, or a non-nucleotide chemical moiety, such as an organic or inorganic compound. A linker may comprise a polymer, such as a polyethylene glycol (PEG), polyethylene (PE), polypropylene (PP), polyvinyl chloride (PVC), polystyrene (PS) poly-L- lysine (PLL), poly (DL-lactic acid) (PLA), poly (DL-lactide-co-glycoside) (PLGA), polyomithine, polyarginine, or other organic or inorganic polymer. A linker may comprise one or more reactive ends, e.g, an amine-reactive group, a carboxyl -reactive group, a sulfhydrylreactive group, a hydroxyl-reactive group, etc. Alternatively, a linker may not comprise a reactive end. In some examples, a linker may be used to join different molecule types, e.g, different biomolecule types such as a peptide with a nucleic acid molecule, a lipid with a peptide, a carbohydrate with a peptide, etc.; non-biomolecule types; or a biomolecule to a nonbiomolecule. For example, a linker may be used to join a binding agent with a tag, a tag with a macromolecule (e.g., peptide, nucleic acid molecule), a macromolecule with a solid support, a tag with a solid support, etc. A linker may join two molecules via enzymatic reaction or chemistry reaction (e.g., click chemistry). A linker may join more than two molecules, e.g., via enzymatic or chemical reactions. A linker may influence reaction kinetics, yield, or product specificity by modulating molecular proximity, steric hindrance, or electronic environment. For example, a linker may enhance or inhibit reaction rates by controlling the spatial arrangement of reactants, stabilizing intermediates, or facilitating catalytic interactions. In some reactions, a linker may stabilize or position reaction intermediates in a favorable orientation, thereby influencing reaction efficiency, pathway selection, or product distribution. In chemical synthesis, a linker may dictate regioselectivity or stereoselectivity by constraining molecular conformation. Additionally, a linker may participate in dynamic structural changes, enabling conformational flexibility or rigidity to promote or suppress specific reaction pathways. A linker may modulate product stability, for example, by reducing susceptibility to hydrolysis or degradation. In some cases, a linker may introduce functional groups that participate in subsequent transformations, thereby influencing multi-step reaction cascades. A linker can be relatively linear or non-linear, e.g, cyclic or circularized, branched, polygonal, etc.

[0085] The term “conjugated” as used herein generally refers to a covalent or ionic interaction between two entities, e.g., molecules, compounds, or combinations thereof.

[0086] As used herein, the term “tag” generally refers to a molecule or moiety that is conjugated to a molecule. A tag may comprise a detectable label, e.g., a fluorophore or fluorescent protein, a radioactive isotope, an enzyme (e.g., a chromogenic or fluorescent protein, proteins that can catalyze chromogenic substrates), a mass tag, a hapten (e.g., biotin, digoxigenin, urushiol, fluorescein), a vibrational or FTIR tag (e.g., alkyne group). A tag may comprise a biomolecule, such as a nucleic acid molecule, a protein, a lipid, a carbohydrate, or a combination thereof. A tag may comprise one or more nucleic acid molecules, which may optionally encode information regarding the tag or the molecule onto which a tag is conjugated (e.g., a binding agent, such as an antibody). For example, a tag may comprise a nucleic acid barcode molecule. A tag may comprise an organic compound or an inorganic compound. As used herein, the term “tag” may also refer to a patterned sequence of signals, wherein signals appear, skip, or disappear at defined intervals, creating a recognizable marker or reference within the signal data. This patterned appearance may encompass periodic or non-periodic intervals between signals, selective omissions, or complex structured combinations of signal presence and absence that collectively form a unique, detectable signature. For example, a tag may comprise a nucleic acid barcode molecule that has a series of distinct vibrational signatures detectable by techniques such as Raman or FTIR spectroscopy.

[0087] As used herein, the term “nucleic acid sequence” or “oligonucleotide sequence” generally refers to a contiguous string of nucleotide bases and may refer to the particular placement of nucleotide bases in relation to each other as they appear in an oligonucleotide. Similarly, the term “polypeptide sequence” or “amino acid sequence” refers to a contiguous string of amino acids and may refer to the particular placement of amino acids in relation to each other as they appear in a polypeptide.

[0088] A “nucleic acid molecule” according to the present invention may include any polymer or oligomer of nucleotides such as pyrimidine and purine bases, such as cytosine, thymine, and uracil, and adenine and guanine, respectively and combinations thereof. The nucleotide sequence may comprise any deoxyribonucleotide, ribonucleotide, hexitol -nucleotide, cyclohexane-nucleotide, peptide nucleic acid component, and any chemical variants thereof, such as methylated, 7-deaza purine analogs, 8-hal opurine analogs, hydroxymethylated or glycosylated forms of these bases, and the like. The polymers or oligomers may be heterogeneous or homogenous in composition, and may be isolated from naturally occurring sources or may be artificially or synthetically produced. A nucleic acid molecule may compriseDNA, RNA, HNA, CeNA or a mixture thereof, and may exist permanently or transitionally in single-stranded or double-stranded form, including homoduplex, heteroduplex, and hybrid states.

[0089] The terms “complementary” or “complementarity” refer to polynucleotides (z.e., a sequence of nucleotides) related by base-pairing rules. For example, the sequence “5'-AGT-3',” is complementary to the sequence “5'- ACT-3'”. Complementarity may be “partial,” in which only some of the nucleic acids’ bases are matched according to the base pairing rules, or there may be “complete” or “total” complementarity between the nucleic acids. The degree of complementarity between nucleic acid strands can have significant effects on the efficiency and strength of hybridization between nucleic acid strands under defined conditions.

[0090] As used herein, the term “hybridization” is used in reference to the pairing of complementary nucleic acids. Hybridization and the strength of hybridization (e.g., the strength of the association between the nucleic acids) is influenced by such factors as the degree of complementary between the nucleic acids, stringency of the conditions involved, and the melting temperature of the formed hybrid. Hybridization methods involve the annealing of one nucleic acid to another, complementary nucleic acid, e.g., based on Watson-Crick base pairing.

[0091] As used herein, the term “proteomics” generally refers to quantitative and / or qualitative analysis of the proteome within a sample, such as biological sample, e.g., from cells, tissues, or bodily fluids. Proteomics may include the analysis of spatial distributions of proteins within a sample (e.g., cell and / or tissues). Proteomics may include studies of the dynamic state of the proteome, e.g., how one or more proteins change in time. A proteome may comprise multiple “-omes”, e.g., a kinome; a secretome; a receptome (e.g., GPCRome); an immunoproteome; a nutriproteome; a proteome subset defined by a post-translational modification (e.g., phosphorylation, ubiquitination, methylation, acetylation, glycosylation, oxidation, lipidation, and / or nitrosylation), such as a phosphoproteome (e.g., phosphotyrosine- proteome, tyrosine-kinome, and tyrosine-phosphatome), a glycoproteome, etc.; a proteome subset associated with a tissue or organ, a developmental stage, or a physiological or pathological condition; a proteome subset associated a cellular process, such as cell cycle, differentiation (or de-differentiation), cell death, senescence, cell migration, transformation, or metastasis; or any combination thereof.

[0092] The terminal amino acid at one end of the peptide chain that has a free amino group may be referred to herein as the “N-terminal amino acid” (NTAA). The terminal amino acid at the other end of the chain that has a free carboxyl group may be referred to herein as the “C- terminal amino acid” (CTAA). The amino acids making up a peptide may be numbered in order,with the peptide being “n” amino acids in length. As used herein, in some instances, NTAA may be considered the nth amino acid (also referred to herein as the “n NTAA”). In such cases, the next amino acid is the n-1 amino acid, then the n-2 amino acid, and so on down the length of the peptide from the N-terminal end to C-terminal end. Alternatively, CTAA may be considered the nth amino acid (also referred to herein as the “n CTAA”). In such cases, the next amino acid is the n-1, then the n-2 amino acid, and so on down the length of the peptide from the C-terminal end to N-terminal end. An NTAA, CTAA, or both may be modified or labeled with a chemical moiety.

[0093] As used herein, the terms “determining,” “measuring,” “assessing,” and “assaying” are used interchangeably and include both quantitative and qualitative determinations.

[0094] As used herein, the term “unique molecular identifier” or “UMI” generally refers to a molecule barcode comprising indexing information. A UMI may comprise a nucleic acid molecule of about 3 to about 150 bases (3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45,46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71,72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97,98, 99, 100, 105, 110, 115, 120, 125, 130, 135, 140, 145, or 150 bases) in length. A UMI may provide a unique identifier tag for each molecule (e.g., peptide, binding agent, a nucleic acid molecule) that comprises or is coupled to a UMI. A UMI may comprise a random sequence (e.g. a random N-mer).

[0095] As used herein, a “derivative” of a nucleic acid molecule generally refers to a nucleic acid molecule that is derived from an originating nucleic acid molecule. The derivative may have the same or substantially the same nucleotide sequence as the originating nucleic acid molecule, or the derivative may comprise a complement or partial complement as the originating nucleic acid molecule. A derivative may be the same type of nucleic acid (e.g., DNA or RNA) as the originating nucleic acid molecule, or the derivative may be a different type of nucleic acid (e.g., cDNA generated from an RNA molecule). A nucleic acid molecule derivative may display sequence identity as the originating nucleic acid molecule. The derivative nucleic acid molecule may also be subjected to additional processing from the originating nucleic acid molecule, e.g., chemical or enzymatic modification, splicing, ligation, polymerization, fragmentation, tagmentation (e.g., using a transposase), digestion, etc.

[0096] A derivative polypeptide or peptide may be derived from an originating polypeptide (or peptide). A derivative may comprise the same amino acid sequence as the originating polypeptide, or the sequence may be different. The derivative polypeptide may result from or besubjected to additional processing from the originating polypeptide, e.g., chemical or enzymatic modification. The derivative polypeptide may comprise one or more tags, nucleic acid molecules, barcode molecules, labels (e.g., detectable labels), fluorophores, probes, linkers, post-translational modifications, chemical protecting groups, or other chemical moieties.

[0097] As used herein, the term “compartment” or “partition” generally refers to a physical area or volume that separates or isolates a subset of molecules from a sample of molecules. For example, a compartment or partition may separate an individual cell from other cells, or a subset of a sample’s proteome from the rest of the sample’s proteome. A compartment or partition may be an aqueous compartment (e.g., microfluidic droplet), a solid compartment (e.g., picotiter well or microtiter well on a plate, tube, vial, gel bead), a liquid-liquid phase separation, a liquid condensate, a subcellular region, or a separated region on a surface. A compartment may comprise one or more beads to which macromolecules may be immobilized. A compartment may be transient.

[0098] In some embodiments, the partitions as described herein are droplets. The terms “drop,” “droplet,” and “microdroplet” are used interchangeably herein, to refer to small, generally spherically structures, containing at least a first fluid phase, e.g., an aqueous phase (e.g., water), bounded by a second fluid phase (e.g., oil) which is immiscible with the first fluid phase. In some embodiments, droplets according to the present disclosure may contain a first fluid phase, e.g., oil, bounded by a second immiscible fluid phase, e.g., an aqueous phase fluid (e.g., water). In some embodiments, the second fluid phase will be an immiscible phase carrier fluid. Thus, droplets according to the present disclosure may be provided as aqueous-in-oil emulsions or oil-in-aqueous emulsions. Droplets may be sized and / or shaped as described herein for discrete entities. For example, droplets according to the present disclosure generally range from 1 pm to 1000 pm, inclusive, in diameter. Droplets according to the present disclosure may be used to encapsulate cells, nucleic acids (e.g., DNA), enzymes, reagents, and a variety of other components. The term droplet may be used to refer to a droplet produced in, on, or by a microfluidic device and / or flowed from or applied by a microfluidic device.

[0099] As used herein, the term “carrier fluid” refers to a fluid configured or selected to contain one or more discrete entities, e.g., droplets, as described herein. A carrier fluid may include one or more substances and may have one or more properties, e.g., viscosity, which allow it to be flowed through a microfluidic device or a portion thereof, such as a delivery orifice. In some embodiments, carrier fluids include, for example: oil or water, and may be in a liquid or gas phase. Suitable carrier fluids are described in greater detail herein.

[0100] As used herein, the term “solid support”, “solid surface”, or “solid substrate” or “substrate” refers to any solid material, including porous and non-porous materials, to which a molecule can be associated directly or indirectly. The molecule may be associated with the substrate by covalent or non-covalent interactions, or a combination thereof. A substrate may be two-dimensional (e.g., planar surface) or three-dimensional (e.g., gel matrix or bead). A solid support may comprise, in non-limiting examples, a bead, a microbead, an array, a glass surface, a silicon surface, a plastic surface, a filter, a membrane, nylon or other polymer, a silicon wafer chip, a flow through chip, a flow cell, a microfluidic device or chip or a surface thereof, a biochip including signal transducing electronics, a channel, a microtiter well, an ELISA plate, a spinning interferometry disc, a nitrocellulose membrane, a nitrocellulose-based polymer surface, a polymer matrix, a nanoparticle, or a microsphere. Materials for a solid support include but are not limited to acrylamide, agarose, cellulose, nitrocellulose, glass, gold, quartz, polystyrene, polyethylene vinyl acetate, polypropylene, polymethacrylate, polyethylene, polyethylene oxide, polysilicates, polycarbonates, Teflon, fluorocarbons, nylon, silicon rubber, polyanhydrides, polyglycolic acid, polylactic acid, polyorthoesters, functionalized silane, polypropylfumerate, collagen, glycosaminoglycans, polyamino acids, dextran, or any combination thereof. Solid supports further include thin film, membrane, bottles, dishes, fibers, woven fibers, shaped polymers such as tubes, (e.g., nanotubes), particles, beads, DNA origami, microspheres, microparticles, or any combination thereof. For example, when solid surface is a bead, the bead can include, but is not limited to, a ceramic bead, polystyrene bead, a polymer bead, a methylstyrene bead, an agarose bead, an acrylamide bead, a solid core bead, a porous bead, a magnetic or paramagnetic bead, a glass bead, or a controlled pore bead, or a combination thereof. A magnetic bead may comprise or be composed of any useful material and may respond to an applied magnetic field. A bead may be spherical or an irregularly shaped. A bead’s size may range from nanometers, e.g., 1 nm, 10 nm, 100 nm, to millimeters, e.g., 1 mm. In certain embodiments, beads range in size from about 0.2 micron to about 200 microns, or from about 0.5 micron to about 5 microns. In some embodiments, beads can be about 1, 1.5, 2, 2.5, 2.8, 3, 3.5, 4, 4.5, 5, 5.5, 6, 6.5, 7, 7.5, 8, 8.5, 9, 9.5, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48,49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74,75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, or100 pm in diameter. In certain embodiments, “a bead” solid support may refer to an individual bead or a plurality of beads. A solid support may assume any useful geometry, e.g., pyramid, cube, cylinder, helix, sphere, spheroid, rod, disc, arrow, spring, teardrop, prism, tetrapod, or anyother useful geometry. A bead may be coated or treated with a range of substances or surface modifications to alter its physical, chemical, or biological properties.

[0101] As used herein, “sequencing” generally refers to determining the order and identity of: (A) nucleotides (base sequences) in a nucleic acid sample, e.g., DNA or RNA; or determining the order and identity of (B) amino acids in all or part of a polymer, such as a protein, peptide, or other multimeric molecule. Many techniques are available for nucleic acid sequencing, such as Sanger sequencing or High Throughput Sequencing technologies (HTS). Sanger sequencing may involve sequencing via detection through (capillary) electrophoresis, in which up to 384 capillaries may be sequence analyzed in one run. High throughput sequencing involves the parallel sequencing of thousands or millions or more sequences at once. HTS can be defined as Next Generation sequencing (NGS), i.e. techniques based on solid phase pyrosequencing or as Next-Next Generation sequencing based on single nucleotide real time sequencing (SMRT). HTS technologies are available such as offered by Roche, Illumina and Applied Biosystems (Life Technologies). Further high throughput sequencing technologies are described by and / or available from Helicos, Pacific Biosciences, Complete Genomics, Ion Torrent Systems, Oxford Nanopore Technologies, Nabsys, ZS Genetics, GnuBio. Additional sequencing methods include Raman sequencing and Infrared (IR) sequencing, which utilizes Raman spectroscopy or IR spectroscopy to detect molecular vibrations associated with specific nucleotide or amino acid sequences, enabling label-free sequencing based on unique vibrational energy signatures. Tunneling current sequencing identifies base sequences or amino acid sequences through electronic signal variations as nucleotides or amino acids pass through a nanoscale gap, detecting characteristic tunneling currents specific to each molecular component.

[0102] As used herein, “next generation sequencing” refers to high-throughput sequencing methods that allow the sequencing of millions to billions of molecules in parallel. Examples of next generation sequencing methods include sequencing by synthesis, sequencing by ligation, sequencing by hybridization, polony sequencing, ion semiconductor sequencing, nanopore sequencing, and pyrosequencing. By attaching primers to a solid substrate and a complementary sequence to a nucleic acid molecule, a nucleic acid molecule can be hybridized to the solid substrate via the primer and then multiple copies can be generated in a discrete area on the solid substrate by using polymerase to amplify (these groupings are sometimes referred to as polymerase colonies or polonies). Consequently, during the sequencing process, a nucleotide at a particular position can be sequenced multiple times (e.g., hundreds or thousands of times) — this depth of coverage is referred to as “deep sequencing.” Examples of high throughput nucleic acid sequencing technology include platforms provided by Illumina, BGI, Qiagen,ThermoFisher, and Roche, including formats such as parallel bead arrays, sequencing by synthesis, sequencing by ligation, capillary electrophoresis, electronic microchips, “biochips,” microarrays, parallel microchips, and single-molecule arrays, zero mode waveguide based sequencing, some of which are reviewed by Service (Science 311 : 1544-1546, 2006).

[0103] As used herein, “analyzing” the macromolecule means to quantify, characterize, distinguish, or a combination thereof, all or a portion of the components of a molecule (e.g., a macromolecule, a biological molecule such as a protein, amino acid, nucleic acid molecule, etc.). For example, analyzing a peptide, polypeptide, or protein may comprise determining all or a portion of the amino acid sequence (contiguous or non-continuous) of the peptide. Analyzing a macromolecule may include partial identification of a component of the macromolecule. For example, partial identification of amino acids in a protein sequence can identify an amino acid in the protein as belonging to a subset of possible amino acids. Analysis may be performed sequentially, e.g., beginning with analysis of the n NTAA, and then proceeding to the next amino acid of the peptide (i.e., n-1, n-2, n-3, and so forth). In such instances, sequencing may be performed by cleavage of the n NTAA, thereby converting the n-1 amino acid of the peptide to an N-terminal amino acid (referred to herein as the “n-1 NTAA”). Similarly, analysis of a peptide may begin from C-terminus towards the N-terminus with each round of cleavage from the C-terminus creating a new CTAA. Cleavage of the n CTAA converts the n-1 amino acid of the peptide to a C-terminal amino acid, referred to herein as an “n-1 CTAA”. Analyzing the peptide may also include determining a presence and frequency of post-translational modifications on the peptide, which may or may not include information regarding the sequential order of the post-translational modifications on the peptide. Analyzing the peptide may also include determining the presence and frequency of epitopes in the peptide, which may or may not include information regarding the sequential order or location of the epitopes within the peptide. Analyzing the peptide may include combining different types of analysis, for example obtaining epitope information, amino acid sequence information, post-translational modification information, or any combination thereof.

[0104] As used herein, the term “analyte” generally refers to a substance that is of interest to be further identified, characterized, or measured. An analyte can be, in non-limiting examples, an ion, chemical, compound, small molecule, element, particle, metal, biomolecule, macromolecule, metabolite, lipid, carbohydrate, peptide or protein, nucleic acid molecule, organelle, or cell. An analyte may be naturally occurring or synthetic. The analyte may be a solid, semi-solid, liquid, semi-liquid, gas, or plasma. The analyte may be characterized qualitatively or quantitatively. A portion of an analyte may be analyzed. For example, an analytemay be a peptide and the constituent amino acids may be analyzed. The analyte may comprise a polymer, also referred to herein as “polymeric analyte”, which generally refers to an analyte of interest that comprises one or more monomers. A polymeric analyte can be, in non-limiting examples, a group of ions, chemicals, compounds, small molecules, elements, particles, metals, or a biomolecule, macromolecule, metabolite, lipid, carbohydrate, peptide or protein, nucleic acid molecule, organelle, or cell.

[0105] As used herein, the term “array” generally refers to a population of molecules that is attached to one or more solid supports such that the molecules at one address can be distinguished from molecules at other addresses. An array can include different molecules that are each located at different addresses on a solid support. Alternatively, an array can include separate solid supports each functioning as an address that bears a different molecule, wherein the different molecules can be identified according to the locations of the solid supports on a surface to which the solid supports are attached, or according to the locations of the solid supports in a liquid such as a fluid stream. The molecules of the array can be, for example, nucleic acids such as SNAPs, polypeptides, proteins, peptides, oligopeptides, enzymes, ligands, or receptors such as antibodies, functional fragments of antibodies or aptamers. The addresses of an array can optionally be optically observable and, in some configurations, adjacent addresses can be optically distinguishable when detected using a method or apparatus set forth herein.

[0106] As used herein, the term “functionalized” refers to any material or substance that has been modified to include a functional group. A functionalized material or substance may be naturally or synthetically functionalized. For example, a polypeptide can be naturally functionalized with a phosphate group, oligosaccharide (e.g., glycosyl, glycosylphosphatidylinositol or phosphoglycosyl), nitrosyl, methyl, acetyl, lipid (e.g., glycosyl phosphatidylinositol, myristoyl or prenyl), ubiquitin or other naturally occurring post- translational modification. A functionalized material or substance may be functionalized for any given purpose, including altering chemical properties (e.g., altering hydrophobicity or changing surface charge density) or altering reactivity (e.g., capable of reacting with a moiety or reagent to form a covalent bond to the moiety or reagent).

[0107] As used herein, the term “click reaction,” “click chemistry,” or “bioorthogonal reaction” refers to single-step, thermodynamically favorable conjugation reaction utilizing biocompatible reagents. A click reaction may utilize no toxic or biologically incompatible reagents (e.g., acids, bases, heavy metals) or generate no toxic or biologically incompatible byproducts. A click reaction may utilize an aqueous solvent or buffer (e.g., phosphate buffersolution, Tris buffer, saline buffer, MOPS, etc.). A click reaction may be thermodynamically favorable if it has a negative Gibbs free energy of reaction, for example a Gibbs free energy of reaction of less than about -5 kiloJoules / mole (kJ / mol), -10 kJ / mol, -25 kJ / mol, -50 kJ / mol, -100 kJ / mol, -200 kJ / mol, -300 kJ / mol, -400 kJ / mol, or less than -500 kJ / mol. Exemplary bioorthogonal and click reactions are described in detail in WO 2019 / 195633A1, which is herein incorporated by reference in its entirety. Exemplary click reactions may include metal-catalyzed azide-alkyne cycloaddition, strain-promoted azide-alkyne cycloaddition, strain-promoted azide- nitrone cycloaddition, strained alkene reactions, thiolene reaction, Diels-Alder reaction, inverse electron demand Diels- Alder reaction, [3+2] cycloaddition, [4+1] cycloaddition, nucleophilic substitution, dihydroxylation, thiolyne reaction, photoclick, nitrone dipole cycloaddition, norbomene cycloaddition, oxanob ornadiene cycloaddition, tetrazine ligation, and tetrazole photoclick reactions. Exemplary functional groups or reactive handles utilized to perform click reactions (also referred to herein as “click chemistry moi eties”) may include alkenes (e.g., linear alkenes or cyclic alkenes such as trans-cyclooctene (TCO)), alkynes (e.g., linear alkynes or cycloalkynes (e.g., cyclooctynes or derivatives thereof, e.g., aza-dimethoxy cyclooctyne (DIMAC), symmetrical pyrrolocyclooctyne (SYPCO), pyrrolocyclooctyne (PYRROC), difluorocyclooctyne (DIFO), a,a-bis(trifluoromethyl)pyrrolocyclooctyne(TRIPCO), bicyclo[6.1.0]nonyne (BCN), dibenzocyclooctyne (DBCO), difluorinated cyclooctyne (DIFO), difluorobenzocyclooctyne (DIFBO), dibenzoazacyclo-octyne (DBACO), difluoro-aza-dibenzocyclooctyne (F2-DIBAC), biaryl-azacyclooctynone (BARAC), difluorodimethoxydibenzocyclooctynol (FMDIBO), difluorodimethoxydibenzocyclooctynone (keto-FMDIBO), and 3,3,6,6-tetramethylthiacycloheptyne (TMTH)), TMTH-sulfoximine (TMTHSI), azides, epoxides, amines, thiols, nitrones, isonitriles, isocyanides, aziridines, activated esters, and tetrazines, triazoles, and combinations, variations, or derivatives thereof. The click chemistry moi eties may be subjected to conditions sufficient to react the first click chemistry moiety to the second click chemistry moiety, e.g., provision of metal catalysts, appropriate solvents, pH, temperature, ionic concentration, or light / energy, for any useful duration of time.

[0108] As used herein, the terms “group” and “moiety” are intended to be synonymous when used in reference to the structure of a molecule. The terms refer to a component or part of the molecule. The terms do not necessarily denote the relative size of the component or part compared to the molecule, unless indicated otherwise. The terms do not necessarily denote the relative size of the component or part compared to any other component or part of the molecule, unless indicated otherwise. A group or moiety can contain one or more atoms.

[0109] As used herein, “primers” generally refer to nucleic acid molecules which can prime the synthesis of a nucleic acid molecule (e.g., DNA or RNA). A primer may be single stranded. A primer may comprise one or more recognition sites for a protein (e.g., a polymerizing enzyme, a restriction enzyme, a cleaving enzyme, a nuclease, etc.) to bind to the primer or a primer hybridized to a template strand. A primer may comprise DNA, RNA, or other nucleic acid analogs or noncanonical bases (e.g., spacer moi eties, uracils, abasic sites). A primer may optionally comprise any number of functional sequences such as sequencing primer sequences e.g., P5 or P7 sequences), sequencing primer-binding sequences, read sequences e.g., R1 or R2 sequences), restriction sites, nuclease-recognition sites, abasic sites, cleavage sites, transposition sites e.g., mosaic end sequences), a barcode sequence, a unique molecular identifier (UMI), etc.

[0110] “Amplification” or “amplifying” generally refers to a polynucleotide amplification reaction, namely, a population of polynucleotides that are replicated from one or more starting sequences. Amplifying may refer to a variety of amplification reactions, including but not limited to polymerase chain reaction (PCR), linear polymerase reactions, nucleic acid sequencebased amplification, rolling circle amplification and similar reactions. An amplification reaction may generate an amplicon. Amplification or amplifying may refer to an increase in quantity of a measurable output, for example, signal amplification.

[0111] An “adapter” as referred to herein, generally refers to a short nucleic acid molecule (e.g., about 10 to about 100 base pairs in length). An adapter may comprise a short doublestranded DNA molecule. An adapter may be attached, e.g., via polymerization or ligation, to an end of a DNA fragments or amplicons. Adapters may comprise synthetic oligonucleotides, e.g., oligonucleotides that have nucleotide sequences which are at least partially complementary to each other. An adapter may have blunt ends, may have staggered ends (also referred to herein as a 3’ or 5’ “overhang sequence” or “sticky end”, or a blunt end and a staggered end. Adapters may be attached (e.g., via ligation) to fragments to provide an adapter-ligated fragment; the adapter-ligated fragment may serve as a starting point for subsequent manipulation e.g., for amplification or sequencing. An adapter may be functionalized, e.g., conjugated with a tag, probe, detectable label, affinity capture reagent (e.g., biotin or streptavidin).

[0112] The term “capture moiety” as used herein generally refers to a molecule that is configured to be coupled to another moiety or molecule. A capture moiety can be a biomolecule, e.g., a lipid, carbohydrate, sugar, amino acid, peptide or protein, nucleotide, nucleic acid molecule, metabolite, or a combination thereof (e.g., glycoproteins, lipoproteins, glycosaminoglycans, etc.). A capture moiety can be a small molecule, organic compound, inorganic compound, metal, polymer, ion, or other molecule or molecular compound. A capturemoiety may comprise a macromolecule. A capture moiety may comprise an enzyme, antibody, antibody fragment, nanobody, aptamer, biotin, streptavidin, avidin, neutravidin, or analogs or derivatives thereof. A capture moiety may comprise more than one molecule, e.g., a dimer, trimer, tetramer, pentamer, hexamer, heptamer, octamer, etc. A capture moiety can be a solid substrate or part of a solid substrate, or the capture moiety can be separate from a substrate, e.g., in a fluidic medium (e.g., air, in a liquid solution). A capture moiety may have specificity to a binding partner or a plurality of binding partners. A capture moiety may be able to bind to one molecule or moiety (univalent), or a plurality of molecules or moieties (multivalent).

[0113] The terms “translocating” and “translocation” as used herein (along with variations, such as “translocate” or “translocates”), generally refers to the movement of a molecule through a medium (e.g., a gas, a liquid, a solid, or a multiphase medium). Translocation of a molecule may occur spontaneously (e.g., through diffusion, Brownian motion, etc.). Alternatively, or in addition to, translocation of a molecule may occur with an application of force or pressure, e.g., using frictional force, tension force, a normal force, air resistance force, spring force, a temperature gradient, gravitational force, electrical force, magnetic force, acoustic force (e.g., acoustophoresis) etc. In some examples, translocation of a molecule may be achieved by application of pressure-driven flow or electrophoretic forces. Translocation may occur through a liquid or through a solid or semi-solid substrate (e.g., through a pore or gap) or adjacent or in proximity to the solid or semi-solid substrate.

[0114] As used herein, the abbreviations for the natural 1 -enantiomeric amino acids are conventional and can be as follows: alanine (A, Ala); arginine (R, Arg); asparagine (N, Asn); aspartic acid (D, Asp); cysteine (C, Cys); glutamic acid (E, Glu); glutamine (Q, Gin); glycine (G, Gly); histidine (H, His); isoleucine (I, He); leucine (L, Leu); lysine (K, Lys); methionine (M, Met); phenylalanine (F, Phe); proline (P, Pro); serine (S, Ser); threonine (T, Thr); tryptophan (W, Trp); tyrosine (Y, Tyr); valine (V, Vai). Unless otherwise specified, X can indicate any amino acid. In some aspects, X can be asparagine (N), glutamine (Q), histidine (H), lysine (K), or arginine (R). References to these amino acids are also in the form of “[amino acid] [residues / residues]” (e.g., lysine residue, lysine residues, leucine residue, leucine residues, etc.).

[0115] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure belongs. Although methods and materials similar or equivalent to those described herein can be used in the practice or testing of the present disclosure, suitable methods and materials are described below.

[0116] “Amino” refers to the -NH2 radical.

[0117] “Cyano” refers to the -CN radical.

[0118] “Nitro” refers to the -NO2 radical.

[0119] “ Oxo” refers to the =0 radical.

[0120] “Hydroxyl” refers to the -OH radical.

[0121] “Alkyl” generally refers to an acyclic (e.g., straight or branched) or cyclic hydrocarbon (e.g., chain) radical consisting of carbon and hydrogen atoms, such as having from one to fifteen carbon atoms (e.g., C1-C15 alkyl). Unless otherwise state, alkyl is saturated or unsaturated e.g., an alkenyl, which comprises at least one carbon-carbon double bond; or an alkynyl, which comprises at least one carbon-carbon triple bond). Disclosures provided herein of an “alkyl” are intended to include independent recitations of a saturated “alkyl,” unless otherwise stated. Alkyl groups described herein are generally monovalent, but may also be divalent (which may also be described herein as “alkylene,” “alkylenyl,” “alkenylene,” “alkenylenyl,” “alkynylene,” “alkynylenyl,” groups). In certain embodiments, an alkyl comprises one to thirteen carbon atoms (e.g., C1-C13 alkyl). In certain embodiments, an alkyl comprises one to eight carbon atoms (e.g., Ci-Cs alkyl). In other embodiments, an alkyl comprises one to five carbon atoms (e.g., C1-C5 alkyl). In other embodiments, an alkyl comprises one to four carbon atoms (e.g., C1-C4 alkyl). In other embodiments, an alkyl comprises one to three carbon atoms e.g., C1-C3 alkyl). In other embodiments, an alkyl comprises one to two carbon atoms e.g., C1-C2 alkyl). In other embodiments, an alkyl comprises one carbon atom e.g., Ci alkyl). In other embodiments, an alkyl comprises five to fifteen carbon atoms e.g., C5-C15 alkyl). In other embodiments, an alkyl comprises five to eight carbon atoms e.g., Cs-Cs alkyl). In other embodiments, an alkyl comprises two to five carbon atoms e.g., C2-C5 alkyl). In other embodiments, an alkyl comprises three to five carbon atoms e.g., C3-C5 alkyl). In other embodiments, the alkyl group is selected from methyl, ethyl, 1-propyl (n-propyl), 1-methylethyl (iso-propyl), 1-butyl (n-butyl), 1 -methylpropyl (sec-butyl), 2- methylpropyl (iso-butyl), 1,1 -dimethylethyl (tert-butyl), 1 -pentyl (n-pentyl). The alkyl is attached to the rest of the molecule by a single, double, or triple bond. In general, alkyl groups are each independently substituted or unsubstituted. Each recitation of “alkyl” provided herein, unless otherwise stated, includes a specific and explicit recitation of an unsaturated “alkyl” group. Similarly, unless stated otherwise specifically in the specification, an alkyl group is optionally substituted by one or more of the following substituents: halo, cyano, nitro, oxo, thioxo, imino, oximo, trimethylsilanyl, -ORa, -SIU, -OC(O)-Ra, -N(Ra)2, -C(O)Ra, -C(O)ORa, - C(O)N(Ra)2, -N(Ra)C(O)ORa, -OC(O)-N(Ra)2, -N(Ra)C(O)Ra, -N(Ra)S(O)tRa(where t is 1 or 2), -S(O)tORa(where t is 1 or 2), -S(O)tRa(where t is 1 or 2) and -S(O)tN(Ra)2 (where t is 1 or 2)where each Rais independently hydrogen, alkyl (optionally substituted with halogen, hydroxy, methoxy, or trifluoromethyl), fluoroalkyl, carbocyclyl (optionally substituted with halogen, hydroxy, methoxy, or trifluoromethyl), carbocyclylalkyl (optionally substituted with halogen, hydroxy, methoxy, or trifluoromethyl), aryl (optionally substituted with halogen, hydroxy, methoxy, or trifluoromethyl), aralkyl (optionally substituted with halogen, hydroxy, methoxy, or trifluoromethyl), heterocyclyl (optionally substituted with halogen, hydroxy, methoxy, or trifluoromethyl), heterocyclylalkyl (optionally substituted with halogen, hydroxy, methoxy, or trifluoromethyl), heteroaryl (optionally substituted with halogen, hydroxy, methoxy, or trifluoromethyl), or heteroarylalkyl (optionally substituted with halogen, hydroxy, methoxy, or trifluoromethyl).

[0122] “Alkylene” or “alkylene chain” generally refers to a straight or branched divalent alkyl group linking the rest of the molecule to a radical group, such as having from one to twelve carbon atoms, for example, methylene, ethylene, propylene, / -propylene, ^-butylene, and the like. In some embodiments, alkylene may encompass unsaturated divalent hydrocarbons, such as alkenylene (e.g., -CH=CH-) or alkynylene (e.g., -C=C-) groups. Unless stated otherwise specifically in the specification, an alkylene chain is optionally substituted as described for alkyl groups herein.

[0123] “Alkenyl” refers to a straight or branched hydrocarbon chain radical group consisting of carbon and hydrogen atoms, containing at least one carbon-carbon double bond, and having from two to twelve carbon atoms. In certain embodiments, an alkenyl comprises two to eight carbon atoms. In other embodiments, an alkenyl comprises two to four carbon atoms. The alkenyl is optionally substituted as described for “alkyl” groups. In some embodiments, two or more carbon-carbon double bonds may form a conjugated system (e.g., alternating single (o) and double (n) bonds such as -CH=CH-CH=CH-). In some embodiments, the alkenyl is optionally substituted with an alkynyl group. In some embodiments, a carbon-carbon single bond (o), a carbon-carbon double bond (TC), and a carbon-carbon triple bond (K), may form a conjugated system (e.g., -CH=CH-C=C-).

[0124] “Alkenylene” or “alkenylene chain” generally refers to a straight or branched divalent alkenyl group linking the rest of the molecule to a radical group, such as having from one to twelve carbon atoms, for example, ethenylene (-CH=CH-), propenylene (-CH2CH=CH-), butenylene (-CH=CHCH2CH2-, or -CH2CH=CHCH2-), and the like. Unless stated otherwise specifically in the specification, an alkenylene chain is optionally substituted as described for alkyl groups herein.

[0125] “Alkynyl” refers to a straight or branched hydrocarbon chain radical group consisting of carbon and hydrogen atoms, containing at least one carbon-carbon triple bond, and having from two to twelve carbon atoms. In certain embodiments, an alkynyl comprises two to eight carbon atoms. In other embodiments, an alkynyl comprises two to four carbon atoms. The alkynyl is optionally substituted as described for “alkyl” groups. In some embodiments, two or more carbon-carbon triple bonds may form a conjugated system (e.g., alternating single (o) and double (71) bonds such as -C=C-C=C-). In some embodiments, an alkynyl is optionally substituted with an alkenyl group. In some embodiments, a carbon-carbon single bond (o), a carbon-carbon double bond (71), and a carbon-carbon triple bond (71), may form a conjugated system (e g., -CH=CH-C=C-).

[0126] “Alkynylene” or “alkynylene chain” generally refers to a straight or branched divalent alkynyl group linking the rest of the molecule to a radical group, such as having from one to twelve carbon atoms, for example, ethynylene (-C=C-), propynylene (-CH2OC-), butynylene (-OCCH2CH2-, or -CtLC^CCHz-), and the like. Unless stated otherwise specifically in the specification, an alkynylene chain is optionally substituted as described for alkyl groups herein.

[0127] “Alkoxy” refers to a radical bonded through an oxygen atom of the formula -O- alkyl, where alkyl is an alkyl chain as defined above.

[0128] “Aryl” refers to a radical derived from an aromatic monocyclic or multicyclic hydrocarbon ring system by removing a hydrogen atom from a ring carbon atom. The aromatic monocyclic or multicyclic hydrocarbon ring system contains hydrogen and carbon from five to eighteen carbon atoms, where at least one of the rings in the ring system is fully unsaturated, ie., it contains a cyclic, delocalized (4n+2) 71-electron system in accordance with the Hiickel theory. The ring system from which aryl groups are derived include, but are not limited to, groups such as benzene, fluorene, indane, indene, tetralin and naphthalene. Unless stated otherwise specifically in the specification, the term "aryl" or the prefix "ar-" (such as in "aralkyl") is meant to include aryl radicals optionally substituted by one or more substituents independently selected from alkyl, alkenyl, alkynyl, halo, fluoroalkyl, cyano, nitro, optionally substituted aryl, optionally substituted aralkyl, optionally substituted aralkenyl, optionally substituted aralkynyl, optionally substituted carbocyclyl, optionally substituted carbocyclylalkyl, optionally substituted heterocyclyl, optionally substituted heterocyclylalkyl, optionally substituted heteroaryl, optionally substituted heteroarylalkyl, -Rb-ORa, -Rb-OC(O)-Ra, -Rb-OC(O)-ORa, -Rb-OC(O)- N(Ra)2, -Rb-N(Ra)2, -Rb-C(O)Ra, -Rb-C(O)ORa, -Rb-C(O)N(Ra)2, -Rb-O-Rc-C(O)N(Ra)2, -Rb- N(Ra)C(O)ORa, -Rb-N(Ra)C(O)Ra, -Rb-N(Ra)S(O)tRa(where t is 1 or 2), -Rb-S(O)tRa(where t is1 or 2), -Rb-S(O)tORa(where t is 1 or 2) and -Rb-S(O)tN(Ra)2 (where t is 1 or 2), where each Rais independently hydrogen, alkyl (optionally substituted with halogen, hydroxy, methoxy, or trifluoromethyl), fluoroalkyl, cycloalkyl (optionally substituted with halogen, hydroxy, methoxy, or trifluoromethyl), cycloalkylalkyl (optionally substituted with halogen, hydroxy, methoxy, or trifluoromethyl), aryl (optionally substituted with halogen, hydroxy, methoxy, or trifluoromethyl), aralkyl (optionally substituted with halogen, hydroxy, methoxy, or trifluoromethyl), heterocyclyl (optionally substituted with halogen, hydroxy, methoxy, or trifluoromethyl), heterocyclylalkyl (optionally substituted with halogen, hydroxy, methoxy, or trifluoromethyl), heteroaryl (optionally substituted with halogen, hydroxy, methoxy, or trifluoromethyl), or heteroarylalkyl (optionally substituted with halogen, hydroxy, methoxy, or trifluoromethyl), each Rbis independently a direct bond, or a straight or branched alkylene, alkenylene, or alkynylene chain, and Rc is a straight or branched alkylene, alkenylene, or alkynylene chain, and where each of the above substituents is unsubstituted unless otherwise indicated.

[0129] “Aralkyl” or “aryl-alkyl” refers to a radical of the formula -Rc-aryl where Rcis an alkylene chain, an alkenylene chain, or an alkynylene chain as defined above, for example, methylene, ethylene, butylene, ethenylene, butenylene, ethynylene, butynylene, and the like. The alkylene chain part of the aralkyl radical is optionally substituted as described above for an alkylene chain. The alkenylene chain part of the aralkyl radical is optionally substituted as described above for an alkenylene chain. The alkynylene chain part of the aralkyl radical is optionally substituted as described above for an alkynylene chain. The aryl part of the aralkyl radical is optionally substituted as described above for an aryl group.

[0130] “Carbocyclyl” or “cycloalkyl” refers to a stable non-aromatic monocyclic or polycyclic hydrocarbon radical consisting of carbon and hydrogen atoms, which includes fused, bridged, or spiro ring systems, having from three to fifteen carbon atoms. In certain embodiments, a carbocyclyl comprises three to ten carbon atoms. In other embodiments, a carbocyclyl comprises five to seven carbon atoms. The carbocyclyl is attached to the rest of the molecule by a single bond, a double bond, or a triple bond. Carbocyclyl or cycloalkyl is saturated (z.e., containing single C-C bonds) or unsaturated (z.e., containing one or more double bonds or triple bonds). Examples of saturated cycloalkyls include, e.g., cyclopropyl, cyclobutyl, cyclopentyl, cyclohexyl, cycloheptyl, and cyclooctyl. An unsaturated carbocyclyl is also referred to as “cycloalkenyl” or “cycloalkynyl.” Examples of monocyclic cycloalkenyls include, e.g., cyclopentenyl, cyclohexenyl, cycloheptenyl, and cyclooctenyl. Examples of monocyclic cycloalkynyls include, e.g., cyclooctynyl, cyclononynyl, and cyclodecynyl. Polycycliccarbocyclyl radicals include, for example, adamantyl, norbornyl (z.e., bicyclo[2.2.1]heptanyl), norbomenyl, decalinyl, 7,7-dimethyl-bicyclo[2.2.1]heptanyl, and the like. Unless otherwise stated specifically in the specification, the term “carbocyclyl” is meant to include carbocyclyl radicals that are optionally substituted by one or more substituents independently selected from alkyl, alkenyl, alkynyl, halo, fluoroalkyl, oxo, thioxo, cyano, nitro, optionally substituted aryl, optionally substituted aralkyl, optionally substituted aralkenyl, optionally substituted aralkynyl, optionally substituted carbocyclyl, optionally substituted carbocyclylalkyl, optionally substituted heterocyclyl, optionally substituted heterocyclylalkyl, optionally substituted heteroaryl, optionally substituted heteroarylalkyl, -Rb-0Ra, -Rb-OC(O)-Ra, -Rb-OC(O)-ORa, -Rb-OC(O)- N(Ra)2, -Rb-N(Ra)2, -Rb-C(O)Ra, -Rb-C(O)ORa, -Rb-C(O)N(Ra)2, -Rb-O-Rc-C(O)N(Ra)2, -Rb- N(Ra)C(O)ORa, -Rb-N(Ra)C(O)Ra, -Rb-N(Ra)S(O)tRa(where t is 1 or 2), -Rb-S(O)tRa(where t is 1 or 2), -Rb-S(O)tORa(where t is 1 or 2) and -Rb-S(O)tN(Ra)2 (where t is 1 or 2), where each Rais independently hydrogen, alkyl (optionally substituted with halogen, hydroxy, methoxy, or trifluoromethyl), fluoroalkyl, cycloalkyl (optionally substituted with halogen, hydroxy, methoxy, or trifluoromethyl), cycloalkylalkyl (optionally substituted with halogen, hydroxy, methoxy, or trifluoromethyl), aryl (optionally substituted with halogen, hydroxy, methoxy, or trifluoromethyl), aralkyl (optionally substituted with halogen, hydroxy, methoxy, or trifluoromethyl), heterocyclyl (optionally substituted with halogen, hydroxy, methoxy, or trifluoromethyl), heterocyclylalkyl (optionally substituted with halogen, hydroxy, methoxy, or trifluoromethyl), heteroaryl (optionally substituted with halogen, hydroxy, methoxy, or trifluoromethyl), or heteroarylalkyl (optionally substituted with halogen, hydroxy, methoxy, or trifluoromethyl), each Rbis independently a direct bond, or a straight or branched alkylene, alkenylene, or alkynylene chain, and Rcis a straight or branched alkylene, alkenylene, or alkenylene chain, and where each of the above substituents is unsubstituted unless otherwise indicated.

[0131] “Carbocyclylalkyl” refers to a radical of the formula -Rc-carbocyclyl where Rcis an alkylene chain, an alkenylene chain, or an alkynylene chain as defined above. The alkylene chain, the alkenylene chain, the alkynylene chain, or the carbocyclyl radical is optionally substituted as defined above.

[0132] “Carbocyclylalkenyl” refers to a radical of the formula -Rc-carbocyclyl where Rcis an alkenylene chain as defined above. The alkenylene chain and the carbocyclyl radical is optionally substituted as defined above.

[0133] “Carbocyclylalkynyl” refers to a radical of the formula -Rc-carbocyclyl where Rcis an alkynylene chain as defined above. The alkynylene chain and the carbocyclyl radical is optionally substituted as defined above.

[0134] “Carbocyclylalkoxy” refers to a radical bonded through an oxygen atom of the formula -O-Rc-carbocyclyl where Rcis an alkylene chain, an alkenylene chain, or an alkynylene chain as defined above. The alkylene chain, the alkenylene chain, the alkynylene chain and the carbocyclyl radical is optionally substituted as defined above.

[0135] “Halo" or “halogen” refers to fluoro, bromo, chloro, or iodo substituents.

[0136] “Haloalkyl” refers to an alkyl radical, as defined above, that is substituted by one or more halogen radicals, as defined above, for example, trihalomethyl, dihalomethyl, halomethyl, and the like. In some embodiments, the haloalkyl is a fluoroalkyl, such as, for example, trifluoromethyl, difluoromethyl, fluoromethyl, 2,2,2-trifluoroethyl, l-fluoromethyl-2-fluoroethyl, and the like. In some embodiments, the alkyl part of the fluoroalkyl radical is optionally substituted as defined above for an alkyl group.

[0137] “Heteroalkyl” refers to an alkyl group as defined above in which one or more skeletal carbon atoms of the alkyl are substituted with one to six heteroatoms selected from boron, nitrogen, oxygen, silicon, phosphorus, and sulfur (with the appropriate number of substituents or valencies - for example, -CH2- may be replaced with -NH- or -O-). For example, each substituted carbon atom is independently substituted with a heteroatom, such as wherein the carbon is substituted with a nitrogen, oxygen, sulfur, or other suitable heteroatom. In some instances, each substituted carbon atom is independently substituted for an oxygen, nitrogen (e.g. -NH-, -N(alkyl)-, or -N(aryl)- or having another substituent contemplated herein), or sulfur (e.g. -S-, -S(=O)-, or -S(=O)2-). In some embodiments, a heteroalkyl is attached to the rest of the molecule at a carbon atom of the heteroalkyl. In some embodiments, a heteroalkyl is attached to the rest of the molecule at a heteroatom of the heteroalkyl. In some embodiments, a heteroalkyl is a C1-C18 heteroalkyl. In some embodiments, a heteroalkyl is a C1-C12 heteroalkyl. In some embodiments, a heteroalkyl is a Ci-Ce heteroalkyl. In some embodiments, a heteroalkyl is a Ci- C4 heteroalkyl. In some embodiments, heteroalkyl includes alkylamino, alkylaminoalkyl, aminoalkyl, heterocycloalkyl, heterocyclyl, and heterocycloalkylalkyl, as defined herein. Unless stated otherwise specifically in the specification, a heteroalkyl group is optionally substituted as defined above for an alkyl group.

[0138] “Heteroalkylene” refers to a divalent heteroalkyl group defined above which links one part of the molecule to another part of the molecule. Unless stated specifically otherwise, a heteroalkylene is optionally substituted, as defined above for an alkyl group.

[0139] “Heterocyclyl” or “heterocycloalkyl” refers to a stable 3- to 18-membered nonaromatic ring radical that comprises two to twelve carbon atoms and from one to six heteroatoms selected from boron, nitrogen, oxygen, silicon, phosphorus, and sulfur (with the appropriate number of substituents or valencies - for example, -CH2- may be replaced with - NH- or -O-). Unless stated otherwise specifically in the specification, the heterocyclyl radical is a monocyclic, bicyclic, tricyclic or tetracyclic ring system, which optionally includes fused, bridged, or spiro ring systems. The heteroatoms in the heterocyclyl radical are optionally oxidized. One or more nitrogen atoms, if present, are optionally quatemized (z.e., a nitrogen atom forms four covalent bonds, resulting in a positively charged quaternary ammonium species N+). The heterocyclyl radical is partially or fully saturated. The heterocyclyl radical is saturated (z.e., containing single C-C bonds) or unsaturated (e.g., containing one or more double bonds or triple bonds in the ring system). In some instances, the heterocyclyl radical is saturated. In some instances, the heterocyclyl radical is saturated and substituted. In some instances, the heterocyclyl radical is unsaturated. Examples of such heterocyclyl radicals include, but are not limited to, dioxolanyl, thienyl[l,3]dithianyl, decahydroisoquinolyl, imidazolinyl, imidazolidinyl, isothiazolidinyl, isoxazolidinyl, morpholinyl, octahydroindolyl, octahydroisoindolyl, 2-oxopiperazinyl, 2-oxopiperidinyl, 2-oxopyrrolidinyl, oxazolidinyl, piperidinyl, piperazinyl, 4-piperidonyl, pyrrolidinyl, pyrazolidinyl, quinuclidinyl, thiazolidinyl, tetrahydrofuryl, trithianyl, tetrahydropyranyl, thiomorpholinyl, thiamorpholinyl, 1-oxo-thiomorpholinyl, and 1,1-dioxo-thiomorpholinyl. Unless stated otherwise specifically in the specification, the term “heterocyclyl” is meant to include heterocyclyl radicals as defined above that are optionally substituted by one or more substituents selected from alkyl, alkenyl, alkynyl, halo, fluoroalkyl, oxo, thioxo, cyano, nitro, optionally substituted aryl, optionally substituted aralkyl, optionally substituted aralkenyl, optionally substituted aralkynyl, optionally substituted carbocyclyl, optionally substituted carbocyclylalkyl, optionally substituted heterocyclyl, optionally substituted heterocyclylalkyl, optionally substituted heteroaryl, optionally substituted heteroarylalkyl, -Rb-ORa, -Rb-OC(O)-Ra, -Rb-OC(O)-ORa, -Rb-OC(O)-N(Ra)2, -Rb-N(Ra)2, -Rb- C(O)Ra, -Rb-C(O)ORa, -Rb-C(O)N(Ra)2, -Rb-O-Rc-C(O)N(Ra)2, -Rb-N(Ra)C(O)ORa, -Rb- N(Ra)C(O)Ra, -Rb-N(Ra)S(O)tRa(where t is 1 or 2), -Rb-S(O)tRa(where t is 1 or 2), -Rb- S(O)tORa(where t is 1 or 2) and -Rb-S(O)tN(Ra)2 (where t is 1 or 2), where each Rais independently hydrogen, alkyl (optionally substituted with halogen, hydroxy, methoxy, or trifluoromethyl), fluoroalkyl, cycloalkyl (optionally substituted with halogen, hydroxy, methoxy, or trifluoromethyl), cycloalkylalkyl (optionally substituted with halogen, hydroxy, methoxy, or trifluoromethyl), aryl (optionally substituted with halogen, hydroxy, methoxy, ortrifluoromethyl), aralkyl (optionally substituted with halogen, hydroxy, methoxy, or trifluoromethyl), heterocyclyl (optionally substituted with halogen, hydroxy, methoxy, or trifluoromethyl), heterocyclylalkyl (optionally substituted with halogen, hydroxy, methoxy, or trifluoromethyl), heteroaryl (optionally substituted with halogen, hydroxy, methoxy, or trifluoromethyl), or heteroarylalkyl (optionally substituted with halogen, hydroxy, methoxy, or trifluoromethyl), each Rbis independently a direct bond, or a straight or branched alkylene, alkenylene, or alkynylene chain, and Rcis a straight or branched alkylene, alkenylene, or alkynylene chain, and where each of the above substituents is unsubstituted unless otherwise indicated.

[0140] ‘W-heterocyclyl” or “N-attached heterocyclyl” refers to a heterocyclyl radical as defined above containing at least one nitrogen and where the point of attachment of the heterocyclyl radical to the rest of the molecule is through a nitrogen atom in the heterocyclyl radical. An / f-heterocyclyl radical is optionally substituted as described above for heterocyclyl radicals. Examples of such A-heterocyclyl radicals include, but are not limited to, 1- morpholinyl, 1-piperidinyl, 1-piperazinyl, 1-pyrrolidinyl, pyrazolidinyl, imidazolinyl, and imidazolidinyl.

[0141] “C-heterocyclyl” or “C-attached heterocyclyl” refers to a heterocyclyl radical as defined above containing at least one heteroatom and where the point of attachment of the heterocyclyl radical to the rest of the molecule is through a carbon atom in the heterocyclyl radical. A C-heterocyclyl radical is optionally substituted as described above for heterocyclyl radicals. Examples of such C-heterocyclyl radicals include, but are not limited to, 2- morpholinyl, 2- or 3- or 4-piperidinyl, 2-piperazinyl, 2- or 3-pyrrolidinyl, and the like.

[0142] “Heterocyclylalkyl” refers to a radical of the formula -Rc-heterocyclyl where Rcis an alkylene, alkenylene, or alkynylene chain as defined above. If the heterocyclyl is a nitrogen-containing heterocyclyl, the heterocyclyl is optionally attached to the alkyl radical at the nitrogen atom. The alkylene, the alkenylene, or the alkynylene chain of the heterocyclylalkyl radical is optionally substituted as defined above. The heterocyclyl part of the heterocyclylalkyl radical is optionally substituted as defined above for a heterocyclyl group.

[0143] “Heterocyclylalkoxy” refers to a radical bonded through an oxygen atom of the formula -O-Rc-heterocyclyl where Rcis an alkylene, an alkenylene chain, or an alkynylene chain as defined above. If the heterocyclyl is a nitrogen-containing heterocyclyl, the heterocyclyl is optionally attached to the alkyl radical at the nitrogen atom. The alkylene chain of the heterocyclylalkoxy radical is optionally substituted as defined above for an alkylenechain. The heterocyclyl part of the heterocyclylalkoxy radical is optionally substituted as defined above for a heterocyclyl group.

[0144] “Heteroaryl” refers to a radical derived from a 3- to 18-membered aromatic ring radical that comprises two to seventeen carbon atoms and from one to six heteroatoms selected from nitrogen, oxygen and sulfur (with the appropriate number of substituents or valencies - for example, -CH2- may be replaced with -NH- or -O-). As used herein, the heteroaryl radical is a monocyclic, bicyclic, tricyclic or tetracyclic ring system, wherein at least one of the rings in the ring system is fully unsaturated, z.e., it contains a cyclic, delocalized (4n+2) 71-electron system in accordance with the Hiickel theory. Heteroaryl includes fused, bridged, or spiro ring systems. The heteroatom(s) in the heteroaryl radical is optionally oxidized. One or more nitrogen atoms, if present, are optionally quatemized (z.e., a nitrogen atom forms four covalent bonds, resulting in a positively charged quaternary ammonium species N+). The heteroaryl is attached to the rest of the molecule through any atom of the ring(s). Examples of heteroaryls include, but are not limited to, azepinyl, acridinyl, benzimidazolyl, benzindolyl, 1,3-benzodioxolyl, benzofuranyl, benzooxazolyl, benzo[d]thiazolyl, benzothiadiazolyl, benzo[b][l,4]dioxepinyl, benzo [b][ 1,4] oxazinyl, 1,4-benzodioxanyl, benzonaphthofuranyl, benzoxazolyl, benzodioxolyl, benzodioxinyl, benzopyranyl, benzopyranonyl, benzofuranyl, benzofuranonyl, benzothienyl (benzothiophenyl), benzothieno[3,2-d]pyrimidinyl, benzotriazolyl, benzo[4,6]imidazo[l,2-a]pyridinyl, carbazolyl, cinnolinyl, cyclopenta[d]pyrimidinyl,6.7-dihydro-5H-cyclopenta[4,5]thieno[2,3-d]pyrimidinyl, 5,6-dihydrobenzo[h]quinazolinyl, 5,6-dihydrobenzo[h]cinnolinyl, 6,7-dihydro-5H-benzo[6,7]cyclohepta[l,2-c]pyridazinyl, dibenzofuranyl, dibenzothiophenyl, furanyl, furanonyl, furo[3,2-c]pyridinyl,5.6.7.8.9.10-hexahydrocycloocta[d]pyrimidinyl, 5, 6, 7, 8, 9, 10-hexahydrocycloocta[d]pyridazinyl,5.6.7.8.9.10-hexahydrocycloocta[d]pyridinyl, isothiazolyl, imidazolyl, indazolyl, indolyl, indazolyl, isoindolyl, indolinyl, isoindolinyl, isoquinolyl, indolizinyl, isoxazolyl,5.8-methano-5,6,7,8-tetrahydroquinazolinyl, naphthyridinyl, 1 ,6-naphthyridinonyl, oxadiazolyl, 2-oxoazepinyl, oxazolyl, oxiranyl, 5,6,6a,7,8,9,10,10a-octahydrobenzo[h]quinazolinyl,1 -phenyl- IH-pyrrolyl, phenazinyl, phenothiazinyl, phenoxazinyl, phthalazinyl, pteridinyl, purinyl, pyrrolyl, pyrazolyl, pyrazolo[3,4-d]pyrimidinyl, pyridinyl, pyrido[3,2-d]pyrimidinyl, pyrido[3,4-d]pyrimidinyl, pyrazinyl, pyrimidinyl, pyridazinyl, pyrrolyl, quinazolinyl, quinoxalinyl, quinolinyl, isoquinolinyl, tetrahydroquinolinyl, 5,6,7,8-tetrahydroquinazolinyl,5.6.7.8-tetrahydrobenzo[4,5]thieno[2,3-d]pyrimidinyl,6.7.8.9-tetrahydro-5H-cyclohepta[4,5]thieno[2,3-d]pyrimidinyl, 5,6,7,8-tetrahydropyrido[4,5-c]pyridazinyl, thiazolyl, thiadiazolyl, triazolyl, tetrazolyl, triazinyl,thieno[2,3-d]pyrimidinyl, thieno[3,2-d]pyrimidinyl, thieno[2,3-c]pridinyl, and thiophenyl (j.e. thienyl). Unless stated otherwise specifically in the specification, the term "heteroaryl" is meant to include heteroaryl radicals as defined above which are optionally substituted by one or more substituents selected from alkyl, alkenyl, alkynyl, halo, fluoroalkyl, haloalkenyl, haloalkynyl, oxo, thioxo, cyano, nitro, optionally substituted aryl, optionally substituted aralkyl, optionally substituted aralkenyl, optionally substituted aralkynyl, optionally substituted carbocyclyl, optionally substituted carbocyclylalkyl, optionally substituted heterocyclyl, optionally substituted heterocyclylalkyl, optionally substituted heteroaryl, optionally substituted heteroarylalkyl, -Rb-0Ra, -Rb-0C(0)-Ra, -Rb-0C(0)-0Ra, -Rb-0C(0)-N(Ra)2, -Rb-N(Ra)2, -Rb- C(0)Ra, -Rb-C(0)0Ra, -Rb-C(0)N(Ra)2, -Rb-0-Rc-C(0)N(Ra)2, -Rb-N(Ra)C(0)0Ra, -Rb- N(Ra)C(0)Ra, -Rb-N(Ra)S(0)tRa(where t is 1 or 2), -Rb-S(O)tRa(where t is 1 or 2), -Rb- S(O)tORa(where t is 1 or 2) and -Rb-S(0)tN(Ra)2 (where t is 1 or 2), where each Rais independently hydrogen, alkyl (optionally substituted with halogen, hydroxy, methoxy, or trifluoromethyl), fluoroalkyl, cycloalkyl (optionally substituted with halogen, hydroxy, methoxy, or trifluoromethyl), cycloalkylalkyl (optionally substituted with halogen, hydroxy, methoxy, or trifluoromethyl), aryl (optionally substituted with halogen, hydroxy, methoxy, or trifluoromethyl), aralkyl (optionally substituted with halogen, hydroxy, methoxy, or trifluoromethyl), heterocyclyl (optionally substituted with halogen, hydroxy, methoxy, or trifluoromethyl), heterocyclylalkyl (optionally substituted with halogen, hydroxy, methoxy, or trifluoromethyl), heteroaryl (optionally substituted with halogen, hydroxy, methoxy, or trifluoromethyl), or heteroarylalkyl (optionally substituted with halogen, hydroxy, methoxy, or trifluoromethyl), each Rbis independently a direct bond, or a straight or branched alkylene, alkenylene or alkynylene chain, and Rcis a straight or branched alkylene, alkenylene or alkynylene chain, and where each of the above substituents is unsubstituted unless otherwise indicated.

[0145] ‘W-heteroaryl” refers to a heteroaryl radical as defined above containing at least one nitrogen and where the point of attachment of the heteroaryl radical to the rest of the molecule is through a nitrogen atom in the heteroaryl radical. An A-heteroaryl radical is optionally substituted as described above for heteroaryl radicals.

[0146] “ C-heteroaryl” refers to a heteroaryl radical as defined above and where the point of attachment of the heteroaryl radical to the rest of the molecule is through a carbon atom in the heteroaryl radical. A C-heteroaryl radical is optionally substituted as described above for heteroaryl radicals.

[0147] “Heteroarylalkyl” refers to a radical of the formula -Rc-heteroaryl, where Rcis an alkylene chain, an alkenylene chain, or an alkynylene chain as defined above. If the heteroaryl is a nitrogen-containing heteroaryl, the heteroaryl is optionally attached to the alkyl radical at the nitrogen atom. The alkylene chain, the alkenylene chain, or the alkynylene chain of the heteroarylalkyl radical is optionally substituted as defined above. The heteroaryl part of the heteroarylalkyl radical is optionally substituted as defined above for a heteroaryl group.

[0148] “Heteroarylalkoxy” refers to a radical bonded through an oxygen atom of the formula -O-Rc-heteroaryl, where Rcis an alkylene chain, an alkenylene chain, or an alkynylene chain as defined above. If the heteroaryl is a nitrogen-containing heteroaryl, the heteroaryl is optionally attached to the alkyl radical at the nitrogen atom. The alkylene chain, the alkenylene chain, or the alkynylene chain of the heteroarylalkoxy radical is optionally substituted as defined above. The heteroaryl part of the heteroarylalkoxy radical is optionally substituted as defined above for a heteroaryl group.

[0149] “Leaving group” or “LG” as used herein may refer to an atom or a group that departs with a pair of electrons in heterolytic bond cleavage. In some instances, a leaving group is capable of being displaced by a nucleophile. Such a leaving group may include, but is not limited to halogen such as chloro, bromo, fluoro, and iodo; alkyl; aryl; aralkyl; carbocyclyl; heteroalkyl; heterocyclyl; and the like. Additional examples may be found in Smith, March’s Advanced Organic Chemistry, 8thEdition (2020), which is incorporated herein by reference. In some instances, a leaving group may stabilize the negative charge upon departure, such leaving group may include, but is not limited to halogen, tosylate (OTs), mesylate (OMs), triflate (OTf), nonaflate (ONf), besylate (OBs), sulfate (-SO4), and phosphate (-PO4). In some instances, a leaving group may dissociate after undergoing chemical modification (e.g., transformation of - OH into a more favorable leaving group -OTs) or under specific reaction conditions (e.g., protonation in acidic conditions), such leaving group may include, but is not limited to, hydroxyl (-OH), amine (-NH2), and alkoxide (-OR). In some instances, the leaving group is an intramolecular leaving group, which participates in a reaction where bond cleavage and new bond formation occur within the same molecule. Examples of such intramolecular leaving groups include phenylthiocarbamoyl (PTC) group elimination in Edman degradation, where the PTC-peptide intermediate undergoes intramolecular cyclization and subsequent cleavage to release the N-terminal amino acid as an anilinothiazolinone (ATZ) derivative.

[0150] “Protecting group” or “PG” refers to a group that can temporarily block a particular functional moiety, e.g., O, S, or N, so that a reaction can be carried out selectively at another reactive site in a multifunctional compound. A protecting group can be added and removed atspecific stages during the synthesis of a compound or when the compound participates in a reaction. Exemplary protecting groups are described in detail in Theodora W. Greene et. al. Protecting Groups in Organic Synthesis, 4thedition (2007), which is herein incorporated by reference in its entirety. Protection of functional groups of a compound may alter other physical properties besides the reactivity of the protected functional group, such as the polarity, solubility, lipophilicity (hydrophobicity), and other properties.

[0151] “Bulky group” or “BG” refers to a substituent or functional group that occupy spatial volume around a molecular center, often impeding steric access to reactive sites. In some embodiments, the bulky group can influence molecular conformations, reaction pathways, and / or steric hindrance effects in chemical reactions. Examples may include, but are not limited to, tert-butyl (-C(CH3)3), triphenylmethyl (- CeHs^), isopropyl (-CH(CH3)2), silyl protecting groups (e.g., trimethylsilyl, -Si(CH3)3), etc. In some embodiments, a bulky group comprises a methyl group. In systems with constrained geometries (e.g., bridgehead positions, cyclic structures, or sterically crowded environments), the presence of a methyl group can significantly influence molecular conformation and restrict access to reactive sites.

[0152] “Electron-withdrawing group” or “EWG” as used herein is a moiety, e.g., an atom or group, which draws electron density from the neighboring atoms towards itself, usually by resonance or inductive effects. Non-limiting examples of EWG are halogen, haloalkyl, -C(O)R', -COOR', -C(O)NH2, -NHC(O)R', -C(O)NR'R", carbonyl, -CF3, -CN, sulfonate, amino, alkylamino, -SO3H, -SO2CF3, -SO2R', SCENR'R", oxo (=0), alkyl ammonium, and -NO2, wherein R' and R" are independently H, alkyl, heteroalkyl, aryl, heteroaryl, each optionally substituted with one or more EWGs.

[0153] “Electron-donating group” or “EDG” as used herein is a moiety, e.g., an atom or group, which releases electron density to the neighboring atoms from itself, usually by resonance or inductive effects. Non-limiting examples of EDG are -NRN1RN2, -OR', -NHC(O)R', -OC(O)R', phenyl, and vinyl, wherein RN1, RN2, and R' are independently H, alkyl, heteroalkyl, aryl, heteroaryl, each optionally substituted with one or more EDGs.

[0154] “ Solvent” refers to a substance that dissolves one or more solutes, resulting in a solution. A solvent may serve as a medium for any reaction or transformation described herein. The solvent may dissolve one or more reactants or reagents in a reaction mixture. The solvent may facilitate the mixing of one or more reagents or reactants in a reaction mixture. The solvent may also serve to increase or decrease the rate of a reaction relative to the reaction in a different solvent. Solvents can be organic solvents and aqueous (water-based) solvents. Organic solvents comprise carbon-based compounds that dissolve organic and inorganic substances and arecommonly used in chemical synthesis, extraction, and purification processes. Aqueous solvents comprise water-miscible organic solvents or buffers to adjust pH, ionic strength, or solubility properties. Solvents can be polar or non-polar, protic or aprotic. Common solvents useful in the methods described herein include, but are not limited to, acetone, acetonitrile, benzene, benzonitrile, 1 -butanol, 2-butanone, butyl acetate, tert-butyl methyl ether, carbon disulfide carbon tetrachloride, chlorobenzene, 1 -chlorobutane, chloroform, cyclohexane, cyclopentane, 1,2-di chlorobenzene, 1,2-di chloroethane, di chloromethane (DCM), N,N-dimethylacetamide N,N-dimethylformamide (DMF), l,3-dimethyl-3,4,5,6-tetrahydro-2-pyrimidinone (DMPU), 1,4- di oxane, 1,3 -di oxane, di ethylether, 2-ethoxy ethyl ether, ethyl acetate, ethyl alcohol, ethylene glycol, dimethyl ether, heptane, n-hexane, hexanes, hexamethylphosphoramide (HMPA), 2- methoxyethanol, 2-methoxyethyl acetate, methyl alcohol, 2-methylbutane, 4-methyl-2- pentanone, 2-methyl-l -propanol, 2-m ethyl -2 -propanol, 1 -methyl -2 -pyrrolidinone, dimethylsulfoxide (DMSO), nitromethane, 1-octanol, pentane, 3-pentanone, 1-propanol, 2- propanol, pyridine, tetrachloroethylene, tetrahydrofuran (THF), 2-methyltetrahydrofuran, toluene, tri chlorobenzene, 1, 1,2-tri chlorotrifluoroethane, 2,2,4-trimethylpentane, trimethylamine, triethylamine, N,N-diisopropylethylamine, diisopropylamine, water, o-xylene, and p-xylene.

[0155] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure belongs. Although methods and materials similar or equivalent to those described herein can be used in the practice or testing of the present disclosure, suitable methods and materials are described below.Sequencing Reagents

[0156] Provided herein are compounds, which may be used as sequencing reagents. The sequencing reagents provide a novel approach to sequencing polymeric analytes, such as peptides, wherein monomers, such as terminal amino acids, of the polymeric analytes are isolated using a sequencing reagent provided herein. See, e.g., FIGs. 1A-2D and 6. In some instances, isolation occurs by tethering of a monomer, such as a terminal amino acid, to a capture moiety (e.g., a capture moiety bound to a substrate, the polymeric analyte of interest, or in solution) using the sequencing reagent and cleavage of the monomer from the polymeric analyte (e.g., peptide). In some instances, sequencing comprises coupling the sequencing reagent to a monomer, thereby generating a sequencing reagent-monomer complex, and detecting the sequencing reagent-monomer complex or derivative thereof. The detection of the sequencingreagent-monomer complex or derivative thereof may be performed using specific binding agents to the sequencing reagent-monomer complexes or derivatives thereof, such as sequencing reagent-amino acid complexes. Alternatively, or in addition to, the detection may be performed using a nanopore device or instrument (e.g., nanopore sequencer). In some embodiments, the isolation of amino acids provided herein avoids a local environment problem in which adjacent amino acids can impact the detection of an amino acid, e.g., by altering the binding (e.g., selectivity or affinity) of a binding agent or detection of convoluted signals arising from multiple amino acids present in the detection area of a nanopore device or instrument.

[0157] Classical peptide sequencing may be completed using Edman degradation, in which the N-terminal amino acid of a peptide is sequentially removed. This is completed using a phenyl isothiocyanate reactive group (PITC), which upon reaction with the N-terminal amino acid, results in a phenylthiocarbamoyl intermediate which is cleaved, creating a cyclic compound. Sequential removal of amino acid residues allows for peptide sequencing without damage of the peptide or protein itself. However, there are a number of issues associated with Edman degradation and the use of PITC as a reactive group. These issues comprise the requirement of harsh reaction conditions (e.g., high heat and acid). These harsh reaction conditions are not amenable to the analysis or use of nucleic acid molecules (e.g., DNA used for barcoding or coupling of monomers to a capture moiety or substrate, described elsewhere herein and exemplified in FIGs. 1A-1D, 2A-2D). Edman degradation also can result in racemization of the cleaved N-terminal amino acid from the peptide, which can limit binding agent recognition and detection (e.g., if the binding agent only recognizes one stereoisomer) or detection and accurate readout of an amino acid using a nanopore device (e.g., nanopore sequencer). Further, in amino acids, such as lysine, where there are amine side chain functionalities, the PITC reactive group can react with the side chain, which decreases the number of amine groups for functionalization (e.g., labeling, tethering to a substrate, or other purpose).

[0158] Provided herein, in some embodiments, is a sequencing reagent comprising a reactive group. In some embodiments, the sequencing reagent comprises an additional reactive moiety. In some embodiments, the sequencing reagent comprises a linker coupled to the reactive group and the additional reactive moiety. In some instances, the sequencing reagents provided herein comprise a guanidinylating agent. The sequencing reagents provided herein can allow for sequential removal of terminal amino acid residues under milder reaction conditions than the high heat and acidic conditions required for Edman degradation. Such milder reaction conditions may be amenable to use of polymerizable molecules, such as nucleic acid molecules, that can be used for further analysis of the removed terminal amino acid residues. In some examples, thesequencing reagents provided herein provide high efficiency, such as faster kinetics than that of traditional Edman degradation. Beneficially, the sequencing reagents provided herein provide for alternative Edman degradation agents that increased reactivity with peptides and high cleavage efficiencies under relatively mild reaction conditions. The sequencing reagents provided herein may comprise electrophilic groups to allow for highly-efficient N-terminal amine conjugation whilst also comprising nucleophilic nitrogen moieties to improve cleavage efficiencies.

[0159] In some embodiments, provided herein is a compound of Formula (la) or Formula (lb):Formula (la) Formula (lb).

[0160] In some embodiments, the compound of Formula (la) or Formula (lb) is a sequencing reagent.

[0161] In some embodiments, LG is a leaving group. In some embodiments, LG is a leaving group described elsewhere herein.

[0162] In some embodiments, the leaving group is optionally substituted. In some embodiments, the leaving group is optionally substituted with one electron-withdrawing group (e.g., electron-withdrawing protecting group). In some embodiments, the leaving group is optionally substituted with one or more electron-withdrawing group (e.g., electron-withdrawing protecting group). In some embodiments, the leaving group is substituted with two electronwithdrawing groups (e.g., electron-withdrawing protecting groups). In some embodiments, the leaving group is substituted with three electron-withdrawing group. In some embodiments, the leaving group is substituted with four electron-withdrawing groups (e.g., electron-withdrawing protecting groups). In some embodiments, the leaving group is substituted with five electronwithdrawing groups (e.g., electron-withdrawing protecting groups). In some instances, an electron withdrawing group is an electron withdrawing group described elsewhere herein.

[0163] In some instances, as used herein, “electron-withdrawing group” may refer to an “electron-withdrawing protecting group.”

[0164] In some embodiments, L1is a linker, a bond, or absent. In some embodiments, L1is a linker. In some embodiments, L1is a bond. In some embodiments, L1is absent.

[0165] In some embodiments, R1, R2, and R3are each independently hydrogen, a protecting group, an amino-protecting group, an electron-withdrawing group, an amino-electron- withdrawing group, or an electron-withdrawing protecting group.

[0166] In some embodiments, R1is hydrogen. In some embodiments, R1is a protecting group. In some embodiments, R1is an amino-protecting group. In some embodiments, R1is an amino-electron-withdrawing group. In some embodiments, R1is an electron-withdrawing group (e.g., electron-withdrawing protecting group).

[0167] In some embodiments, R2is hydrogen. In some embodiments, R2is a protecting group. In some embodiments, R2is an amino-protecting group. In some embodiments, R2is an amino-electron-withdrawing group. In some embodiments, R2is an electron-withdrawing group (e.g., electron-withdrawing protecting group).

[0168] In some embodiments, R3is hydrogen. In some embodiments, R3is a protecting group. In some embodiments, R3is an amino-protecting group. In some embodiments, R3is an amino-electron-withdrawing group. In some embodiments, R3is an electron-withdrawing group (e.g., electron-withdrawing protecting group).

[0169] In some embodiments, B is a reactive moiety or a polymer. In some embodiments, B is a reactive moiety, such as a reactive moiety described elsewhere herein. In some embodiments, B is a polymer.

[0170] In some embodiments, provided herein is a compound represented by Formula (I-Aa) or Formula (I-Ab):Formula (I-Aa) Formula (I-Ab).

[0171] In some embodiments, provided herein is a compound represented by Formula (I-Aa). In some embodiments, provided herein is a compound represented by Formula (I-Ab). In some embodiments, the compound of Formula (I-Aa) or Formula (I-Ab) is a sequencing reagent.

[0172] In some instances, R1, R2, R3, L1, and B are described hereinabove.

[0173] In some embodiments, Ring A is C4-C12 heteroaryl or C4-C12 heterocycle, optionally substituted with one or more electron-withdrawing groups. In some embodiments, Ring A is C4- C12 heteroaryl. In some embodiments, Ring A is C4-C12 heteroaryl, substituted with one or more electron-withdrawing groups. In some embodiments, Ring A is C4-C12 heterocycle. In someembodiments, Ring A is C4-C12 heterocycle, substituted with one or more electron- withdrawing groups.

[0174] In some embodiments, provided herein is a compound of Formula (Ila) or Formula (lib):Formula (Ila) Formula (lib).

[0175] In some embodiments, provided herein is a compound of Formula (Ila). In some embodiments, provided herein is a compound of Formula (lib).

[0176] In some embodiments, the compound of Formula (Ila) or Formula (lib) is a sequencing reagent.

[0177] In some embodiments, LG is a leaving group, such as a leaving group described elsewhere herein.

[0178] In some embodiments, the leaving group is optionally substituted.

[0179] In some embodiments, the leaving group is optionally substituted with one electronwithdrawing group (e.g., electron-withdrawing protecting group). In some embodiments, the leaving group is substituted with one electron-withdrawing group (e.g., electron-withdrawing protecting group). In some embodiments, the leaving group is optionally substituted with one or more electron-withdrawing group (e.g., electron-withdrawing protecting group). In some embodiments, the leaving group is substituted with two electron- withdraw! ng groups (e.g., electron-withdrawing protecting groups). In some embodiments, the leaving group is substituted with three electron-withdrawing groups (e.g., electron-withdrawing protecting groups). In some embodiments, the leaving group is substituted with four electron-withdrawing groups (e.g., electron-withdrawing protecting groups). In some embodiments, the leaving group is substituted with five electron -withdrawing groups (e.g., electron-withdrawing protecting groups). In some instances, an electron withdrawing group is an electron withdrawing group described elsewhere herein.

[0180] In some embodiments, L1is a linker, a bond, or absent. In some embodiments, L1is a linker. In some embodiments, L1is a bond. In some embodiments, L1is absent.

[0181] In some embodiments, R1is hydrogen, a protecting group, an amino-protecting group, an electron-withdrawing group, an amino-electron-withdrawing group, or an electron-withdrawing protecting group. In some embodiments, R1is hydrogen. In some embodiments, R1is a protecting group. In some embodiments, R1is an amino-protecting group. In some embodiments, R1is an amino-electron-withdrawing group. In some embodiments, R1is an electron-withdrawing group (e.g., electron-withdrawing protecting group).

[0182] In some embodiments, R2is hydrogen, a protecting group, an amino-protecting group, an electron-withdrawing group, an amino-electron-withdrawing group, or an electronwithdrawing protecting group. In some embodiments, R2is hydrogen. In some embodiments, R2is a protecting group. In some embodiments, R2is an amino-protecting group. In some embodiments, R2is an amino-electron-withdrawing group. In some embodiments, R2is an electron-withdrawing group (e.g., electron-withdrawing protecting group).

[0183] In some embodiments, B is a reactive moiety or a polymer. In some embodiments, B is a reactive moiety, such as a reactive moiety described elsewhere herein. In some embodiments, B is a polymer.

[0184] In some embodiments, n is an integer from 0 to 3. In some embodiments, n is 0. In some embodiments, n is 1. In some embodiments, n is 2. In some embodiments, n is 3.

[0185] In some embodiments, provided herein is a compound of Formula (II-Aa) or Formula (II- Ab):Formula (II- Aa) Formula (Il-Ab).

[0186] In some embodiments, provided herein is a compound of Formula (II-Aa). In some embodiments, provided herein is a compound of Formula (Il-Ab).

[0187] In some embodiments, the compound of Formula (II-Aa) or Formula (Il-Ab) is a sequencing reagent.

[0188] In some instances, R1, R2, L1, and B are described hereinabove.

[0189] In some embodiments, Ring A is C4-C12 heteroaryl or C4-C12 heterocycle, optionally substituted with one or more electron-withdrawing groups. In some embodiments, Ring A is C4- C12 heteroaryl. In some embodiments, Ring A is C4-C12 heteroaryl, substituted with one or more electron-withdrawing groups. In some embodiments, Ring A is C4-C12 heterocycle. In some embodiments, Ring A is C4-C12 heterocycle, substituted with one or more electron- withdrawing groups.

[0190] In some embodiments, provided herein is a compound of Formula (Illa), Formula (Illb), Formula (IVa), or Formula (IVb):Formula (Illa), Formula (Illb), Formula (IVa), Formula (IVb).

[0191] In some embodiments, provided herein is a compound of Formula (Illa). In some embodiments, provided herein is a compound of Formula (Illb). In some embodiments, provided herein is a compound of Formula (IVa). In some embodiments, provided herein is a compound of Formula (IVb).

[0192] In some embodiments, the compound of Formula (Illa), Formula (Illb), Formula (IVa), or Formula (IVb) is a sequencing reagent.

[0193] In some embodiments, LG is a leaving group, such as a leaving group described elsewhere herein.

[0194] In some embodiments, the leaving group is optionally substituted. In some embodiments, the leaving group is optionally substituted with one electron-withdrawing group (e.g., electron-withdrawing protecting group). In some embodiments, the leaving group is optionally substituted with one or more electron-withdrawing group (e.g., electron-withdrawing protecting group). In some embodiments, the leaving group is substituted with two electronwithdrawing groups (e.g., electron-withdrawing protecting groups). In some embodiments, the leaving group is substituted with three electron-withdrawing groups (e.g., electron-withdrawing protecting groups). In some embodiments, the leaving group is substituted with four electronwithdrawing groups (e.g., electron-withdrawing protecting groups). In some embodiments, the leaving group is substituted with five electron-withdrawing groups (e.g., electron-withdrawing protecting groups). In some instances, an electron withdrawing group is an electron withdrawing group described elsewhere herein.

[0195] In some embodiments, L1is a linker, a bond, or absent. In some embodiments, L1is a linker. In some embodiments, L1is a bond. In some embodiments, L1is absent.

[0196] In some embodiments, L2is a hydrogen or a linker. In some embodiments, L2is a hydrogen. In some embodiments, L2is a linker.

[0197] In some embodiments, L3and L4are independently linkers, taken together to form heterocycle. In some embodiments, L3and L4are taken together to form Cs-Cs heterocycle.

[0198] In some embodiments, B is a reactive moiety or a polymer. In some embodiments, B is a reactive moiety, such as a reactive moiety described elsewhere herein. In some embodiments, B is a polymer.

[0199] In some embodiments, R1is hydrogen, a protecting group, an amino-protecting group, an electron-withdrawing group, an amino-electron-withdrawing group, or an electronwithdrawing protecting group. In some embodiments, R1is hydrogen. In some embodiments, R1is a protecting group. In some embodiments, R1is an amino-protecting group. In some embodiments, R1is an amino-electron-withdrawing group. In some embodiments, R1is an electron-withdrawing group (e.g., electron-withdrawing protecting group).

[0200] In some embodiments, R2is hydrogen, a protecting group, an amino-protecting group, an electron-withdrawing group, an amino-electron-withdrawing group, or an electronwithdrawing protecting group. In some embodiments, R2is hydrogen. In some embodiments, R2is a protecting group. In some embodiments, R2is an amino-protecting group. In some embodiments, R2is an amino-electron-withdrawing group. In some embodiments, R2is an electron-withdrawing group (e.g., electron-withdrawing protecting group).

[0201] In some embodiments, provided herein is a compound of Formula (III-Aa), Formula (III-Ab), Formula (IV-Aa), or Formula (IV-Ab):Formula (III-Aa), Formula (III-Ab), Formula (IV-Aa), Formula (IV-Ab).

[0202] In some embodiments, provided herein is a compound of Formula (III-Aa). In some embodiments, provided herein is a compound of Formula (III-Ab). In some embodiments, provided herein is a compound of Formula (IV-Aa). In some embodiments, provided herein is a compound of Formula (IV-Ab).

[0203] In some embodiments, the compound of Formula (III-Aa), Formula (III-Ab), Formula (IV-Aa), or Formula (IV-Ab) is a sequencing reagent.

[0204] In some instances, R1, R2, L1, L2, L3, L4, and B are described hereinabove.

[0205] In some embodiments, Ring A is C4-C12 heteroaryl or C4-C12 heterocycle, optionally substituted with one or more electron-withdrawing groups. In some embodiments, Ring A is C4- C12 heteroaryl. In some embodiments, Ring A is C4-C12 heteroaryl, substituted with one or moreelectron-withdrawing groups. In some embodiments, Ring A is C4-C12 heterocycle. In some embodiments, Ring A is C4-C12 heterocycle, substituted with one or more electron- withdrawing groups.

[0206] In some embodiments, L3and L4are taken together to form Cs-Cs heterocycle.

[0207] In some embodiments, provided herein is a compound of Formula (IV-Ba) or Formula(IV-Bb):Formula (IV-Ba) Formula (IV-Bb).

[0208] In some embodiments, provided herein is a compound of Formula (IV-Ba). In some embodiments, provided herein is a compound of Formula (IV-Bb).

[0209] In some embodiments, the compound of Formula (IV-Ba) or Formula (IV-Bb) is a sequencing reagent.

[0210] In some instances, R1, R2, L1, and B are described hereinabove.

[0211] In some instances, n is an integer from 0 to 3. In some instances, n is 0. In some instances, n is 1. In some instances, n is 2. In some instances, n is 3.

[0212] In some embodiments, provided herein is a compound of Formula (V):Formula (V).

[0213] In some embodiments, the compound of Formula (V) is a sequencing reagent.

[0214] In some embodiments, each LG is independently a leaving group, optionally substituted with one or more electron-withdrawing groups. In some embodiments, each LG is a leaving group, such as a leaving group described elsewhere herein. In some embodiments, each LG is independently a leaving group, substituted with one or more electron-withdrawing groups.

[0215] In some embodiments, the leaving group is optionally substituted. In some embodiments, the leaving group is optionally substituted with one electron-withdrawing group. In some embodiments, the leaving group is optionally substituted with one or more electronwithdrawing group. In some embodiments, the leaving group is substituted with one electron-withdrawing groups. In some embodiments, the leaving group is substituted with two electronwithdrawing groups. In some embodiments, the leaving group is substituted with three electronwithdrawing groups. In some embodiments, the leaving group is substituted with four electronwithdrawing groups. In some embodiments, the leaving group is substituted with five electronwithdrawing groups. In some instances, an electron withdrawing group is an electron withdrawing group described elsewhere herein.

[0216] In some embodiments, L1is a linker, a bond, or absent. In some embodiments, L1is a linker. In some embodiments, L1is a bond. In some embodiments, L1is absent.

[0217] In some embodiments, B is a reactive moiety or a polymer. In some embodiments, B is a reactive moiety, such as a reactive moiety described elsewhere herein. In some embodiments, B is a polymer.

[0218] In some embodiments, provided herein is a compound of Formula (V-A) or Formula (V-B):Formula (V-A) Formula (V-B).

[0219] In some embodiments, provided herein is a compound of Formula (V-A). In some embodiments, provided herein is a compound of Formula (V-B).

[0220] In some embodiments, the compound of Formula (V-A) or Formula (V-B) is a sequencing reagent.

[0221] In some instances, LG, L1, and B are described hereinabove.

[0222] In some embodiments, each Ring A is independently C4-C12 heteroaryl or C4-C12 heterocycle, optionally substituted with one or more electron-withdrawing groups. In some embodiments, each Ring A is independently C4-C12 heteroaryl. In some embodiments, each Ring A is independently C4-C12 heteroaryl, substituted with one or more electron-withdrawing groups. In some embodiments, each Ring A is independently C4-C12 heterocycle. In some embodiments, each Ring A is independently C4-C12 heterocycle, substituted with one or more electronwithdrawing groups.

[0223] In some embodiments, provided herein is a compound of Formula (VI) or Formula (VII):Formula (VI) Formula (VII).

[0224] In some embodiments, provided herein is a compound of Formula (VI). In some embodiments, provided herein is a compound of Formula (VII).

[0225] In some embodiments, the compound of Formula (VI) or Formula (VII) is a sequencing reagent.

[0226] In some embodiments, each LG is a leaving group, such as a leaving group described elsewhere herein.

[0227] In some embodiments, the leaving group is optionally substituted. In some embodiments, the leaving group is optionally substituted with one electron-withdrawing group (e.g., electron-withdrawing protecting group). In some embodiments, the leaving group is optionally substituted with one or more electron-withdrawing group (e.g., electron-withdrawing protecting group). In some embodiments, the leaving group is substituted with two electronwithdrawing groups (e.g., electron-withdrawing protecting groups). In some embodiments, the leaving group is substituted with three electron-withdrawing groups (e.g., electron-withdrawing protecting groups). In some embodiments, the leaving group is substituted with four electronwithdrawing groups (e.g., electron-withdrawing protecting groups). In some embodiments, the leaving group is substituted with five electron-withdrawing groups (e.g., electron-withdrawing protecting groups). In some instances, an electron withdrawing group is an electron withdrawing group described elsewhere herein.

[0228] In some embodiments, L1is a linker, a bond, or absent. In some embodiments, L1is a linker. In some embodiments, L1is a bond. In some embodiments, L1is absent.

[0229] In some embodiments, L2is a hydrogen or a linker. In some embodiments, L2is a hydrogen. In some embodiments, L2is a linker.

[0230] In some embodiments, L3and L4are independently linkers, taken together to form heterocycle. In some embodiments, L3and L4are taken together to form Cs-Cs heterocycle.

[0231] In some embodiments, B is a reactive moiety or a polymer. In some embodiments, B is a reactive moiety, such as a reactive moiety described elsewhere herein. In some embodiments, B is a polymer.

[0232] In some embodiments, provided herein is a compound of Formula (VI-A) or Formula (VII- A):Formula (VI- A) Formula (VII- A).

[0233] In some embodiments, provided herein is a compound of Formula (VI-A). In some embodiments, provided herein is a compound of Formula (VII-A).

[0234] In some embodiments, the compound of Formula (VI-A) or Formula (VII-A) is a sequencing reagent.

[0235] In some instances, L1, L2, L3, L4, and B are described hereinabove.

[0236] In some embodiments, Ring A is C4-C12 heteroaryl or C4-C12 heterocycle, optionally substituted with one or more electron-withdrawing groups. In some embodiments, Ring A is C4- C12 heteroaryl. In some embodiments, Ring A is C4-C12 heteroaryl, substituted with one or more electron-withdrawing groups. In some embodiments, Ring A is C4-C12 heterocycle. In some embodiments, Ring A is C4-C12 heterocycle, substituted with one or more electron- withdrawing groups.

[0237] In some embodiments, L3and L4are taken together to form Cs-Cs heterocycle.

[0238] In some embodiments, the leaving group is electron-withdrawing.

[0239] In some embodiments, the leaving group is C4-C12 heteroaryl. In some embodiments, the leaving group is an unsubstituted C4-C12 heteroaryl. In some embodiments, the leaving group is C4-C12 heteroaryl optionally substituted with one electron-withdrawing group. In some embodiments, the leaving group is C4-C12 heteroaryl optionally substituted with one or more electron-withdrawing groups. In some embodiments, the leaving group is C4-C12 heteroaryl optionally substituted with two electron-withdrawing groups. In some embodiments, the leaving group is C4-C12 heteroaryl optionally substituted with three electron-withdrawing groups. In some embodiments, the leaving group is C4-C12 heteroaryl optionally substituted with four electron-withdrawing groups. In some embodiments, the leaving group is C4-C12 heteroaryl optionally substituted with five electron-withdrawing groups.

[0240] In some embodiments, the leaving group is C4-C12 heterocycle. In some embodiments, the leaving group is an unsubstituted C4-C12 heterocycle. In some embodiments, the leaving group is C4-C12 heterocycle optionally substituted with one electron-withdrawing group. In some embodiments, the leaving group is C4-C12 heterocycle optionally substituted with one or more electron-withdrawing groups. In some embodiments, the leaving group is C4-C12 heterocycle optionally substituted with two electron-withdrawing groups. In some embodiments,the leaving group is C4-C12 heterocycle optionally substituted with three electron-withdrawing groups. In some embodiments, the leaving group is C4-C12 heterocycle optionally substituted with four electron-withdrawing groups. In some embodiments, the leaving group is C4-C12 heterocycle optionally substituted with five electron-withdrawing groups.

[0241] In some embodiments, the leaving group is C4-C6 heteroaryl. In some embodiments, the leaving group is an unsubstituted C4-C6 heteroaryl. In some embodiments, the leaving group is C4-C6 heteroaryl optionally substituted with one electron-withdrawing group. In some embodiments, the leaving group is C4-C6 heteroaryl optionally substituted with one or more electron-withdrawing groups. In some embodiments, the leaving group is C4-C6 heteroaryl optionally substituted with two electron-withdrawing groups. In some embodiments, the leaving group is C4-C6 heteroaryl optionally substituted with three electron-withdrawing groups. In some embodiments, the leaving group is C4-C6 heteroaryl optionally substituted with four electron-withdrawing groups. In some embodiments, the leaving group is C4-C6 heteroaryl optionally substituted with five electron-withdrawing groups.

[0242] In some embodiments, the leaving group is an azole. In some embodiments, the leaving group is an unsubstituted azole. In some embodiments, the leaving group is an azole optionally substituted with one electron-withdrawing group. In some embodiments, the leaving group is an azole optionally substituted with one or more electron-withdrawing groups. In some embodiments, the leaving group is an azole optionally substituted with two electronwithdrawing groups. In some embodiments, the leaving group is an azole optionally substituted with three electron-withdrawing groups.

[0243] In some embodiments, the leaving group is an azine. In some embodiments, the leaving group is an unsubstituted azine. In some embodiments, the leaving group is an azine optionally substituted with one electron-withdrawing group. In some embodiments, the leaving group is an azine optionally substituted with one or more electron-withdrawing groups. In some embodiments, the leaving group is an azine optionally substituted with two electronwithdrawing groups. In some embodiments, the leaving group is an azine optionally substituted with three electron-withdrawing groups.

[0244] In some embodiments, the leaving group is a diazole. In some embodiments, the leaving group is an unsubstituted diazole. In some embodiments, the leaving group is a diazole optionally substituted with one electron-withdrawing group. In some embodiments, the leaving group is a diazole optionally substituted with one or more electron-withdrawing groups. In some embodiments, the leaving group is a diazole optionally substituted with two electron-withdrawing groups. In some embodiments, the leaving group is a diazole optionally substituted with three electron-withdrawing groups.

[0245] In some embodiments, the leaving group is a triazole. In some embodiments, the leaving group is an unsubstituted triazole. In some embodiments, the leaving group is a triazole optionally substituted with one electron-withdrawing group. In some embodiments, the leaving group is a triazole optionally substituted with one or more electron-withdrawing groups. In some embodiments, the leaving group is a triazole optionally substituted with two electronwithdrawing groups. In some embodiments, the leaving group is a triazole optionally substituted with three electron-withdrawing groups.

[0246] In some embodiments, the leaving group is a tetrazole. In some embodiments, the leaving group is an unsubstituted tetrazole. In some embodiments, the leaving group is a tetrazole optionally substituted with one electron-withdrawing group. In some embodiments, the leaving group is a tetrazole optionally substituted with one or more electron-withdrawing groups. In some embodiments, the leaving group is a tetrazole optionally substituted with two electron-withdrawing groups. In some embodiments, the leaving group is a tetrazole optionally substituted with three electron-withdrawing groups.

[0247] In some embodiments, the leaving groupsomeX-X ' 'k 'X" embodiments, the leaving group is. In some embodiments, the leaving group is

[0248] In some embodiments, each X is independently selected from N, CH, or C-EWG (e.g., C substituted with an electron withdrawing group provided elsewhere herein). In some embodiments, X is N. In some embodiments, X is CH. In some embodiments, X is C-EWG. In some embodiments, EWG is an electron-withdrawing group. In some embodiments, at least one X is N. In some embodiments, one X is N. In some embodiments, at least two X are N.

[0249] In some embodiments, each Y is independently selected from N, CH, or C-EWG (e.g., C substituted with an electron withdrawing group provided elsewhere herein). In some embodiments, Y is N. In some embodiments, Y is CH. In some embodiments, Y is C-EWG. Insome embodiments, EWG is an electron-withdrawing group. In some embodiments, at least one Y is N. In some embodiments, one Y is N. In some embodiments, at least two Y are N.X3-X4 / / x2x5

[0250] In some embodiments, the leaving group is. In other embodiments, the leaving group ifurther embodiments, Ring

[0251] In some embodiments, each of X1-X5 are independently selected from N, CH, or C- EWG (e.g., C substituted with an electron withdrawing group provided elsewhere herein). In some embodiments, at least one of X1-X5 is N. In some embodiments, at least one of X1-X5 isCH. In some embodiments, at least one of X1-X5 is C-EWG.

[0252] In some embodiment, each of Yi-YHs independently selected from N, CH, or C-EWG. In some embodiments, at least one of Yi-YHs N. In some embodiments, at least one ofYi-YHs CH. In some embodiments, at least one of Yi-YHs C-EWG.

[0253] In some embodiments, the leaving groupoptionally substituted with one or more electron-withdrawing groups. In some embodiments,is unsubstituted. In some embodiments, is substituted with one electron-" withdrawing group. In some embodiments, is substituted with more than one electron- withdrawing group. In some embodiments, is unsubstituted. In some embodiments, - is substituted with one electron-withdrawing group. In some embodiments, is itNA"hr substituted with more than one electron-withdrawing group. In some embodiments, ~J~>- isunsubstituted. In some embodiments, is substituted with one electron-withdrawinggroup. In some embodiments,is substituted with more than one electron-withdrawinggroup. In some embodiments,is unsubstituted. In some embodiments,isN-Nsubstituted with one electron-withdrawing group. In some embodiments,is substituted with more than one electron-withdrawing group. In some embodiments, the leaving group is

[0254] In some embodiments, each X is independently selected from N, CH, or C-EWG. In some embodiments, X is N. In some embodiments, X is CH. In some embodiments, X is C- EWG. In some embodiments, EWG is an electron-withdrawing group. In some embodiments, at least one X is N. In some embodiments, one X is N. In some embodiments, more than one X is N.

[0255] In some embodiments, the leaving groupunsubstituted. In some embodiments,substituted with one electron-withdrawing group. Insubstituted with more than one electron-withdrawing group. In some embodiments, the leaving group, optionally substituted with at least one electron-withdrawing group.

[0256] In some embodiments, the leaving groupoptionally substituted with at least one electron-withdrawing group.

[0257] In some embodiments, the leaving groupoptionally substituted with at least one electron-withdrawing group.

[0258] In some embodiments, the leaving groupoptionally substituted with at least one electron-withdrawing group.

[0259] In some embodiments, the leaving group, optionally substituted with at least one electron-withdrawing group.

[0260] In some embodiments, the leaving groupoptionally substituted with at least one electron-withdrawing group.

[0261] In some embodiments, the leaving groupoptionally substituted with at least one electron-withdrawing group.

[0262] In some embodiments, the leaving group is C4-C6 heteroaryl fused with C4-C6 aryl, optionally substituted with one or more electron-withdrawing groups. In some embodiments, the leaving group is C4-C6 heteroaryl fused with C4-C6 aryl. In some embodiments, the C4-C6 heteroaryl fused with C4-C6 aryl is unsubstituted. In some embodiments, the C4-C6 heteroaryl fused with C4-C6 aryl is substituted with one or more electron-withdrawing groups. In some embodiments, the C4-C6 heteroaryl fused with C4-C6 aryl is substituted with one electronwithdrawing group. In some embodiments, the C4-C6 heteroaryl fused with C4-C6 aryl is substituted with two electron-withdrawing groups. In some embodiments, the C4-C6 heteroaryl fused with C4-C6 aryl is substituted with three electron-withdrawing groups. In some embodiments, the C4-C6 heteroaryl fused with C4-C6 aryl is substituted with four electronwithdrawing groups.

[0263] In some embodiments, the leaving group is diazole, triazole, or tetrazole, fused with aryl, optionally substituted with one or more electron- withdrawing groups. In some embodiments, the leaving group is diazole fused with aryl. In some embodiments, the diazole fused with aryl is unsubstituted. In some embodiments, the diazole fused with aryl is substitutedwith one or more electron-withdrawing groups. In some embodiments, the diazole fused with aryl is substituted with one electron-withdrawing group. In some embodiments, the leaving group is triazole fused with aryl. In some embodiments, the triazole fused with aryl is unsubstituted. In some embodiments, the triazole fused with aryl is substituted with one or more electron-withdrawing groups. In some embodiments, the triazole fused with aryl is substituted with one electron-withdrawing group. In some embodiments, the leaving group is tetrazole fused with aryl. In some embodiments, the tetrazole fused with aryl is unsubstituted. In some embodiments, the tetrazole fused with aryl is substituted with one or more electron-withdrawing groups. In some embodiments, the tetrazole fused with aryl is substituted with one electronwithdrawing group.

[0264] In some embodiments, the leaving group

[0265] In some embodiments, Ring A is C4-C12 heteroaryl. In some embodiments, Ring A is an unsubstituted C4-C12 heteroaryl. In some embodiments, Ring A is C4-C12 heteroaryl optionally substituted with one electron-withdrawing group. In some embodiments, Ring A is C4-C12 heteroaryl optionally substituted with one or more electron-withdrawing groups.

[0266] In some embodiments, Ring A is C4-C12 heterocycle. In some embodiments, Ring A is an unsubstituted C4-C12 heterocycle. In some embodiments, Ring A is C4-C12 heterocycle optionally substituted with one electron-withdrawing group. In some embodiments, Ring A is C4- C12 heterocycle optionally substituted with one or more electron-withdrawing groups.

[0267] In some embodiments, Ring A is C4-C6 heteroaryl. In some embodiments, Ring A is an unsubstituted C4-C6 heteroaryl. In some embodiments, Ring A is C4-C6 heteroaryl optionally substituted with one electron-withdrawing group. In some embodiments, Ring A is C4-C6 heteroaryl optionally substituted with one or more electron-withdrawing groups. In some embodiments, Ring A is C4-C6 heteroaryl optionally substituted with two electron-withdrawing groups. In some embodiments, Ring A is C4-C6 heteroaryl optionally substituted with threeelectron-withdrawing groups. In some embodiments, Ring A is C4-C6 heteroaryl optionally substituted with four electron-withdrawing groups.

[0268] In some embodiments, Ring A is an azole. In some embodiments, Ring A is an unsubstituted azole. In some embodiments, Ring A is an azole optionally substituted with one electron-withdrawing group. In some embodiments, Ring A is an azole optionally substituted with one or more electron-withdrawing groups. In some embodiments, Ring A is azole optionally substituted with two electron-withdrawing groups. In some embodiments, Ring A is azole optionally substituted with three electron-withdrawing groups. In some embodiments, Ring A is azole optionally substituted with four electron-withdrawing groups.

[0269] In some embodiments, Ring A is an azine. In some embodiments, Ring A is an unsubstituted azine. In some embodiments, Ring A is an azine optionally substituted with one electron-withdrawing group. In some embodiments, Ring A is an azine optionally substituted with one or more electron-withdrawing groups. In some embodiments, Ring A is azine optionally substituted with two electron-withdrawing groups. In some embodiments, Ring A is azine optionally substituted with three electron-withdrawing groups. In some embodiments, Ring A is azine optionally substituted with four electron-withdrawing groups.

[0270] In some embodiments, Ring A is a diazole. In some embodiments, Ring A is an unsubstituted diazole. In some embodiments, Ring A is a diazole optionally substituted with one electron-withdrawing group. In some embodiments, Ring A is a diazole optionally substituted with one or more electron-withdrawing groups. In some embodiments, Ring A is diazole optionally substituted with two electron-withdrawing groups. In some embodiments, Ring A is diazole optionally substituted with three electron- withdrawing groups. In some embodiments, Ring A is diazole optionally substituted with four electron-withdrawing groups.

[0271] In some embodiments, Ring A is a triazole. In some embodiments, Ring A is an unsubstituted triazole. In some embodiments, Ring A is a triazole optionally substituted with one electron-withdrawing group. In some embodiments, Ring A is a triazole optionally substituted with one or more electron-withdrawing groups. In some embodiments, Ring A is triazole optionally substituted with two electron-withdrawing groups. In some embodiments, Ring A is triazole optionally substituted with three electron-withdrawing groups. In some embodiments, Ring A is triazole optionally substituted with four electron-withdrawing groups.

[0272] In some embodiments, Ring A is a tetrazole. In some embodiments, Ring A is an unsubstituted tetrazole. In some embodiments, Ring A is a tetrazole optionally substituted with one electron-withdrawing group. In some embodiments, Ring A is a tetrazole optionally substituted with one or more electron-withdrawing groups. In some embodiments, Ring A istetrazole optionally substituted with two electron-withdrawing groups. In some embodiments, Ring A is tetrazole optionally substituted with three electron-withdrawing groups. In some embodiments, Ring A is tetrazole optionally substituted with four electron-withdrawing groups.X-X k' ''X'

[0273] In some embodiments, Ring A is.

[0274] In some embodiments, each X is independently selected from N, CH, or C-EWG. In some embodiments, X is N. In some embodiments, X is CH. In some embodiments, X is C- EWG. In some embodiments, EWG is an electron-withdrawing group. In some embodiments, at least one X is N. In some embodiments, one X is N. In some embodiments, more than one X is N.

[0275] In some embodiments, Ringoptionallysubstituted with one or more electron-withdrawing groups. In some embodiments, ~~ -~ isunsubstituted. In some embodiments,is substituted with one electron-withdrawing group.In some embodiments, —h™ issubstituted with more than one electron-withdrawing group. Insome embodiments, ~™L~- is unsubstituted. In some embodiments,is substituted with oneelectron-withdrawing group. In some embodiments, ~~J™~ is substituted with more than oneelectron-withdrawing group. In some embodiments, -J— is unsubstituted. In someembodiments,is substituted with one electron-withdrawing group. In someembodiments,is substituted with more than one electron-withdrawing group. In someembodiments,is unsubstituted. In some embodiments,is substituted with oneelectron-withdrawing group. In some embodiments,is substituted with more than one electron-withdrawing group.X3-X4 / / x2x5xf

[0277] In some embodiments, Ring A is . In other embodiments, Ring A isfurther embodiments, Ring

[0278] In some embodiments, each of X1-X5 are independently selected from N, CH, or C-EWG (e.g., C substituted with an electron withdrawing group provided elsewhere herein). In some embodiments, at least one of X1-X5 is N. In some embodiments, at least one of X1-X5 isCH. In some embodiments, at least one of X1-X5 is C-EWG.

[0279] In some embodiment, each of Yi-YHs independently selected from N, CH, or C- EWG. In some embodiments, at least one of Yi-YHs N. In some embodiments, at least one of Yi-YHs CH. In some embodiments, at least one of Yi-YHs C-EWG.

[0280] In some embodiments, each X is independently selected from N, CH, or C-EWG. In some embodiments, X is N. In some embodiments, X is CH. In some embodiments, X is C-EWG. In some embodiments, at least one X is N. In some embodiments, one X is N. In some embodiments, more than one X is N.

[0281] In some embodiments, each Y is independently selected from N, CH, or C-EWG. In some embodiments, Y is N. In some embodiments, Y is CH. In some embodiments, Y is C-EWG. In some embodiments, at least one Y is N. In some embodiments, one Y is N. In some embodiments, more than one Y is N.

[0282] In some embodiments, EWG is an electron-withdrawing group (e.g., electronwithdrawing protecting group) (e.g., such as an electron withdrawing group described elsewhere herein).unsubstituted. In some embodiments,substituted with one electron-withdrawing group(e.g., electron-withdrawing protecting group). In somesubstituted with more than one electron-withdrawing group.

[0284] In some embodiments, Ring, optionally substituted with at least one electron-withdrawing group.

[0285] In some embodiments, Ringoptionally substituted with at least one electron-withdrawing group.

[0286] In some embodiments, Ringoptionally substituted with at least one electron-withdrawing group.

[0287] In some embodiments, Ringoptionally substituted with at least one electron-withdrawing group.

[0288] In some embodiments, Ring, optionally substituted with at least one electron-withdrawing group.

[0289] In some embodiments, Ringoptionally substituted with at least one electron-withdrawing group.

[0290] In some embodiments, Ringoptionally substituted with at least one electron-withdrawing group.

[0291] In some embodiments, Ring

[0292] In some embodiments, Ring A is C4-C6 heteroaryl fused with C4-C6 aryl, optionally substituted with one or more electron-withdrawing groups. In some embodiments, Ring A is C4- Ce heteroaryl fused with C4-C6 aryl. In some embodiments, the C4-C6 heteroaryl fused with C4- Ce aryl is unsubstituted. In some embodiments, the C4-C6 heteroaryl fused with C4-C6 aryl is substituted with one or more electron-withdrawing groups. In some embodiments, the C4-C6 heteroaryl fused with C4-C6 aryl is substituted with one electron- withdraw! ng group. In some embodiments, the C4-C6 heteroaryl fused with C4-C6 aryl is substituted with two electronwithdrawing groups. In some embodiments, the C4-C6 heteroaryl fused with C4-C6 aryl is substituted with three electron-withdrawing groups. In some embodiments, the C4-C6 heteroaryl fused with C4-C6 aryl is substituted with four electron-withdrawing groups. In some embodiments, the C4-C6 heteroaryl fused with C4-C6 aryl is substituted with five electronwithdrawing groups.

[0293] In some embodiments, Ring A is diazole, triazole, or tetrazole, fused with aryl, optionally substituted with one or more electron-withdrawing groups.

[0294] In some embodiments, Ring A is diazole fused with aryl. In some embodiments, the diazole fused with aryl is unsubstituted. In some embodiments, the diazole fused with aryl is substituted with one or more electron-withdrawing groups. In some embodiments, the diazole fused with aryl is substituted with one electron-withdrawing group. In some embodiments, the diazole fused with aryl is substituted with two electron-withdrawing groups. In some embodiments, the diazole fused with aryl is substituted with three electron-withdrawing groups. In some embodiments, the diazole fused with aryl is substituted with four electron-withdrawing groups. In some embodiments, the diazole fused with aryl is substituted with five electronwithdrawing groups.

[0295] In some embodiments, Ring A is triazole fused with aryl. In some embodiments, the triazole fused with aryl is unsubstituted. In some embodiments, the triazole fused with aryl is substituted with one or more electron-withdrawing groups. In some embodiments, the triazole fused with aryl is substituted with one electron-withdrawing group. In some embodiments, Ring A is tetrazole fused with aryl. In some embodiments, the tetrazole fused with aryl is unsubstituted. In some embodiments, the tetrazole fused with aryl is substituted with one or more electron-withdrawing groups. In some embodiments, the tetrazole fused with aryl is substituted with one electron-withdrawing group. In some embodiments, the tetrazole fused with aryl is substituted with two electron-withdrawing groups. In some embodiments, the tetrazole fused with aryl is substituted with three electron-withdrawing groups. In some embodiments, the tetrazole fused with aryl is substituted with four electron-withdrawing groups. In some embodiments, the tetrazole fused with aryl is substituted with five electron-withdrawing groups.

[0296] An electron-withdrawing group as used herein may refer to halogen, haloalkyl, nitro, sulfonate, amino, alkylamino, cyano, or a carbonyl.

[0297] In some embodiments, the electron-withdrawing group is each independently selected from halogen, haloalkyl, nitro, sulfonate, amino, alkylamino, cyano, or a carbonyl. In some embodiments, the electron-withdrawing group is halogen. In some embodiments, the electron-withdrawing group is fluorine. In some embodiments, the electron-withdrawing group is chlorine. In some embodiments, the electron-withdrawing group is haloalkyl. In some embodiments, the electron-withdrawing group is CF3. In some embodiments, the electronwithdrawing group is nitro. In some embodiments, the electron-withdrawing group is sulfonate. In some embodiments, the electron-withdrawing group is amino. In some embodiments, the electron-withdrawing group is alkylamino. In some embodiments, the electron-withdrawing group is cyano. In some embodiments, the electron-withdrawing group is carbonyl. In some embodiments, the carbonyl is an acetyl group. In some embodiments, the carbonyl is an acetylderivative. In some embodiments, the carbonyl is a carboxylic acid. In some embodiments, the carbonyl is an aldehyde.

[0298] In some embodiments, at least one electron-withdrawing group comprises Ci-Ce haloalkyl. In some embodiments, one electron-withdrawing group is Ci-Ce haloalkyl. In some embodiments, more than one electron-withdrawing group is Ci-Ce haloalkyl.

[0299] In some embodiments, at least one electron- withdrawing group comprises CF3. In some embodiments, one electron-withdrawing group is CF3. In some embodiments, more than one electron-withdrawing group is CF3.

[0300] In some embodiments, at least one electron-withdrawing group comprises nitro. In some embodiments, one electron-withdrawing group is nitro. In some embodiments, more than one electron-withdrawing group is nitro.

[0301] In some embodiments, at least one electron-withdrawing group comprises halogen. In some embodiments, one electron-withdrawing group is halogen. In some embodiments, more than one electron-withdrawing group is halogen. In some embodiments, the halogen is chlorine. In some embodiment, the halogen is fluorine. In some embodiment, the halogen is bromine.

[0302] In some embodiments, the linkers described herein are configured to enable nucleic acid attraction. In some embodiments, the nucleic acid attraction is ionic. In some embodiments, the nucleic acid attraction is non-ionic. In some embodiments, the nucleic acid attraction is covalent. In some embodiments, the nucleic acid attraction is van der Walls forces. In some embodiments, the nucleic acid attraction is hydrogen bonding. In some embodiments, the nucleic acid attraction is dipole-dipole forces.

[0303] In some embodiments, the linkers provided herein are configured to provide steric space. In some embodiments, the linker is configured to provide steric space for click chemistry. In some embodiments, the linker is configured to provide steric space for click chemistry via B.

[0304] In some embodiments, the linker is electron-withdrawing. In some embodiments, a first end of the linker is electron-withdrawing. In some embodiments, both a first end and a second end of the linker are electron-withdrawing.

[0305] In some embodiments, one linker is present. In some embodiments, more than one linker is present (e.g., both L1and L2are present).

[0306] In some embodiments, each linker is independently selected from optionally substituted alkyl, optionally substituted heteroalkyl, an oligonucleotide, DNA, RNA, or a peptide. In some embodiments, the linker is a substituted Ci-Ce alkyl. In some embodiments, the linker is an unsubstituted alkyl. In some embodiments, the linker is an unsubstituted Ci-Ce alkyl. In some embodiments, the linker is methylene. In some embodiments, the linker is ethylene. Insome embodiments, the linker is a substituted heteroalkyl. In some embodiments, the linker is a substituted Ci-Ce heteroalkyl. In some embodiments, the linker is an unsubstituted heteroalkyl. In some embodiments, the linker is an unsubstituted Ci-Ce heteroalkyl. In some embodiments, the linker is an unsubstituted alkoxy. In some embodiments, the linker is an unsubstituted Ci-Ce alkoxy. In some embodiments, the linker is a substituted alkoxy. In some embodiments, the linker is a substituted Ci-Ce alkoxy. In some embodiments, the linker is an oligonucleotide. In some embodiments, the linker is DNA. In some embodiments, the linker is RNA. In some embodiments, the linker is a peptide.

[0307] In some embodiments, the linker is a polyethylene glycol or derivative thereof. In some embodiments, the linker is.

[0308] In some embodiments, x is an integer greater than 0. In some embodiments, x is 1. In some embodiments, x is 2. In some embodiments, x is 3. In some embodiments, x is greater than 3. In some embodiments, x is 4. In some embodiments, x is 5.

[0309] In some embodiments, x is at least 1 (e.g., at least 2, 5, 10, 20, 30, 50, 70, 90, or 100). In other embodiments, x is at most 100 (e.g., at most 90, 70, 50, 30, 20, 10, or 5). In some embodiments, x is 1 to about 100, 1 to about 90, 1 to about 50, about 10 to about 100, about 10 to about 50, or about 10 to about 30.

[0310] In some embodiments, at least one linker is Ci-Ce alkyl. In some embodiments, at least one linker is -CH2-.

[0311] In some embodiments, L1and L2are taken together. In some embodiments, L1and L2are taken together to form Cs-Cs heterocycle.

[0312] In some embodiments, at least one of R1, R2, and R3is an amine-protecting group. In some embodiments, R1, R2, and R3are each independently hydrogen or an amine-protecting group. In some embodiments, R1is hydrogen. In some embodiments, R1is an amine-protecting group. In some embodiments, R2is hydrogen. In some embodiments, R2is an amine-protecting group. In some embodiments, R3is hydrogen. In some embodiments, R3is an amine-protecting group.

[0313] In some embodiments, the protecting group or amino-protecting group is carbobenzyl oxy (CBz) (e.g., benzyl carbamate), acetamide (Ac), trifluoroacetyl (TFAc), phthalimide, benzyl, benzylamine (Bn), benzoyl (Bz), triphenylmethyl (Tr) (e.g., triphenylmethylamine), benzylideneneamine, tosyl (Ts) (e.g., p-toluenesulfonamide), Methylsulfonylethoxycarbonyl (Msc), 2-(methylsulfonyl)ethoxy-methanone, tertbutyloxycarbonyl (Boc) (e.g., t-butyl carbamate), or fluorenylmethyloxycarbonyl (Fmoc) (e.g.,9-fluorenylmethyl carbamate), 2,7-disulfo-9-fluorenylmethoxycarbonyl (Smoc), or acetyl. In some embodiments, the protecting group is trifluoroacetyl (TFAc). In some embodiments, the protecting group is 2-(methylsulfonyl)ethoxy-methanone. In some embodiments, the protecting group is 2,7-disulfo-9-fluorenylmethoxycarbonyl (Smoc). In some embodiments, the protecting group is fluorenylmethyloxy carbonyl.

[0314] In some embodiments, the reactive moiety comprises a click chemistry moiety. The click chemistry moiety may comprise any functional group suitable for click chemistry according to one of skill in the art. In some embodiments, the click chemistry moiety is an azide. In some embodiments, the click chemistry moiety is a sulfonyl azide. In some embodiments, the click chemistry moiety is an alkyne. In some embodiments, the click chemistry moiety is a thio acid. In some embodiments, the click chemistry moiety is a tetrazine. In some embodiments, the click chemistry moiety is bicyclo[6.1.0]nonyne (BCN). In some embodiments, the click chemistry moiety is a diarylcyclooctyne (e.g., DBCO). In some embodiments, the click chemistry moiety is transcyclooctene (TCO). In some embodiments, the click chemistry moiety is cyclopropane. In some embodiments, the click chemistry moiety is norbornene. In some embodiments, the click chemistry moiety is a spiroalkene.

[0315] In some embodiments, the reactive moiety comprises a phosphite. The reactive moiety may comprise any functional group suitable to react, given the environment and reagents. In some embodiments, the reactive moiety comprises a phosphine. In some embodiments, the reactive moiety comprises a pentafluorophenyl ester (PFP). In some embodiments, the reactive moiety comprises a tetrafluorophenyl ester (TFP). In some embodiments, the reactive moiety comprises a 4-sulfo-2,3,5,6-tetrafluorophenyl ester (STP). In some embodiments, the reactive moiety comprises a thio-phthalimide.

[0316] In some embodiments, the compounds provided herein are compounds of Table 1.Table 1

[0317] In some embodiments, the compound is represented byIn some embodiments, the compound is represented byrepresentedsome embodiments, the compound is representedsome embodiments, the compound is,some embodiments, the compound is represented byembodiments, the compound is representedembodiments, the compound is representedembodiments, the compound is representedsome embodiments, the compound is representedsome embodiments, the compound is representedsome embodiments, the compound is representedsome embodiments, the compound is represented

[0318] In some embodiments, the compound is a sequencing reagent.

[0319] In some embodiments, the (e.g., amine-) modified DNA comprises a structure selected from:

[0320] In some embodiments, the compound is configured to react with the click chemistry moiety, as described elsewhere herein.

[0321] The sequencing reagents provided herein, such as sequencing reagents of Formula (la), (lb), (I-Aa), (I-Ab), (Ila), (lib), (II-Aa), (Il-Ab), (Illa), (Illb), (III-Aa), (III-Ab), (IVa), (IVb), (IV-Aa), (IV-Ab), (IV-Ba), (IV-Bb), (V), (V-A), or (V-B), may conjugate to and cleave N-terminal amino acids. Cleavage and detection may comprise conjugating the guanidyl group (e.g., guani dination) to a peptide or terminal residue of an amino acid of known molecular weight and measuring the molecular weight before and after cleavage using mass spectrometry. In some instances, one or more products of the reach on(s) may be performed, e.g., using high performance liquid chromatography (HPLC). The expected reduction in molecular weight by loss of one amino acid may signify a successful cleavage. Mass spectrometry may be performed on the peptide only without the sequencing reagent, peptide with the sequencing reagent conjugated to the N-terminal amino acid, and the cleaved peptide after the sequencing reagent removes the N-terminal amino acid.Method of Using the Sequencing Reagent Using Intramolecular Expansion or Local Tethering

[0322] Provided herein are methods of using any of the sequencing reagents provided herein.

[0323] In some embodiments, the method comprises providing a polymeric analyte. In some embodiments, the polymeric analyte comprises a polypeptide. In some embodiments, the polymeric analyte is coupled to a substrate.

[0324] In some embodiments, the method comprises contacting the polymeric analyte with the sequencing reagent, wherein the sequencing reagent binds to a monomer of the polymeric analyte thereby generating a sequencing reagent-monomer complex. In some embodiments, the monomer comprises a terminal amino acid residue.

[0325] In some embodiments, the method comprises cleaving the sequencing reagent- monomer complex from the polymeric analyte, thereby providing a cleaved sequencing reagent- monomer complex. In some embodiments, cleaving the sequencing reagent-monomer complex from the polymeric analyte is performed chemically or enzymatically. In some embodiments, cleaving the sequencing reagent-monomer complex from the polymeric analyte is performed chemically in presence of a base.

[0326] In some embodiments, the method further comprises coupling the sequencing reagent or the sequencing reagent-monomer complex to a capture moiety. In some embodiments, the capture moiety comprises a DNA molecule. In some embodiments, the capture moiety or the sequencing reagent comprises a modified DNA molecule. In some embodiments, the modified DNA molecule comprises a click chemistry moiety-modified DNA molecule. In some embodiments, the sequencing reagent is coupled to the modified DNA molecule via click chemistry.

[0327] In some embodiments, the method further comprises detecting the cleaved sequencing reagent-monomer complex or derivative thereof, wherein detecting comprises contacting the sequencing reagent-monomer complex or derivative thereof with a binding agent. In some embodiments, the binding agent comprises an antibody, nanobody, single chain variable fragment (scFv), or aptamer. In some embodiments, the binding agent comprises a polymerizable molecule. In some embodiments, the polymerizable molecule comprises a nucleic acid molecule.

[0328] In some embodiments, the method further comprises deprotecting the protecting groups (e.g., R1 and / or R2) of the sequencing reagent. In some embodiments, the deprotecting is performed before cleaving the sequence reagent-monomer complex from the polymeric analyte. In some embodiments, the deprotecting is performed before coupling the sequencing reagent to the capture moiety. In some embodiments, the deprotecting is performed in the presence of a base.

[0329] In some embodiments, the method further comprises contacting the polymeric analyte with an additional sequencing reagent, wherein the additional sequencing reagent binds to an additional monomer of the polymeric analyte, thereby generating an additional sequencing reagent-monomer complex; coupling the additional sequencing reagent to the cleaved sequencing reagent-monomer complex; and cleaving the additional sequencing reagent- monomer complex from the polymeric analyte, thereby providing a stacked sequencing reagent- monomer complex.

[0330] In some embodiments, the method further comprises detecting the cleaved sequencing reagent-monomer complex or derivative thereof. In some embodiments, detecting is performed using a nanopore. In some embodiments, detecting is performed by measuring a signal from the nanopore as the cleaved sequencing reagent-monomer complex or derivative thereof translocates through the nanopore, thereby generating a measured signal, and using the measured signal to identify the monomer.

[0331] Provided herein are methods, systems, compositions, and kits for characterizing polymeric analytes, such as proteins. The methods, systems, compositions, and kits of the present disclosure provide for the analysis of individual or clusters of monomers of a polymeric analyte, e.g., amino acids of a protein, thereby providing information on the identity or sequence of monomers (e.g., amino acids in a protein, also referred to herein as “protein sequencing”). Alternatively, or in addition to, the present disclosure may provide for methods of identifying a polymeric analyte without determining the identity of the individual monomers. In some aspects, a method of the present disclosure comprises providing a polymeric analyte, such as a peptide, and determining the identity of the individual (or clusters of) monomers comprised by the polymeric analyte. The present disclosure also provides for approaches for processing and analyzing polymeric analytes, e.g., peptides, polymers, nucleic acid molecules, etc. in a highly- parallelized and accurate manner. Systems and methods of the present disclosure may comprise cleaving a monomer from the polymeric analyte, coupling (e.g., via local tethering) of the monomer to a capture moiety, and detecting the monomer.

[0332] In some aspects, the methods, systems, compositions, and kits provided herein entail coupling polymerizable molecules to monomers (e.g., amino acids), thereby generating modified monomers (e.g, modified amino acids), and detecting the modified monomers (e.g, modified amino acids) or derivative thereof. In some embodiments, the methods provided herein entail repeating one or more operations on a single or plurality of peptides to generate a plurality of modified monomers, which may be discrete or which may be tethered together (e.g., in a stacked plurality of modified monomers), and analyzing or detecting the plurality of modifiedmonomers. In some embodiments, the detecting is performed using nanoscale objects (e.g., nanopores or nanogaps), imaging (e.g., fluorescence imaging), spectroscopy or spectrometry (e.g., mass spectrometry), nucleic acid sequencing, or a combination thereof.

[0333] In one aspect, provided herein is a method for sequencing a peptide comprising a plurality of amino acids, the method comprising (a) providing a plurality of modified amino acids generated from at least a subset of the plurality of amino acids, and (b) determining an amino acid identity (e.g., an amino acid type) of each modified amino acid of the plurality of modified amino acids. The methods provided herein may advantageously provide for highly- accurate identification and sequencing of individual amino acids from a peptide; for instance, sequencing of the plurality of modified amino acids may have an average read accuracy that is greater than 80% for at least 2 different modified amino acid types. The plurality of modified amino acids may be derived from two or more contiguous or non-contiguous amino acids of a peptide. In some embodiments, the modified amino acids herein comprise polymerizable molecules.

[0334] In some embodiments, the methods provided herein comprise performing an intramolecular expansion of a polymeric analyte, e.g., a peptide, to generate a modified monomer (e.g., modified amino acid). Such an intramolecular expansion process may comprise providing a sequencing reagent, coupling the sequencing reagent to a monomer of the polymeric analyte to generate a monomer-sequencing reagent complex, coupling the sequencing reagent or the monomer-sequencing reagent complex to a capture moiety, and cleaving the monomer from the polymeric analyte, thereby yielding a modified monomer. In some instances, the method may further comprise repeating the intramolecular expansion or one or more operations of the intramolecular expansion on the next monomer of the polymeric analyte to generate another modified monomer. In some instances, the monomer-sequencing reagent complex or the modified monomer resulting from a round or cycle of intramolecular expansion may be coupled to that of the previous round or cycle of intramolecular expansion to generate a stacked plurality of modified monomers, e.g., linked by a polymerizable molecule backbone. The modified monomer or stacked plurality of modified monomers is detected using a suitable approach to output the identity of the modified monomer (e.g., an amino acid type of a modified amino acid). The detection may be performed using imaging, nanopore sequencing, spectrometry or spectroscopy, or other suitable detection technique.

[0335] Modified amino acids: In some aspects, the modified monomer is a modified amino acid. The modified amino acid or derivative thereof may originate from or be part of a protein orpeptide; for example, the modified amino acid may comprise or be derived from an amino acid located at a terminus (N-terminus or C-terminus) of a peptide, or the modified amino acid may comprise or be derived from an amino acid located within the peptide. The modified amino acid may comprise a proteinogenic amino acid or derivative thereof with any number of modifications. Examples of modifications include, in non-limiting examples, chemical modifications (e.g., protecting groups), biological modifications (e.g., post-translational modifications, modifications introduced by enzymatic treatment or digestion), physical modifications (e.g., mutations introduced by irradiation, heat, etc.), and the like. In some instances, the modified amino acid or derivative thereof comprises or is coupled to a binding agent, such as an antibody, antibody fragment, nanobody, aptamer, peptide, a small molecule, an inorganic compound, a polymer, or any variations or combinations thereof. The modified amino acid may comprise a covalent or noncovalent modification. In some instances, the modified amino acid comprises a non-naturally occurring chemical modification. For example, the modified amino acid may comprise a protecting group, such as, in non-limiting examples, a methyl, formyl, ethyl, acetyl, t-butyl, anisyl, benzyl, tifluoroacetyl, N-hydroxysuccinimide, t- butyloxycarbonyl, benzoyl, 4-methyl benzyl, thioanizyl, thiocresyl, benzyloxymethyl, 4- nitrophenyl, benzyloxycarbonyl, 2-nitrobenzoyl, 2-nitrophenylsulphenyl, 4-toluenesulphonyl, pentafluorophenyl, diphenylmethyl, 2-chlorobenzyloxycarbonyl, 2,4,5-trichlorophenyl, 2- bromobenzyloxycarbonyl, 9-fluorenylmethyloxycarbonyl, triphenylmethyl, or 2,2, 5,7,8- pentamethyl-chroman-6-sulphonyl group. In some instances, the modified amino acid comprises a proteinogenic amino acid or derivative thereof that is coupled to a polymerizable molecule. In some instances, the modified amino acid comprises a proteinogenic amino acid or derivative thereof, a sequencing reagent, and a polymerizable molecule.

[0336] The modified amino acid may comprise any useful modification. Modifications may be naturally-occurring e.g., post translational modifications) or non-naturally occurring, such as by labeling or tagging, e.g., with an amino acid- or amine-reactive agent or linker comprising the amino acid- or amine-reactive agent. Examples of amino acid- or amine-reactive agents include isothiocyanate (e.g., PITC, NITC), l-fluoro-2,-4-dinitrobenzene (DNFB), dansyl chloride, 4- sulfonyl-2-nitrobfluorobenzene (SNFB), an acetylating agent, an acylating agent, an alkylating agent, a guanidination agent, a thioacetylation agent, a thioacylation agent, a thiobenzoylation agent, or a derivative or combination thereof. Alternatively, or in addition to, the one or more modified amino acids may comprise an adduct e.g., a polymer such as PEG, a polymerizable molecule such as a nucleic acid molecule, a nanoparticle or nanotube, a peptide or protein), a lipid, a carbohydrate, a metabolite, a fluorophore, a hapten, a quencher, a tag e.g., a fluorescenttag, a magnetic tag, a radioactive tag), a barcode, or other moiety. In some instances, a modified amino acid may comprise a modification that facilitates recruitment of an enzyme (or ribozyme or DNAzyme) to recognize or cleave a terminal amino acid, e.g., a NTAA or CTAA of a peptide. For example, a terminal amino acid of a peptide may be modified with a saccharide in order to recruit a lectin or lectin-bound protease. In another example, one or more modified amino acids may comprise or be coupled to a nucleic acid molecule having a first sequence that is complementary to a second sequence comprised by an oligo-bound protease. Hybridization of the first sequence to the second sequence may facilitate local recruitment of the protease to the amino acid to be cleaved. In yet another example, a peptide may be modified with phenylisothiocyanate (PITC), which may allow for recruitment and cleavage of the modified amino acid by an Edmanase. In other examples, a peptide may be modified with functional moiety that is recognized by a specific cleaving enzyme; recognition and binding of the cleaving enzyme to the functional moiety may result in cleavage of the modified amino acid. In some examples, modifications to amino acids may include epitope tags, which can facilitate binding of a binding agent to the modified amino acid. Examples of such epitope tags include fluorophores, nucleic acid molecules, peptides, haptens, polymers, chemical moieties, or other adduct molecule.

[0337] The methods described herein may further comprise generating the modified amino acid from a peptide, such a as a sample peptide or peptide analyte. In some instances, the modified amino acid comprises a proteinogenic amino acid or derivative thereof, a sequencing reagent, and a polymerizable molecule. In one example, an amino acid of a peptide may be contacted with a sequencing reagent that comprises (i) first reactive moiety capable of reacting with the amino acid and (ii) a second reactive moiety. Prior to, during, or subsequent to the reaction of the first reactive moiety with the amino acid, a polymerizable molecule comprising a third reactive moiety that is capable of reacting with the second reactive moiety may be provided. The second and third reactive moieties may comprise click chemistry moieties that can react with one another (e.g., azide and DBCO, azide and BCN, alkyne and DBCO, TCO and tetrazine, etc.). The reaction of the amino acid with the sequencing reagent and the sequencing reagent to the polymerizable molecule may thus yield a modified amino acid comprising the amino acid, the sequencing reagent, and the polymerizable molecule. Alternatively, or in addition to, the polymerizable molecule may comprise the amino acid- reactive moiety (optionally via a linker) and may react directly with the amino acid. In some instances, cleavage of the amino acid from the peptide may be performed, and the modified amino acid may comprise the cleaved product comprising the cleaved, and optionally derivatized, amino acid,the linker (if present), and the polymerizable molecule. In some instances, the amino acid is a terminal amino acid.

[0338] Alternatively, or in addition to, generating the modified amino acid may comprise (i) contacting an amino acid of a peptide with a polymerizable molecule comprising an amino acid reactive group and (ii) polymerizing the polymerizable molecule, thereby generating the modified amino acid. In one such example, the polymerizable molecule may comprise a modified nucleotide comprising an amino acid reactive moiety (e.g., isothiocyanate, guanidinylating group, dithioester, xanthate, etc.). The amino acid reactive group may react with the amino acid (e.g., a terminal amino acid). The modified nucleotide may then be subject to a nucleic acid reaction such as ligation (e.g., using ligase or chemical ligation such as click chemistry) or an extension reaction, e.g., using polymerase or terminal deoxynucleotidyl transferase (TdT), thereby generating the modified amino acid. See, e.g., FIG. IB (inset).

[0339] A modified amino acid may comprise a single amino acid or a plurality of amino acids (e.g., dipeptide, tripeptide, etc.). For example, the modified amino acid may comprise 1, 2, 3, 4, 5, 6, 7, 8, 9, 10 or greater amino acids. In instances where the modified amino acid comprises more than one amino acid, the plurality of amino acids comprised by the modified amino acid may be of the same amino acid type (e.g., leu-leu, val-val, ile-ile) or different amino acid types (e.g., val-pro, arg-his, gly-leu). The plurality of amino acids comprised by the modified amino acid may comprise any number of modifications, e.g., as described elsewhere herein.

[0340] Detection or Identification: In some embodiments, the modified monomers (e.g., modified amino acids) may be detected or identified, e.g., determining an amino acid type of one or more of the modified amino acids. Detection may be performed using any useful technique, such as imaging, using a nanopore sensor or sequencer, mass spectrometry, or other detection technique.

[0341] Nanopore sensing / sequencing: In some embodiments, the modified monomers (e.g., a modified amino acid or derivative thereof) may be subjected to nanopore sensing or sequencing to characterize and / or determine the identity (e.g., an amino acid type) of the modified amino acid and optionally, of the polymerizable molecule. The nanopore sensing or sequencing may be performed using a commercially available nanopore system, e.g., Oxford Nanopore Technologies, Genia Technologies, NobleGen, Northern Nanopore Instruments, Norcada, or Quantum Biosystem. Nanopore sequencing may be performed to determine theidentity of different components of the modified amino acids; for example, a modified amino acid comprising a derivatized amino acid (e.g., a thiocyanate-conjugated amino acid or a thiocarbamyl, thiazolinone, or thiohydantoin derivative) coupled to a polymerizable molecule (e.g., a nucleic acid molecule) may be subjected to nanopore sequencing, which may output the identity of which amino acid type (e.g., which of the 20 proteinogenic amino acids or post- translationally modified amino acids) the modified amino acid comprises or is derived from or a subset of types (e.g., a hydrophobic residue, a charged residue, etc.). Optionally, the identity of individual monomers of the polymerizable molecule (e.g., the nucleic acid sequence of the nucleic acid molecule) may also be identified using nanopore sequencing. For example, the modified amino acids or derivatives thereof may be translocated through or adjacent to a nanopore. As the modified amino acids or derivatives thereof translocate through or adjacent to the nanopore, one or more signals (e.g., current signal, current blockage, impedance, inductance, etc.) generated from both the polymerizable molecule and the amino acid may be generated and collected. Such one or more signals may then be deconvolved using computational approaches to determine the identity of the polymerizable molecules (e.g., nucleic acid sequences) and the amino acid types of the modified amino acids or derivatives thereof. Beneficially, the use of nanopore or nanogap sequencing for generation of multiplexed data may obviate the need for multiple analysis techniques or instruments. In instances where the polymerizable molecule comprises a nucleic acid molecule that encodes additional information (e.g., comprises barcode sequences, UMIs, cycle information, spatial information etc.), the sequencing of both the derivatized amino acid and the nucleic acid molecules may generate multiplexed information.

[0342] The one or more signals output from the nanopore may be measured and used to determine an identity (e.g., the amino acid type) of a modified amino acid. Any useful measurement can be made, e.g, a conductance, current, current blockage, current density, current change, voltage, impedance, resistance, inductance, capacitance, frequency, phase, power, electric field, magnetic field, or other parameter within or adjacent to the nanopore. The measurement may be a function of time. In some instances, the current signal may be an ionic current signal, a cross-pore or transverse-to-pore drain current, or a source current. As a molecule (e.g, modified monomer or modified amino acid) enters the nanopore, nanogap, or nanochannel, a change in the conductance, current, impedance, or other parameter may occur and provide information (e.g., size, charge, aspect ratio, volume, hydrophobicity, chemical structure) on the molecule. Each amino acid or modified amino acid, or a subset of amino acids or modified amino acids, may generate a unique signal that is distinguishable from other amino acids or modified amino acids. As such, the unique signal signatures may be assigned to theamino acids or modified amino acids in order to determine the identity of the amino acids, polymerizable molecules, or both.

[0343] Additional example methods and systems of nanopore sequencing of polymeric analytes are described in US Pat. Nos. 12,259,393 and 12,259,394 and International Patent App. No. PCT / US2025 / 013852, filed on January 20, 2025, each of which is incorporated by reference herein in its entirety.

[0344] FIGs. 1A-1D schematically show example workflows for generating a modified monomer from a polymeric analyte (e.g., a modified amino acid from a peptide) either on a substrate (FIG. 1A) or with or without a substrate (FIGS. 1B-1D). In workflow 100a of FIG. 1A, a polymeric analyte 103 (e.g., a peptide) and a capture moiety 105 are provided, which optionally are coupled to a substrate 101, such as a bead, a flow cell, a particle, etc. The capture moiety 105 may comprise a first nucleic acid molecule (e.g., DNA). In process 106, a sequencing reagent 109 and a polymerizable molecule, e.g., a linking nucleic acid molecule 111, are provided. In some instances, the sequencing reagent 109 is pre-tethered to the polymerizable molecule (depicted as a linking nucleic acid molecule 111); alternatively, the sequencing reagent 109 and the polymerizable molecule may be provided separately and joined at any useful or convenient step. In process 106, the sequencing reagent 109 may couple to a monomer, e.g., an amino acid (e.g., NTAA) of the polymeric analyte 103 (e.g., a peptide) to generate a monomer- sequencing reagent complex. In process 112, the monomer-sequencing reagent complex may couple to the capture moiety 105, thereby generating a monomer-capture moiety complex. Coupling of the monomer-sequencing reagent complex to the capture moiety 105 may be mediated by the polymerizable molecule, e.g., the linking nucleic acid molecule 111. Optionally, the monomer-sequencing reagent complex and the capture moiety 105 may be covalently linked together (e.g., the linking nucleic acid molecule 111 may be covalently linked to the capture moiety 105), using chemical (e.g., click chemistry) or enzymatic (e.g., a ligase, polymerase) approaches. Alternatively, or in addition to, the polymerizable molecule may comprise a first sequence that is complementary and may hybridize to a second sequence of the capture moiety 105 (not shown), or the polymerizable molecule may be linked to the capture moiety 105 via a splint or bridge molecule, which may comprise sequences that are complementary to the first sequence of the polymerizable molecule and the second sequence of the capture moiety 105 (not shown). In process 113, the monomer may be cleaved from the polymeric analyte 103, thereby providing a modified monomer that comprises the cleaved monomer-capture moiety complex; the modified monomer may comprise the cleaved monomer coupled to the sequencing reagent 109, the polymerizable molecule (shown as a linking nucleic acid molecule 111), the capturemoiety 105, or a combination thereof (e.g., the modified monomer may comprise the cleaved monomer, the sequencing reagent, and the polymerizable molecule).

[0345] Any of the operations, e.g., 106, 112, and 113 may be iterated and repeated any number of times (“rounds”) using additional sequencing reagents 109 and polymerizable molecules (e.g., linking nucleic acid molecules optionally comprising cycle / round information), and tethering the additional polymerizable molecules together (e.g., tethering an additional polymerizable molecule to the polymerizable molecule of the monomer-capture moiety complex). The polymerizable molecules may be coupled using any useful approach, e.g., chemical or enzymatic ligation, linkers, polymerization, or other approach. For instance, the polymerizable molecules may comprise nucleic acid molecules, which may be coupled together using a nucleic acid reaction, e.g., a ligation (e.g., chemical ligation using click chemistry, enzymatic ligation using a ligase), or extension reaction (e.g., using a polymerase). Multiple rounds may be performed until all or a subset of the monomers in the polymeric analyte 103 are cleaved and tethered together. In some instances, processes 106, 112, and 113 may be iterated to generate a stacked plurality of modified amino acids 123 comprising a set of cleaved monomers, e.g., a concatenated set of modified monomers that each comprise a polymerizable molecule coupled thereto. The stacked plurality of modified amino acids 123 may comprise a stacked set of polymerizable molecules (e.g., via coupling of the polymerizable molecule from a second round to the polymerizable molecule from a first round and coupling of the polymerizable molecule from the third round to that of the second round, and so on) from the individual modified monomers. The stacked plurality of modified amino acids 123 may comprise polymerizable molecules that are coupled or concatenated in a linear fashion (e.g., polymerizable molecules that form a linear polymerizable molecule “backbone”), or in a nonlinear fashion (e.g., random or semi-random arrangement, branched coupling). The stacked plurality of modified amino acids 123 may be circularized; for example, the stacked plurality of modified amino acids 123 forming a linear polymerizable backbone may be joined at the ends to form a circularized product. In some instances, the stacked plurality of amino acids 123 may be cleaved from the substrate. For example, the capture moiety 105 or the polymerizable molecule (e.g., linking nucleic acid molecule 111) may comprise a restriction or cleavable site that can be cleaved upon addition of the proper cleaving reagent, e.g., a restriction enzyme, a reducing agent (for disulfide bonds), etc. Alternatively, or in addition to, the stacked set of polymerizable molecules may be subject to amplification to generate amplicons of the polymerizable molecules coupled to or comprised by the stacked plurality of modified amino acids. The stacked plurality of modified amino acids 123 may be subjected to analysis or characterization, e.g., by contactingthe stacked plurality of modified amino acids 123 with a library of binding agents, which can bind to their respective monomer targets (e.g., an amino acid type), and detecting the binding agents, or by using a nanopore sequencer.

[0346] FIG. IB schematically illustrates another example workflow of generating a modified monomer, e.g., modified amino acid, in presence or absence of a substrate. In such an example workflow 100b, a polymeric analyte 103, e.g., a peptide, and a capture moiety 105 are provided. The capture moiety 105 may comprise a first nucleic acid molecule (e.g., DNA molecule) and may comprise identifying information of the polymeric analyte 103, e.g., an identifying barcode sequence which may identify the polymeric analyte or an originating sample, partition, or other useful identifying characteristics. The capture moiety 105 may additionally comprise a releasable or cleavable moiety. The polymeric analyte and the capture moiety 105 may, in some instances, be coupled to a substrate (FIG. IB inset), e.g., via an anchor molecule (e.g., anchor nucleic acid molecule). The polymeric analyte may be coupled to the capture moiety, either directly or indirectly (e.g., both the polymeric analyte and the capture moiety may be coupled to a substrate). For instance, the capture moiety 105 may comprise a nucleic acid sequence that is complementary to an anchor sequence on a substrate (e.g., bead, flat surface), or the capture moiety may be coupled to the substrate via a linker (not shown). Alternatively, the capture moiety 105 may comprise a nucleic acid sequence that is partially complementary to a splint oligonucleotide that is also partially complementary to an anchor oligonucleotide on the substrate, and proximity ligation (e.g., using ligase or chemical ligation) may be performed to couple the capture moiety 105 to the substrate (not shown). In process 106, a sequencing reagent 109 and polymerizable molecule, such as a linking nucleic acid molecule 111, are provided. In some instances, the sequencing reagent 109 is pre-tethered to the polymerizable molecule (linking nucleic acid molecule 111); alternatively, the sequencing reagent 109 and the polymerizable molecule (linking nucleic acid molecule 111) may be provided separately. The polymerizable molecule may comprise identifying temporal information, e.g., the cycle or round in which it is provided. In process 106, the sequencing reagent 109 may couple to a monomer, e.g., an amino acid (e.g., NTAA) of the polymeric analyte 103 (e.g., peptide) to generate a monomer-sequencing reagent complex. In process 112, the monomer-sequencing reagent complex may couple to the capture moiety 105. Coupling of the monomer-sequencing reagent complex to the capture moiety 105 may be mediated by the polymerizable molecule and optionally an additional polymerizable molecule 116. In some instances, the additional polymerizable molecule 116 may comprise a cleavable tag moiety, which can enable rapid and precise purification. For example, the additional polymerizablemolecule 116 may comprise a photocleavable biotin moiety, which may allow for purification of the monomer-sequencing reagent complex using streptavidin pulldown. Subsequently, the photocleavable biotin moiety may be cleaved and removed at any convenient or useful step. Alternatively, the capture moiety 105 may be directly hybridized or ligated to the polymerizable molecule (linking nucleic acid molecule 111). Optionally, the monomer-sequencing reagent complex and the capture moiety may be covalently linked together (e.g., via ligation or an extension reaction). Alternatively, or in addition to, the polymerizable molecule may comprise a first sequence that is complementary and may hybridize to a second sequence of the capture moiety 105 (not shown), or the polymerizable molecule may be linked to the capture moiety 105 via a splint or bridge molecule, which may comprise sequences that are complementary to the first sequence of the polymerizable molecule and the second sequence of the capture moiety 105 (not shown). In process 113, the monomer may be cleaved from the polymeric analyte 103 to generate the modified monomer that comprises the monomer-capture moiety complex. The modified monomer may comprise the cleaved monomer, the sequencing reagent 109, the polymerizable molecule (e.g., linking nucleic acid molecule 111), the capture moiety 105, or a combination thereof (e.g., the cleaved monomer, the sequencing reagent, and the polymerizable molecule, just the cleaved monomer, or just the cleaved monomer-sequencing reagent complex). Processes 106, 112, and 113 may be iterated and repeated any number of times (“rounds”) using additional sequencing reagents 109 and polymerizable molecules, and tethering the additional polymerizable molecules together (e.g., tethering an additional (e.g., second) polymerizable molecule to the monomer-capture moiety complex, tethering a polymerizable molecule from the third round to that of the second round, and so on). Multiple rounds may continue until all or a subset of the monomers of the polymeric analyte 103 are tethered together. For example, the process may be iterated to generate a stacked plurality of modified monomers (e.g., stacked plurality of amino acids 123) comprising a set of concatenated modified monomers, e.g., concatenated monomer-sequencing reagent-polymerizable complexes. The polymerizable molecules (e.g., linking nucleic acid molecules 111) of the stacked plurality of modified monomers (e.g., stacked plurality of amino acids 123) may be identical molecules (e.g., same nucleic acid sequence), or they may be different. The stacked plurality of modified amino acids 123 may comprise polymerizable molecules that are coupled or concatenated in a linear fashion (e.g., polymerizable molecules that form a linear polymerizable molecule “backbone”), or in a non-linear fashion (e.g., random or semi-random arrangement, branched coupling). The stacked plurality of modified amino acids 123 may be circularized; for example, the stacked plurality of modified amino acids 123 forming a linear polymerizable backbone may be joined at the ends toform a circularized product. After any useful number of rounds, the stacked plurality of modified monomers (e.g., stacked plurality of modified amino acids 123) may be cleaved from or at the capture moiety 105, e.g., using the cleavable moiety. Alternatively, or in addition to, the stacked set of polymerizable molecules may be subject to amplification to generate amplicons of the polymerizable molecules coupled to or comprised by the stacked plurality of modified amino acids. In instances where a substrate is used, the stacked plurality of modified amino acids may be removed from the substrate, e.g., via cleavage of the capture moiety or dehybridization (e.g., of a nucleic acid capture moiety that is annealed to an anchor nucleic acid molecule of the substrate). The stacked plurality of modified amino acids 123 may be subjected to analysis or characterization, e.g., by contacting the stacked plurality of modified amino acids 123 with a library of binding agents, which can bind to their respective monomer targets (e.g., an amino acid type), and detecting the binding agents or using a nanopore sequencer.

[0347] Alternatively, or in addition to, generating the modified monomer (e.g., modified amino acid) may comprise providing a polymerizable molecule comprising a reactive moiety, reacting the reactive moiety with the monomer, and polymerizing the polymerizable molecule. For example, referring to FIG. IB (inset), in process 106’, a polymerizable molecule comprising an incorporable monomer, e.g., a modified nucleotide, that has an amino acid reactive moiety (e.g., isothiocyanate, such as PITC) is provided and reacted with the modified amino acid. In process 107, a linking polymerizable molecule 111 (e.g., linking nucleic acid molecule) is provided. The linking polymerizable molecule 111 may act as a splint molecule to hybridize the capture moiety 105 or portion thereof. The linking polymerizable molecule 111 may additionally comprise a terminated end such that it is not extendable. In process 112’, an extension reaction is performed, e.g., using polymerase or a TdT enzyme to elongate the capture moiety 105 to comprise a complementary sequence to a portion of the linking polymerizable molecule 111. The extension reaction may incorporate canonical or noncanonical nucleotides (e.g., hexaphosphate nucleotides, deaza-nucleotides), ribonucleotides, or a combination thereof. In some instances, the reaction may be catalyzed by addition of cations (e.g., manganese ions, magnesium ions, calcium ions, etc.). The linking polymerizable molecule 111 may, in some instances, comprise identifying temporal information, e.g., the cycle or round in which it is provided, which may be copied to the extended nucleic acid molecule. In some embodiments, the modified nucleotide may be ligated to the capture moiety 105. The remaining operations of workflow 100b, e.g., process 113, iteration, etc. may be performed.

[0348] FIG. 1C schematically illustrates another example workflow of generating a modified monomer, e.g., modified amino acid, or a detectable product comprising the modifiedmonomer, in presence or absence of a substrate. In such an example workflow 100c, similar to that of 100b, a polymeric analyte 103, e.g., a peptide, and a capture moiety 105 are provided. The capture moiety 105 may comprise a first nucleic acid molecule (e.g., DNA molecule) and may comprise identifying information of the polymeric analyte 103, e.g., an identifying barcode sequence such as a peptide-identifying barcode, a sample barcode, a partition barcode. The capture moiety 105 may additionally comprise a releasable or cleavable moiety. The polymeric analyte and the capture moiety 105 may, in some instances, be coupled to a substrate (FIG. 1C inset). The polymeric analyte may be coupled to the capture moiety, either directly or indirectly (e.g., both the polymeric analyte and the capture moiety may be coupled to a substrate). For instance, the capture moiety 105 may comprise a nucleic acid sequence that is complementary to a sequence on a substrate (e.g., bead, flat surface), or the capture moiety may be coupled to the substrate via a linker (not shown). In process 106, a sequencing reagent 109 and polymerizable molecule, such as a linking nucleic acid molecule 111, are provided. In some instances, the sequencing reagent 109 is pre-tethered to the polymerizable molecule (linking nucleic acid molecule 111); alternatively, the sequencing reagent 109 and the polymerizable molecule (linking nucleic acid molecule 111) may be provided separately. The polymerizable molecule may comprise identifying temporal information, e.g., the cycle or round in which it is provided. In process 106, the sequencing reagent 109 may couple to a monomer, e.g., an amino acid (e.g., NTAA) of the polymeric analyte 103 (e.g., peptide) to generate a monomer-sequencing reagent complex. In process 112, the monomer-sequencing reagent complex may couple to the capture moiety 105, e.g., a portion of the capture moiety 105 may hybridize to the polymerizable molecule (linking nucleic acid molecule 111). In some instances, a nucleic acid extension reaction, e.g., using a polymerase, may be performed, e.g., to copy a sequence (e.g., a barcode sequence) of the capture moiety 105 to the polymerizable molecule (linking nucleic acid molecule 111). In process 113, the monomer may be cleaved from the polymeric analyte 103 to generate the modified monomer that comprises the cleaved monomer, the sequencing reagent 109, the polymerizable molecule (e.g., linking nucleic acid molecule 111), the capture moiety 105, or a combination thereof (e.g., the cleaved monomer, the sequencing reagent, and the polymerizable molecule, just the cleaved monomer, or just the cleaved monomer-sequencing reagent complex). In some instances, the cleavage of the monomer results in formation of a detectable product 131, which comprises a complementary sequence 105’ to a portion of the capture moiety 105, the linking nucleic acid molecule 111, the sequencing reagent 109, and the cleaved monomer (e.g., cleaved amino acid). The detectable product 131 may subsequently bedehybridized or removed (e.g., via cleavage) from the capture moiety 105, thereby regenerating the capture moiety 105, which can be used in subsequent iterations.

[0349] Processes 106, 112, and 113 may be iterated and repeated any number of times (“rounds”) using additional sequencing reagents 109 and polymerizable molecules and tethering the additional polymerizable molecules to the capture moiety 105 (e.g., via ligation, hybridization, extension), cleaving the additional monomers, removing the detectable product, and repeating. Multiple rounds may continue until all or a subset of the monomers of the polymeric analyte 103 are generated into discrete detectable products. For example, the process may be iterated to generate a plurality of detectable products 131 that each comprise a modified monomer. After any useful number of rounds, the detectable products may be collected, optionally processed, and analyzed, e.g., using a nanopore sequencer. In some instances, the discrete detectable products may be combined or concatemerized, e.g., hybridized or ligated to one another, to generate a stacked plurality of modified amino acids. The stacked plurality of modified amino acids 123 may comprise polymerizable molecules that are coupled or concatenated in a linear fashion (e.g., polymerizable molecules that form a linear polymerizable molecule “backbone”), or in a non-linear fashion (e.g., random or semi-random arrangement, branched coupling, circularized molecule). The stacked plurality of modified amino acids 123 may be circularized; for example, the stacked plurality of modified amino acids 123 forming a linear polymerizable backbone may be joined at the ends to form a circularized product.

[0350] In some instances, a branched polymerizable molecule may be used or constructed to generate the modified monomer. FIG. ID schematically illustrates another example workflow of generating a modified monomer, e.g., modified amino acid using a branched linking polymerizable molecule. In such an example workflow lOOd, similar to that of 100b and 100c, a polymeric analyte 103, e.g., a peptide, and a capture moiety 105 are provided. The capture moiety 105 may comprise a first nucleic acid molecule (e.g., DNA molecule) and may comprise identifying information of the polymeric analyte 103, e.g., an identifying barcode sequence. The capture moiety 105 may additionally comprise a releasable or cleavable moiety. The polymeric analyte and the capture moiety 105 may, in some instances, be coupled to a substrate (see, e.g., FIG. 1C inset) The polymeric analyte may be coupled to the capture moiety, either directly or indirectly (e.g., both the polymeric analyte and the capture moiety may be coupled to a substrate). For instance, the capture moiety 105 may comprise a nucleic acid sequence that is complementary to a sequence on a substrate (e.g., bead, flat surface). In process 106, a sequencing reagent 109 and polymerizable molecule, such as a linking nucleic acid molecule 111, are provided. In some instances, the polymerizable molecule is a branched polymerizablemolecule (e.g., a branched DNA molecule). In one such example, the branched DNA molecule may comprise a first nucleic acid molecule that comprises a first click chemistry moiety (e.g., an ethynyl or octadiynyl nucleobase or nucleotide analogue) that can conjugate to a second nucleic acid molecule comprising a second click chemistry moiety (e.g., azide), thereby generating the branched DNA molecule.

[0351] In some instances, the sequencing reagent 109 is pre-tethered to the polymerizable molecule; alternatively, the sequencing reagent 109 and the polymerizable molecule (e.g., linking nucleic acid molecule 111) may be provided separately. The polymerizable molecule may comprise identifying temporal information, e.g., the cycle or round in which it is provided. In process 106, the sequencing reagent 109 may couple to a monomer, e.g., an amino acid (e.g., NTAA) of the polymeric analyte 103 (e.g., peptide) to generate a monomer-sequencing reagent complex. In process 112, the monomer-sequencing reagent complex may couple to the capture moiety 105, e.g., a portion of the capture moiety 105 may hybridize to the polymerizable molecule (linking nucleic acid molecule 111, hybridization not shown) or alternatively, the polymerizable molecule may be coupled to the capture moiety 105 using a splint molecule (not shown) and ligated. In another example, the coupling of the capture moiety 105 and the polymerizable molecule (linking nucleic acid molecule 111) may be performed using click chemistry. In some instances, a nucleic acid extension reaction, e.g., using a polymerase, may be performed (not shown). In some instances, a different order of operations may be performed; for instance, the coupling of the polymerizable molecule to the capture moiety 105 may occur first, followed by coupling of the sequencing reagent 109 to the polymerizable molecule.

[0352] In process 113, the monomer is cleaved from the polymeric analyte 103 to generate the modified monomer that comprises the cleaved monomer, the sequencing reagent 109, the polymerizable molecule (e.g., linking nucleic acid molecule 111), which altogether may be coupled to the capture moiety 105. Processes 106, 112, and 113 may be iterated and repeated any number of times (“rounds”) using additional sequencing reagents 109 and polymerizable molecules, and tethering the additional polymerizable molecules together (e.g., tethering a second polymerizable molecule provided in a second round to polymerizable molecule of the first round). Multiple rounds may continue until all or a subset of the monomers of the polymeric analyte 103 are tethered together. For example, the process may be iterated to generate a stacked plurality of modified monomers (e.g., stacked plurality of amino acids 123) comprising a set of concatenated modified monomers, e.g., concatenated monomer-sequencing reagent-polymerizable complexes. For example, the polymerizable molecule from a second round may be coupled to the polymerizable molecule from the first round, and the polymerizablemolecule from a third round may couple that of the second round, and so on. Accordingly, in some embodiments, the polymerizable molecules from the multiple rounds may thus form a polymerizable molecule backbone that comprises pendant, individual modified monomers. The polymerizable molecules (e.g., linking nucleic acid molecules 111) of the stacked plurality of modified monomers (e.g., stacked plurality of amino acids 123) may be identical molecules (e.g., same nucleic acid sequence), or they may be different. The stacked plurality of modified amino acids 123 may comprise polymerizable molecules that are coupled or concatenated in a linear fashion (e.g., polymerizable molecules that form a linear polymerizable molecule “backbone”), or in a non-linear fashion (e.g., random arrangement, branched coupling). The stacked plurality of modified amino acids 123 may be circularized; for example, the stacked plurality of modified amino acids 123 forming a linear polymerizable backbone may be joined at the ends to form a circularized product. After any useful number of rounds, the stacked plurality of modified monomers (e.g., stacked plurality of amino acids 123) may be cleaved at or from the capture moiety 105, e.g., using the cleavable moiety. Alternatively, or in addition to, the stacked set of polymerizable molecules may be subject to amplification to generate amplicons of the polymerizable molecules coupled to or comprised by the stacked plurality of modified amino acids. The stacked plurality of modified amino acids 123 may be subjected to analysis or characterization, e.g., by contacting the stacked plurality of modified amino acids 123 with a library of binding agents, which can bind to their respective monomer targets (e.g., an amino acid type), and detecting the binding agents, or using a nanopore sequencer.

[0353] In some instances, the polymerizable molecule, e.g., linking nucleic acid molecule 111 comprises temporal information on the cycle in which it is provided; as such, the temporal information may be used for reconstructing the sequence of the peptide or for quality control. For example, referring to FIG. IB, if a particular cycle number is missing, then it can be inferred that an amino acid is missing or was not present in the peptide, that cleavage of the amino acid did not occur, or other error. For reconstruction purposes, e.g., referring to FIG. 1C, the presence of a temporal (e.g., round or cycle number) barcode and a peptide-identifying barcode can be used to attribute a particular identified modified amino acid to the order or position (using the temporal barcode) in which it occurs in a specific peptide (using the peptide- identifying barcode).

[0354] Accordingly, in some aspects of the present disclosure, a method for processing a peptide may comprise (a) providing the peptide and a sequencing reagent, wherein the sequencing reagent is coupled to a polymerizable molecule (e.g., nucleic acid molecule); (b) coupling the sequencing reagent to the amino acid of the peptide to generate an amino acid-sequencing reagent complex; (c) optionally, coupling the polymerizable molecule to a capture moiety, e.g., via hybridization or ligation; (d) cleaving the amino acid from the peptide to yield a modified amino acid comprising the cleaved amino acid, the sequencing reagent, and the polymerizable molecule; and (e) optionally, performing an extension reaction (e.g., nucleic acid extension reaction) of the polymerizable molecule, thereby generating a detectable product comprising the modified amino acid.

[0355] It will be appreciated that the modified amino acid may comprise other additional modifications that are not depicted, such as posttranslational modifications or chemical modifications (e.g., protecting groups), described elsewhere herein. In some instances, the modified amino acid comprises a derivatized amino acid. For instance, the sequencing reagent may comprise a PITC moiety as the amino acid reactive group, and upon conjugation of PITC to an amino acid (e.g., NTAA) of a peptide under mildly basic conditions, a phenylthiocarbamoyl (PTC) derivative of the amino acid is generated. The PTC-derivatized amino acid may be treated with acid (e.g., TFA or a Lewis acid) to generate a cleaved cyclic 2-anilino-5(4)- thiazolinone (ATZ)-derivatized amino acid, leaving a new N-terminus on the remaining peptide. The ATZ- derivatized amino acid may be converted to a phenylthiohydantoin (PTH) derivative or PTC derivative.

[0356] It will also be appreciated that the polymerizable molecules and capture moieties described herein may comprise other molecule types, e.g., peptides, lipids, carbohydrates, polymers (both naturally occurring and synthetic), or a combination thereof. For instance, referring again to FIGS. 1A-1D, the capture moiety 105 and the polymerizable molecule (e.g., shown as a linking nucleic acid molecule 111) may each comprise a peptide, optionally comprising a peptide barcode sequence. In such examples, process 112 (coupling of the monomer-sequencing reagent complex to the capture moiety) may be mediated using a peptide enzyme such as sortase A. In one such example, the capture moiety (or the polymerizable molecule) may comprise an oligo-glycine or poly-glycine peptide sequence at the C-terminus and the polymerizable molecule (or capture moiety) may comprise a sortase recognition sequence (e.g., LPXTG, where X is any amino acid) at the N-terminus and optionally, an oligoglycine or poly-glycine peptide sequence at the C-terminus (to facilitate further attachment). Sortase A may then be used to catalyze the formation of a peptide bond between the capture moiety and the polymerizable molecule. Accordingly, subsequent to cleavage of the terminal monomer (e.g., terminal amino acid of a peptide analyte), and iteration of the workflow, a stacked plurality of modified monomers 123 that are connected via a peptide backbone may be generated. In some instances, the polymerizable molecule and / or capture moiety may comprise aprotecting group that can be deprotected at any useful or convenient step. Beneficially, the use of a peptide backbone may allow for performing harsher reaction conditions (e.g., traditional Edman degradation using strong acids) and for assisting in nanopore readout of the individual monomers spaced along the peptide backbone, e.g., as described by K. Motone et al. Multi-pass, single-molecule nanopore reading of long protein strands. Nature (2024), which is incorporated by reference herein in its entirety.

[0357] A “derivative” of the modified amino acid may generally refer to a molecule that is derived from the modified amino acid. A derivative may be a product of a reaction (e.g., chemical, enzymatic) or interaction of the modified amino acid with another molecule. Further processing of the modified amino acid may optionally be performed to arrive at the “derivative thereof.” For example, further chemical or enzymatic treatment, cleavage of the amino acid, an extension or amplification reaction, cleavage or removal from a substrate, physical processes such as mechanical shearing or fragmentation may be performed on the modified amino acid to obtain a derivative of the amino acid or derivative of the modified amino acid. In an example, a modified amino acid may be derivatized by chemical reaction, e.g., using maleimide to react with cysteine residues, NHS esters or isothiocyanates to react with lysine residues, etc. The derivative may result from performing a nucleic acid reaction (e.g., nucleic acid extension reaction, amplification, ligation, transposition, hybridization, dehybridization, etc.). In some instances, a modified amino acid may refer to a stacked plurality of modified amino acids, as described elsewhere herein.

[0358] Iteration: In some instances, one or more of the operations described herein may be iterated or repeated. Iteration of the operations may allow for sequential processing, analysis, or identification of the individual monomers of the polymeric analyte, which can allow for reconstruction of the entire polymeric analyte. For example, referring to FIGs. 1A-1D, the operations of the workflow 100a, 100b, 100c, and lOOd may be conducted to generate a modified amino acid sequentially for each terminal monomer (e.g., NTAA) of the polymeric analyte (e.g., peptide). In some instances, for analyzing peptides, the individual modified amino acids may then be immobilized to a substrate (e.g., the same or separate substrate as shown in FIG. 1A), linearized, optionally immobilized (e.g., at another end), and detected. Alternatively, or in addition to, the individual modified amino acids may be combined, e.g., via hybridization or ligation of the polymerizable molecules to generate the stacked plurality of modified amino acids. As such, the stacked plurality of modified amino acids may comprise a plurality of polymerizable molecules from multiple rounds or cycles. In some instances, the polymerizablemolecule of the second (or third, fourth, fifth, . . . nth) cycle may be configured to only couple to the first (or second, third, fourth, . . .n-lth) polymerizable molecule. For example, the first cycle polymerizable molecule may comprise a unique binding sequence that is absent on the capture moiety of the substrate, and to which the second cycle polymerizable molecule can bind. Accordingly, the second cycle polymerizable molecule may only bind to the first cycle polymerizable molecule and not to any of the capture moieties. In the event that an amino acid is missed during a cycle (a “null” event, e.g., not cleaved, not coupled to the sequencing reagent or polymerizable molecule, etc.), a bridging polymerizable molecule may be provided that encodes for a null event but comprises the unique binding sequence, such that subsequent rounds may continue, even if a null event occurs. Alternatively, the polymerizable molecules across cycles or rounds may comprise the same binding sequence (e.g., an adapter sequence).

[0359] In some instances, a plurality of click chemistry moieties may be used in the polymerizable molecules for different cycles. For instance, during a first cycle, the polymerizable molecule (e.g., a linking nucleic acid molecule) may comprise tetrazine, which may react with a sequencing reagent provided in the first cycle that comprises TCO. During the second cycle, the polymerizable molecule and the sequencing reagent may comprise moieties for an orthogonal click chemistry to that of the first cycle, e.g., sulfur fluoride exchange (SuFEx) click chemistry. Accordingly, temporal information may be provided by the different click chemistry moieties, in addition to or alternatively to using cycle barcodes.

[0360] In some instances, the polymerizable molecules (e.g., linking nucleic acid molecules 111) may additionally encode temporal information, e.g., the cycle or iteration number, such that the order of the individual monomers may be determined. For example, for a given peptide, the terminal amino acid may be coupled to a polymerizable molecule that comprises a barcode sequence that identifies the cycle number (e.g., cycle 1) (not shown). The information encoded by the barcode sequence may be coupled to an adjacent (additional) capture moiety (not shown). Following cleavage of the monomer from the polymeric analyte (e.g., as shown in process 113 of FIG. IB and FIG. 1C), the workflow may be repeated for the n-1 terminal amino acid, which may again be coupled to a capture moiety via a sequencing reagent and barcoded polymerizable molecule and cleaved. The barcoded polymerizable molecule may comprise the cycle number (e.g., cycle 2).

[0361] In some instances, temporal information may be provided separately. For example, prior to, during, or subsequent to coupling of a polymerizable molecule 111 to the capture moiety, a temporal barcode may be provided that can couple to the polymerizable molecule 111 or to the capture moiety 105, or a combination thereof. The temporal barcode may comprise anyuseful agent, including a nucleic acid molecule, a peptide, a lipid, a carbohydrate, an enzyme (e.g., a chromogenic or fluorogenic enzyme) or a ribozyme or DNAzyme, a fluorophore, a dye, an intercalating agent, a dideoxynucleotide, a fluorescent nucleic acid molecule or nucleotide, a radioisotope, a mass tag, or other detectable label that can indicate the time or cycle (or iteration) number in which it is provided. In some instances, the temporal barcode comprises a cycle-specific nucleic acid barcode molecule, which can couple to the polymerizable molecule 111 or to the capture moiety 105. The temporal barcode may comprise any additional useful functional sequences, e.g., primer sites, sequencing sites, restriction sites, abasic or cleavable sites, etc. In some instances, the temporal barcode may comprise an amplification site that allows for bridge amplification of the temporal barcode and optionally, the coupled polymerizable molecules, to other capture or polymerizable molecules.

[0362] Detection Using Binding Agents: In some instances, detecting comprises using monomer-specific binding agents to recognize and bind the cleaved monomers. In some embodiments, the monomer-specific binding agents are used for direct or indirect detection; for example, the binding agents may comprise a detectable label (e.g., fluorophore, mass tag, radioisotope) that can be detected directly, or the binding agent may comprise a polymerizable molecule with encoded information, which may be transferred via coupling to or copying of the encoded information onto the capture moiety or an additional polymerizable molecule. In some instances, the additional polymerizable molecule or the capture moiety is located adjacent to the polymeric analyte. Alternatively, or in addition to, the binding agents may be used to sort the cleaved monomers, e.g., into separate partitions or compartments for downstream labeling, e.g., with identifying barcode molecules. The operations may be iterated or repeated any number of times to obtain information on all or a subset of the monomers of the polymeric analyte and optionally, the sequence of the monomers relative to the polymeric analyte. The information may be read out from the polymerizable molecules using, for example, conventional nextgeneration sequencing or nanopore sequencing approaches. Beneficially, by cleaving the monomer from the polymeric analyte, the monomers are removed from the adjacent monomers, and binding agents can specifically bind to individual monomers without the influence of the adjacent surrounding monomers. As such, the methods disclosed herein enable more accurate molecular identification and polymeric sequencing, which has applications in diagnosing disease, monitoring protein dynamics or protein interactions, single-cell proteomics, developing or characterizing therapeutics, and more.

[0363] Methods of the present disclosure for processing a polymeric analyte comprising a plurality of monomers may comprise cleaving of a monomer of the polymeric analyte and coupling the monomer to a capture moiety (e.g., coupled to a substrate, the polymeric analyte, or provided in solution) for subsequent processing or analysis. In one example, a method of the present disclosure may comprise: providing (i) a polymeric analyte comprising a plurality of monomers and (ii) a capture moiety; coupling a monomer of the plurality of monomers to the capture moiety to generate a monomer-capture moiety complex; cleaving the monomer; contacting the cleaved monomer-capture moiety complex with the binding agent; and coupling a first polymerizable molecule to a second polymerizable molecule or to the capture moiety. In some instances, the first polymerizable molecule is coupled to the binding agent and comprises information about the binding agent, such as the identity of the binding agent or its cognate molecule, which information may be transferred to the second polymerizable molecule or to the capture moiety. In some instances, the capture moiety may be a copy or identical molecule as the second polymerizable molecule. Alternatively, or in addition to, the binding agent may be used to sort a mixture of cleaved monomers by identity or type, and subsequent to sorting, an identifying label or barcode that identifies the monomer type may be coupled to the capture moiety. Example methods and systems of such processing approaches and systems are described in U.S. Pat. No. 11,499,979, International Pat. Pub. No. WO / 2023 / 114732, International Pat. Pub. No. WO / 2023 / 196642, and U.S. Pat. App. No. 18 / 740,088, filed June 11, 2024, and L. Zheng et al. 2024. Peptide sequencing via reverse translation of peptides into DNA. BioRxiv., each of which is incorporated by reference herein in its entirety.

[0364] FIG. 2A shows an example workflow of analyzing a polymeric analyte. In workflow 200, a polymeric analyte 203 is provided, sequentially disassembled into individual monomers via contacting with capture moieties and cleavage, and the individual cleaved monomers are complexed with the capture moieties are contacted with binding agents comprising polymerizable molecules that identify or encode for the binding agents. The polymerizable molecule of a binding agent is coupled to an additional polymerizable molecule, and the monomer is cleaved from the capture moiety or blocked to prevent downstream recognition from additional binding agents. In FIG. 2A Panel A, a substrate 201 is coupled to a polymeric analyte 203 (e.g., a peptide to be sequenced), a capture moiety 205 (e.g., a nucleic acid molecule), and an additional polymerizable molecule 207 (e.g., an additional nucleic acid molecule). In some instances, the capture moiety 205 and the additional polymerizable molecule 207 are identical molecules (e.g., comprise the same sequence). The polymeric analyte may be contacted with a sequencing reagent 209, which comprises a terminal monomer-coupling group(e.g., an amino acid-reactive group such as a guanidinyl group) and a click chemistry moiety (e.g., azide). In some instances, the sequencing reagent 209 couples to a terminal monomer (e.g., terminal amino acid such as the N-terminal amino acid (NTAA)). In FIG. 2A Panel B, a linking nucleic acid molecule 211 comprising a click chemistry moiety (e.g., alkyne) is reacted with the sequencing reagent 109 and covalently linked. In some instances, the linking nucleic acid molecule 211 and the sequencing reagent 209 are provided pre-coupled (see, e.g, FIG. 3). In FIG. 2A Panel C, the linking nucleic acid molecule 211 is coupled to the capture moiety 205, thereby generating a monomer-capture moiety complex. The coupling may be mediated by hybridization of the linking nucleic acid molecule 211 to the capture moiety 205 (hybridization not shown), or by using a splint oligonucleotide 213 comprising sequences complementary to a sequence of the linking nucleic acid molecule 211 and the capture moiety 205. In some instances (not shown), the linking nucleic acid molecule 211 comprises a self-splinting sequence, such that the linking nucleic acid molecule 211 may couple to the capture moiety 205 in the absence of a separate splint molecule. A ligase may be used to covalently link the linking nucleic acid molecule 211 to the capture moiety 205. Alternatively, the linking nucleic acid molecule 211 may comprise a first reactive moiety (e.g., click chemistry moiety not shown) that can react with a second reactive moiety (not shown) of the capture moiety 205. In FIG. 2A Panel D, the system is subjected to conditions sufficient to cleave the terminal monomer (e.g., amino acid) from the polymeric analyte 203 (e.g., peptide), thereby generating a cleaved monomer-capture moiety complex. The conditions may include performing an Edman degradation reaction. The cleavage of the monomer from the polymeric analyte results in a cleaved monomer-capture moiety complex comprising the cleaved monomer, sequencing reagent 209, linking nucleic acid molecule 211, and capture moiety 205. In FIG. 2A Panel E, a binding agent 215 (e.g., antibody) comprising another polymerizable molecule 217 (e.g., nucleic acid molecule) is provided. The binding agent 215 may be specific or partially specific to the monomer (e.g., to an amino acid type) or to the monomer-sequencing reagent complex (e.g., guanidinyl-amino acid complex). The polymerizable molecule 217 of the binding agent may comprise information on the identity of the binding agent or the specific monomer (e.g., single amino acid) to which the binding agent binds. The polymerizable molecule 217 of the binding agent may comprise additional sequences, such as a barcode sequence, UMI, restriction site, transposition site, a sequence to represent a cycle or iteration number, or other functional sequence. The polymerizable molecule 217 of the binding agent may couple to the additional polymerizable molecule 207 that is coupled to the substrate 201. In some instances, an extension reaction may be performed (e.g., using a polymerase), to copy the sequence of the polymerizable molecule 217 of the bindingagent to the additional polymerizable molecule 207 that is coupled to the substrate 201. Alternatively, the polymerizable molecule 217 of the binding agent may be ligated to the additional polymerizable molecule 207, either chemically (e.g., via complementary click chemistry) or enzymatically (e.g., using a ligating enzyme, ribozyme or DNAzyme). Optionally, the polymerizable molecule 217 may be cleaved from the binding agent 215 (not shown), e.g., the polymerizable molecule 217 may comprise a releasable or cleavable moiety. In FIG. 2A Panel F, the monomer may be decoupled (e.g., removed or cleaved) from the capture moiety 205. For example, the monomer, sequencing reagent 209, and all or a portion of the linking nucleic acid molecule 211 may be cleaved (depicted as a star). The cleavage may be performed chemically, mechanically, or enzymatically. In an example of enzymatic cleavage, the linking nucleic acid molecule 211 may comprise a restriction site or other cleavage site (e.g., a uracil), and cleavage occurs by introduction of a restriction enzyme or cleaving enzyme (e.g., uracil DNA glycosylase) to cleave the restriction / cleavage site. Alternatively, the cleaved monomer- capture moiety complex may be blocked with a blocking agent (not shown). The workflow 200 may then be iterated or repeated to sequence all or a portion of the polymeric analyte 203.

[0365] FIG. 2B schematically shows another example of processing and characterizing a polymeric analyte 203, e.g., a peptide. A polymeric analyte 203 may be tagged (or provided pretagged) with a capture moiety 205, e.g., a polymerizable molecule such as a nucleic acid molecule. The capture moiety 205 may, for example, be attached at a terminus of a peptide (e.g., as shown at the C-terminus) or at an internal residue. The capture moiety 205 may, in some instances, comprise a barcode sequence. In some instances, the capture moiety 205 may not be attached to the peptide but may be associated to the peptide (e.g., via an indirect interaction). In other examples (not shown), the peptide is coupled (directly or indirectly) to a non-nucleic acid molecule (e.g., a polymerizable molecule such as an additional peptide, or other detectable label such as a mass tag, fluorophore, radioisotope, etc.). A sequencing reagent 209 comprising a linking nucleic acid molecule 211 is provided. The linking nucleic acid molecule 211 may comprise any useful sequences, such as a primer sequence, a barcode sequence, a UMI, restriction site, etc. In some instances, the linking nucleic acid molecule 211 comprises a cleavable moiety, e.g., a restriction site, an abasic site, a uracil, a transposition site, etc. In process 210, the sequencing reagent 209 couples to a monomer (e.g., terminal amino acid) of the polymeric analyte to generate a monomer-sequencing reagent complex. Prior to, during, or subsequent to the coupling of the sequencing reagent to the monomer, the linking nucleic acid molecule 211 may couple to the capture moiety 205, e.g., via hybridization (not shown), ligation, or splinted ligation (not shown), thereby generating a monomer-capture moietycomplex. In process 213, the monomer is cleaved (e.g., chemically or enzymatically) from the rest of the polymeric analyte 203, thereby providing a cleaved monomer-capture moiety complex comprising the cleaved monomer coupled to the sequencing reagent 209, linking nucleic acid molecule 211, and capture moiety 205. In some instances, the cleaved monomer- capture moiety complex remains coupled to the polymeric analyte 203 via the linking nucleic acid molecule 211 and the capture moiety 205. The cleaved monomer-capture moiety complex can be further processed for downstream analysis.

[0366] The downstream processing and analysis may comprise sorting, detection, or both. In some instances, a plurality of binding agents 215 is provided; the binding agents 215 may recognize and bind different monomer types (e.g., different amino acid types). The binding agents 215 may be contacted with a plurality of cleaved monomer-capture moiety complexes comprising different monomer types (e.g., different amino acids) and bind to their respective targets. In some instances (not shown), the binding agents 215 comprise polymerizable molecules, e.g., nucleic acid molecules, that identify the binding agent or its cognate molecule; the polymerizable molecules may be coupled or transferred to the capture moiety 205 (not shown), e.g., via nucleic acid extension, ligation, transposition, etc. Alternatively or in addition to, in process 219, the binding agents 215 may be separated or sorted, e.g., using complementary nucleic acid sequences to the nucleic acid molecules of the binding agents, into separate compartments (not shown). Alternatively, or in addition to, the binding agent may comprise a sorting tag, e.g., a reporter molecule, mass tag, fluorophore or fluorescent protein, which can enable sorting of the different binding agent types. In one such example, a first binding agent against a first monomer type (e.g., an amino acid residue) may comprise a GFP tag and a second binding agent against a second monomer type may comprise an RFP tag that is sortable by fluorescence (e.g., using FACS) or affinity sorting (e.g., using beads with anti-GFP and anti- RFP antibodies).

[0367] Optionally, subsequent to process 219, the sequencing reagent 209 or the linking nucleic acid molecule 211 or portion thereof may be removed from the monomer-sequencing reagent complex or cleaved monomer-sequencing reagent complex, e.g., via restriction digest or cleavage of a uracil (e.g. using UDG or USER enzymes) of the linking nucleic acid molecule 211.

[0368] In some instances, barcoding of the cleaved monomer-capture moiety complexes may be performed. For example, as described above, the binding agents 215 may comprise polymerizable barcode molecules, e.g., nucleic acid barcode molecules, that identify the binding agent or its cognate molecule; the polymerizable barcode molecules may be coupled ortransferred to the capture moiety 205 (not shown), thereby barcoding the capture moiety. Alternatively, or in addition to, subsequent to sorting in process 219, the sorted cleaved monomer-capture moiety complexes may be barcoded. As each compartment comprises a known monomer (e.g., amino acid) type according to the binding profile (e.g., specificity to a particular monomer type) of the binding agent, the cleaved monomer-capture moiety complex may be labeled with an identifying polymerizable molecule 217 (e.g., nucleic acid barcode molecule) that comprises the identity of the particular monomer type of the cleaved monomer- capture moiety complex, thereby generating a barcoded capture moiety.

[0369] Subsequent to barcoding, the contents of the separate compartments may then be pooled, and the process repeated to iteratively cleave and attach an identifying barcode for each monomer in the polymeric analyte. The subsequent barcode molecules may attach to the barcoded capture moiety to generate a multi -barcoded additional polymerizable molecule (e.g., a concatenated or stacked barcoded nucleic acid molecule). Alternatively, or in addition to, the barcodes may be added onto separate polymerizable molecules, e.g, amplicons (not shown) of the capture moiety 205. In some instances, the polymerizable molecules 217 or the linking nucleic acid molecule 211 may comprise temporal information (e.g, a round or cycle number). After any useful number of rounds or iterations, the barcoded additional polymerizable molecule (or multi -barcoded additional polymerizable molecule) may be removed and sequenced, e.g., using NGS approaches, thereby outputting the identity of each monomer type that has been processed and the order or position in which they occur, based on the temporal information, in the polymeric analyte.

[0370] FIG. 2C schematically shows another example workflow of sequencing a polymeric analyte. A plurality of polymeric analytes (e.g., peptides) may be coupled to a substrate 201, along with a plurality of capture moi eties. For illustration purposes, further sequencing workflow operations are shown for a single polymeric analyte 203; however, it will be appreciated that the workflow operations of FIGs. 2A-2C may be performed on all or a subset of the polymeric analytes in parallel. In process 210, a sequencing reagent 209 is provided. The sequencing reagent 209 is capable of coupling to (i) a monomer of the polymeric unit (e.g., a terminal amino acid, such as the N-terminal amino acid) and (ii) a capture moiety 205, which may be used to locally tether the monomer adjacent to the polymeric analyte. In an example, the sequencing reagent 209 may comprise an amino acid-reactive group such as guanidinylating agents, which enables the sequencing reagent to couple to an N-terminal amino acid. The sequencing reagent may additionally comprise a substrate-tethering moiety that can couple to the capture moiety. For example, the sub state-tethering moiety may comprise a click chemistrymoiety that can couple to a click chemistry capture moiety, e.g., through an azide-alkyne or azide-cycloalkyne reaction. In other examples, the substrate-tethering moiety may comprise a nucleic acid molecule that can couple, e.g., via hybridization, ligation, or both, to a nucleic acid capture moiety (e.g., as shown in FIG. 2A). In process 212, the substrate-tethering moiety of the sequencing reagent may couple to the capture moiety. In process 213, the monomer may be cleaved from the polymeric analyte. In some embodiments, cleavage of the monomer may be mediated by a stimulus, e.g., chemical reaction or pH change (such as addition of acid). Subsequent to cleavage, a binding agent 215 may be provided. The binding agent, e.g., an antibody or antibody fragment, may be specific to a particular monomer of a plurality of monomers, e.g., specific to a particular amino acid type or derivative thereof. The binding agent may comprise a detectable label (e.g., fluorophore, radioisotope, mass tag, etc.). The detectable label may be detected (e.g., using microscopy or imaging). To determine the identity of the cleaved terminal monomer (e.g., amino acid). Subsequently, the sequencing reagent may be removed or cleaved, and the process may be repeated or reiterated to sequence the remaining monomers of the polymeric analyte.

[0371] FIG. 2D shows another workflow of sequencing a polymeric analyte. A polymeric analyte 203 and a capture moiety 205 are provided, which optionally are coupled to a substrate 201. The capture moiety 205 may comprise a first nucleic acid molecule (e.g., DNA). In process 106, a sequencing reagent 207 and a polymerizable molecule, e.g., a linking nucleic acid molecule 211 are provided. In some instances, the sequencing reagent 209 is pre-tethered to the polymerizable molecule (depicted as a linking nucleic acid molecule 211); alternatively, the sequencing reagent 209 and the polymerizable molecule may be provided separately. In process 206, the sequencing reagent 209 may couple to a monomer, e.g., an amino acid (e.g., NTAA) of the polymeric analyte 203 (e.g., a peptide) to generate a monomer-sequencing reagent complex. In process 212, the monomer-sequencing reagent complex may couple to the capture moiety 205, thereby generating a monomer-capture moiety complex. Coupling of the monomer- sequencing reagent complex to the capture moiety 205 may be mediated by the polymerizable molecule, e.g., the linking nucleic acid molecule 211. Optionally, the monomer-sequencing reagent complex and the capture moiety 205 may be covalently linked together (e.g., the linking nucleic acid molecule 211 may be covalently linked to the capture moiety 205), using chemical (e.g., click chemistry) or enzymatic (e.g., a ligase) approaches. Alternatively, or in addition to, the polymerizable molecule may comprise a first sequence that is complementary and may hybridize to a second sequence of the capture moiety 205 (not shown), or the polymerizable molecule may be linked to the capture moiety 205 via a splint or bridge molecule, which maycomprise sequences that are complementary to the first sequence of the polymerizable molecule and the second sequence of the capture moiety 205 (not shown). In process 213, the monomer may be cleaved from the polymeric analyte 203, thereby providing a cleaved monomer-capture moiety complex comprising the cleaved monomer coupled to the sequencing reagent 209, polymerizable molecule (shown as a linking nucleic acid molecule 211), and capture moiety 205. In process 114, a binding agent 215 (e.g., an antibody, binding protein, etc.) may be contacted with the monomer-capture moiety complex. The binding agent may be configured to recognize all or a portion of the monomer-capture moiety complex. For example, the binding agent may recognize the monomer, the monomer-sequencing reagent complex, or the entirety of the monomer-capture moiety complex. In one example, the sequencing reagent may comprise a guanidinyl moiety, and the binding agent may recognize the guanidinyl-amino acid, or a derivative thereof, e.g., a phenylthiocarbamyl, a thiazolone, or a phenylthiohydantoin. In some instances, the binding agent 215 may comprise a detectable moiety (not shown) or may be contacted with an additional binding agent (e.g., a secondary antibody) which may optionally comprise a detectable moiety (not shown).

[0372] Any of the processes, e.g., 206, 212, 213, or 114 may be iterated and repeated any number of times (“rounds”) using additional sequencing reagents 209 and polymerizable molecules (optionally comprising cycle / round information), and tethering the additional polymerizable molecules together (e.g., tethering an additional polymerizable molecule to the polymerizable molecule of the monomer-capture moiety complex). Multiple rounds may be performed until all or a subset of the monomers in the polymeric analyte 203 are cleaved and tethered together. In some instances, processes 206, 212, and 213 may be iterated to generate a stacked polymerizable molecule 223 comprising a set of cleaved monomers, e.g., a concatenated set of monomer-sequencing reagent-polymerizable complexes. The stacked set of polymerizable molecules may then be contacted with a library of binding agents, which can bind to their respective monomer targets (e.g., an amino acid type).

[0373] Further downstream analysis may be performed, e.g., using a nanopore or nanogap system. In one such example, the stacked polymerizable molecule 223, which optionally may be coupled to binding agents, may be prepared and translocated through a nanopore sequencing system, which may output the identity of the polymerizable molecules (e.g., a nucleic acid sequence), the monomer type, and, if prevalent, the individual binding agents.

[0374] In some instances, the polymerizable molecule comprises temporal information on the cycle in which it is provided; as such, the temporal information may be used for conducting quality control. For example, if a missing cycle number is missing, then it can be inferred that anamino acid is missing or was not present in the peptide, that cleavage of the amino acid did not occur, or other error.

[0375] Iteration: In some instances, one or more of the operations described herein may be iterated or repeated. Iteration of the operations may allow for sequential processing, analysis, or identification of the individual monomers of the polymeric analyte, which can allow for reconstruction of the entire polymeric analyte. For example, referring to FIG. 2A, the operations of the workflow 200 may be conducted to encode the identity (e.g., via the polymerizable molecule 217) of a terminal amino acid (e.g., NTAA) onto the additional polymerizable molecule 207. The operations of workflow 200 may then be repeated to encode the identities of the n-1 terminal amino acid, the n-2 terminal amino acid, the n-3 terminal amino acid, etc. until the entire or portion of the peptide is processed. The encoding may occur on the same (additional) polymerizable molecule 207, e.g., to generate a stacked polymerizable molecule comprising multiple polymerizable molecules from multiple binding agents, or the encoding may occur on additional polymerizable molecules (not shown) present on the substrate. In the former situation, in some instances, the polymerizable molecule of the second (or third, fourth, fifth, . . . nth) cycle may be configured to only couple to the first (or second, third, fourth, . . .n- 1th) polymerizable molecule. For example, the first cycle binding agent polymerizable molecule may comprise a unique binding sequence that is absent on the additional polymerizable (or capture) molecules of the substrate, and to which the second cycle binding agent polymerizable molecule can bind. Accordingly, the second cycle binding agent polymerizable molecule can only bind to the first cycle binding agent polymerizable molecule and not to any of the additional polymerizable (or capture) molecules of the substrate. In the event that no binding occurs (a “null” event), a bridging polymerizable molecule may be provided that encodes for a null binding event but comprises the unique binding sequence, such that subsequent rounds may continue, even if a binding agent does not bind the cleaved monomer.

[0376] Similarly, any of the operations depicted in FIGS. 2B-2D may be iterated to sequentially analyze all or a subset of the monomers of the polymeric analyte. For instance, referring to FIG. 2B, the operations of the workflow may be conducted to encode the identity (e.g., via the polymerizable molecule 217) of a terminal monomer (e.g., NTAA) onto an additional polymerizable molecule (not shown) or onto the barcoded capture moiety. The operations may be repeated to encode the identities of the n-1 terminal amino acid, the n-2 terminal amino acid, the n-3 terminal amino acid, etc. until the entire or portion of the peptide is processed. The encoding may occur on the same capture moiety 205, e.g., to generate a stackedpolymerizable molecule comprising multiple barcode polymerizable molecules 217, or the encoding may occur on additional polymerizable molecules (not shown). In the former situation, in some instances, the polymerizable molecule of the second (or third, fourth, fifth, . . . nth) cycle may be configured to only couple to the first (or second, third, fourth, . . .n-lth) polymerizable molecule, as described above.

[0377] In instances where one or multiple additional polymerizable molecules are used, the polymerizable molecules 217 (e.g., coupled to the binding agents or provided separately after sorting) may additionally encode temporal information, e.g., the cycle or iteration number, such that the order of the individual monomers may be determined. For example, for a given peptide, the terminal amino acid may be coupled to a capture moiety and cleaved, then contacted with a binding agent comprising a barcode sequence that identifies (i) the identity of the amino acid (e.g., any one of twenty proteinogenic amino acids) and (ii) the cycle number (e.g., cycle 1) (not shown). The information encoded by the barcode sequence may be coupled to an adjacent (additional) polymerizable molecule (not shown) or to the capture moiety 205 or barcoded capture moiety. Following cleavage of the monomer from the capture moiety (e.g., as shown in process 221 of FIG. 2B), the workflow may be repeated for the n-1 terminal amino acid, which may again be coupled to a capture moiety, cleaved, and contacted with a binding agent, which may comprise an additional barcode sequence that identifies (i) the identity of the amino acid (e.g., any one of twenty proteinogenic amino acids) and (ii) the cycle number (e.g., cycle 2). The information encoded by the additional barcode sequence may be transferred to the same capture moiety 205 or barcoded capture moiety, or to an additional polymerizable molecule (not shown). In the former situation, the polymerizable molecule may then comprise information on the (i) the identity of the terminal amino acid, (ii) the cycle number of the terminal amino acid (cycle 1), (iii) the identity of the n-1 terminal amino acid, and (iv) the cycle number of the n-1 terminal amino acid (cycle 2), and so forth. Alternatively, or in addition to, the temporal information may be provided on another molecule, such as the capture moiety 205, the linking nucleic acid molecule 211, additional polymerizable molecule 207, etc.

[0378] Alternatively, or in addition to, the binding agent may comprise a sorting tag, and a barcode sequence may be provided subsequent to sorting. For example, the binding agents may be used to sort different cleaved amino acid-sequencing reagent complexes (as shown in process 219), and a polymerizable molecule 217 comprising the identity of the amino acid type may be provided for each sorted, cleaved amino acid-sequencing reagent complex. The polymerizable molecule 217 may also comprise temporal information, e.g., the round or cycle in which it is provided.

[0379] In some instances, temporal information may be provided separately. For example, prior to, during, or subsequent to coupling of a polymerizable molecule 217 comprising barcode information to either an additional polymerizable molecule 207 (FIG. 2A) or the capture moiety (FIG. 2B), a temporal barcode may be provided that can couple to the polymerizable molecule 217 (FIGs. 2A-2B), the additional polymerizable molecule 207 (FIG. 2A), the capture moiety 205 (FIGs. 2B, 2D), or a combination thereof. The temporal barcode may comprise any useful agent, including a nucleic acid molecule, a peptide, a lipid, a carbohydrate, an enzyme (e.g., a chromogenic or fluorogenic enzyme) or a ribozyme or DNAzyme, a fluorophore, a dye, an intercalating agent, a dideoxynucleotide, a fluorescent nucleic acid molecule or nucleotide, a radioisotope, a mass tag, or other detectable label that can indicate the time or cycle (or iteration) number in which it is provided. In some instances, the temporal barcode comprises a cycle-specific nucleic acid barcode molecule, which can couple to the polymerizable molecule 217 (comprising the identity of the monomer) or to a terminal polymerizable molecule of a stacked polymerizable molecule comprising polymerizable molecules from multiple rounds or iterations. The temporal barcode may comprise any additional useful functional sequences, e.g., primer sites, sequencing sites, restriction sites, abasic or cleavable sites, etc. In some instances, the temporal barcode may comprise an amplification site that allows for bridge amplification of the temporal barcode and optionally, the coupled polymerizable molecules, to other capture or polymerizable molecules.

[0380] Polymeric Analytes: The polymeric analyte described herein may be a biomolecule, macromolecule, or synthetic molecule. The polymeric analyte may be a biomolecule or other biological molecule that comprises one or more monomers. Non-limiting examples of polymeric biomolecules include nucleic acid molecules (e.g., DNA molecule, RNA molecule, DNA:RNA hybrids, aptamers), peptides and proteins, polysaccharides, lipid polymers (e.g., diglycerides, triglycerides and other fatty acids). The polymeric analyte may be a synthetic molecule, e.g., a peptoid or synthetic polymer, or a peptidomimetic (e.g., a peptoid, a beta-peptide, a D-peptide peptidomimetic). Non-limiting examples of synthetic polymers include acrylics, nylons, silicones, viscose, rayon, polyesters, polycarboxylic acids, polyvinyl acetate, polyacrylamide, polyacrylate, polyethylene glycol, polyurethane, polylactic acid, silica, polystyrene, polyacrylonitrile, polybutadiene, polycarbonate, polyethylene terephthalate, poly(chlorotrifluoroethylene), poly(ethylene oxide), polyethylene terephthalate), polyethylene, polyisobutylene, poly(methyl methacrylate), poly(oxymethylene), polyformaldehyde, polypropylene, polystyrene, poly(tetrafluoroethylene), poly(vinyl acetate), poly(vinyl alcohol),poly(vinyl chloride), poly(vinylidene dichloride), poly(vinylidene difluoride), poly(vinyl fluoride), or a combination thereof. The polymeric analytes may comprise a single polymer type (e.g., a homopolymer) or more than one polymer type (e.g., a copolymer) and may comprise random or arranged monomers. The polymeric analytes may be a block polymer, alternating copolymer, periodic copolymer, statistical copolymer, stereoblock copolymer, gradient copolymer, branched copolymer, graft copolymer, etc.

[0381] The polymeric analytes may be any size or comprise a range of sizes. The polymeric analyte may be about 1 nanometer (nm), about 5 nm, about 10 nm, about 20 nm, about 30 nm, about 40 nm, about 50 nm, about 60 nm, about 70 nm, about 80 nm, about 90 nm, about 100 nm, about 200 nm, about 300 nm, about 400 nm, about 500 nm, about 600 nm, about 700 nm, about 800 nm, about 900 nm, about 1 micrometer (gm), about 10 pm, about 100 pm, about 1 millimeter mm in size or greater. A plurality of polymeric analytes may comprise polymeric analytes of similar size or within a range of sizes, e.g., between about 10 nm to about 100 nm, between about 50 nm to about 1 pm. Similarly, the polymeric analytes may have any molecular weight or range of molecular weights. The polymeric analytes may be about 10 daltons (Da), 100 Da, 500 Da, 1 kilodalton (kDa), 10 kDa, 100 kDa, 1,000 kDa, 10,000 kDa, 100,000 kDa, or greater. The polymeric analytes may comprise polymeric analytes of similar molecular weight or within a range of molecular weights.

[0382] The monomers of the polymeric analytes may comprise any size or range of sizes that is less than that of the entire polymeric analyte. A monomer may be about 0.1 nanometer (nm), about 0.5 nm, 1 about 1 nm, about 5 nm, about 10 nm, about 20 nm, about 30 nm, about 40 nm, about 50 nm, about 60 nm, about 70 nm, about 80 nm, about 90 nm, about 100 nm, about 200 nm, about 300 nm, about 400 nm, about 500 nm, about 600 nm, about 700 nm, about 800 nm, about 900 nm, about 1 micrometer (pm), about 10 pm, about 100 pm, about 1 millimeter mm in size or greater. The monomers may have any molecular weight or range of molecular weights. The monomers may be about 1 dalton (Da), 10 Da, 100 Da, 500 Da, 1 kilodalton (kDa), 10 kDa, 100 kDa, 1,000 kDa, 10,000 kDa, 100,000 kDa, or greater. The monomers or polymeric analytes may range in size of molecular weight; for example, a polymeric analyte may comprise a peptide comprising amino acid monomers, which may vary in molecular weight from 75 Da (glycine) to 204 Da (tryptophan).

[0383] The polymeric analytes may comprise any number of monomers. The polymeric analytes may comprise about 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1,000, 2,000, 3,000, 4,000, 5,000, 6,000, 7,000, 8,000, 9,000, 10,000, 50,000, 100,000 or more monomers. The polymeric analytes may comprise at leastabout 2, at least about 3, at least about 4, at least about 5, at least about 6, at least about 7, at least about 8, at least about 9, at least about 10, at least about 20, at least about 30, at least about 40, at least about 50, at least about 60, at least about 70, at least about 80, at least about 90, at least about 100, at least about 500, at least about 1,000, at least about 5,000, at least about 10,000, at least about 50,000 at least about 100,000 or greater monomers. Alternatively, the polymeric analytes may comprise at most about 100,000, at most about 50,000, at most about 10,000, at most about 5,000, at most about 1,000, at most about 500, at most about 100, at most about 50, at most about 10, at most about 5, or fewer monomers. The polymeric analytes may comprise a range of monomers; for example, a polymeric analyte may comprise about 5 monomers whereas another polymeric analyte may comprise about 500 monomers.

[0384] In some instances, the polymeric analyte comprises a peptide comprising amino acid monomeric units. The peptide may be naturally occurring or synthetic. The peptide may comprise any number of amino acids. The amino acids may be one of 20 proteinogenic amino acids and may comprise any number of post-translational modifications. The peptides or any of the constituent amino acids may be processed, e.g., contacted with protecting groups, alkylated, beta-elimination of phosphate groups, etc., as is described elsewhere herein. In some instances, the peptides are derived from larger peptides or proteins and are fragmented.

[0385] Substrates'. One or more operations described herein may be performed using a substrate. For example, one or more molecules described herein (e.g., polymeric analyte such as a peptide, capture moiety, polymerizable molecule) may be coupled to a substrate. In some instances, the polymeric analyte, capture moiety, and one or more polymerizable molecules (e.g., the first or second polymerizable molecule), or a combination thereof may be provided coupled to one or more substrates. In one example, the polymeric analyte and a capture moiety are coupled to a substrate. The substrate may comprise one or more anchor molecules that may couple to the polymeric analyte or the capture moiety. In some instances, more than one substrate may be used. In such cases, the substrates may comprise the same material or different material.

[0386] The substrate may be made from any suitable material, e.g., glass, silicon, gel (e.g., a hydrogel including reversible hydrogels), polymer, etc., as is described elsewhere herein. In some instances, the substrate may be a bead or a gel bead (e.g., polyacrylamide, agarose, or TentaGel® bead). The substrate may comprise a flow cell, a microfluidic device, or one or more surfaces disposed thereon. The substrate may be functionalized. One or more molecules, e.g., a capture moiety and the polymeric analyte (e.g., a peptide) may be coupled to the substrate via acovalent or non-covalent interaction. The capture moiety and polymeric analyte (e.g., peptide) can be coupled to the substrate using any suitable chemistry, e.g., click chemistry moieties (e.g., alkyne-azide coupling), photoreactive groups (e.g., benzophenone), l-ethyl-3-(3- dimethylaminopropyl)carbodiimide hydrochloride (EDC) (e.g., to couple amino-oligos or peptides), N-hydroxy sulfosuccinimide (NHS), Sulfo-NHS, or NHS-esters (e.g., to couple sulfhydryl oligos), maleimides, hydrazines, hydroxyl amines, thiols, biotin-streptavidin interactions, cystamine, glutaraldehyde, formaldehyde, succinimidyl 4-(N- maleimidomethyl)cyclohexame-l -carboxylate (SMCC), Sulfo-SMCC, 4-(4,6-Dimethoxy-l,3,5- triazin-2-yl)-4-methylmorpholinium chloride (DMTMM), silane (e.g., amino silanes), combinations thereof, etc. In some instances, the substrate may be functionalized to comprise a coupling chemistry to couple the polymeric analyte or the capture moiety. In one non-limiting example, a substrate (e.g., bead or surface) may comprise an alkyne such as dibenzocyclooctyne (DBCO), which may be configured to react to an amine (e.g., DBCO-alcohol, DBCO-Boc, DBCO-NHS), a carboyxl or carbonyl (e.g., DBCO, DBCO-silane), a sulfhydryl, etc. An azide- functionalized nucleic acid or protein may react with DBCO to link the nucleic acid or protein to the DBCO substrate. In other examples, linkers such as bifunctional linkers may be used to attach a molecule to a substrate; such bifunctional linkers may comprise the same reactive moiety on both ends or a different moiety at each end (e.g., heterobifunctional linker). Additional examples of sequencing reagents are described elsewhere herein.

[0387] In some instances, a molecule (e.g., polymeric analytes such as peptides, capture moieties, polymerizable molecules) may be coupled to the substrate or an anchor molecule of the substrate using an enzymatic approach, e.g., as described elsewhere herein. For example, a chemical linker or moiety such as a click chemistry moiety may be attached to a polymeric analyte e.g., peptide) using an enzyme. The chemical linker or moiety may be able to react with another chemical linker or moiety e.g., click chemistry moiety) of a substrate, capture moiety, or polymerizable molecule.

[0388] The substrates may be coupled to any useful number of molecules e.g., polymeric analytes, modified monomers, stacked plurality of modified monomers, capture moieties, polymerizable molecules). For example, the substrate may be coupled to 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 500, 1000, 5000, 10000, 100000, 500000, 1000000, 10000000 or more molecules. In some instances, a substrate may comprise a plurality of polymeric analytes e.g., peptides) and a plurality of anchor molecules, which may be provided at any useful ratio or density. For example, the ratio of polymeric analytes, modified monomers, or stacked plurality of modified monomers to anchor molecules may be about 1 : 1, 1 :5, 1 : 10,1 :20, 1 : 100, 1 : 1000, 1 : 10,000, 1 : 100,000, 1 : 1,000,000 or lower. In some instances, the ratio of polymeric analytes to anchor molecules may be at most about 1 : 1, at most about 1 :5, at most about 1 : 10, at most about 1 :20, at most about 1 : 100, at most about 1 : 1000, at most about 1 : 10,000, at most about 1 : 100,000, at most about 1 : 1,000,000 or lower.

[0389] Similarly, the molecules (e.g., polymeric analytes, modified monomers, stacked plurality of modified monomers, capture moieties, anchor molecules, or polymerizable molecules) may be coupled to the substrate at any useful density, for example about 1 molecule / square micron (gm2), about 10 molecules / pm2, about 100 molecules / pm2, about 1,000 molecules / pm2, about 10,000 molecules / pm2, about 100,000 molecules / pm2, about 1,000,000 molecules / pm2, about 10,000,000 molecules / pm2, about 100,000,000 molecules / pm2, about 1,000,000,000 molecules / pm2, about 10,000,000,000 molecules / pm2, about 100,000,000,000 molecules / pm2, or greater. The polymeric analytes, capture moieties, anchor molecules, and / or polymerizable molecules may be coupled to the substrate at a range of densities, e.g., from about 100 to about 10,000 molecules / pm2, or from about 10 to about 1,000 molecules / pm2. The density of the polymeric analytes, capture moieties, anchor molecules, and / or polymerizable molecules may be the same or different. For example, the density of the polymerizable molecules may be 1-fold, 2-fold, 3-fold, 4-fold, 5-fold, 6-fold, 7-fold, 8-fold, 9-fold, 10-fold, 100-fold, 1000-fold, 10,000-fold, 100,000-fold, 1,000,000-fold or greater-fold lower than that of the polymeric analyte.

[0390] In some instances, the molecules coupled to the substrate may be spaced apart at a designated or controlled distance. For example, the average spacing or distance between the anchor molecules, the capture moieties, or the processed polymeric analyte, e.g., the stacked plurality of modified monomers, may be spaced at a pitch of about 1 nanometer (nm), about 2 nm, about 3 nm, about 4 nm, about 5 nm, about 6 nm, about 8 nm, about 9 nm, about 10 nm, about 20 nm, about 30 nm, about 40 nm, about 50 nm, about 60 nm, about 70 nm, about 80 nm, about 90 nm, about 100 nm, about 500 nm, about 1pm, about 5 pm, about 10 pm or greater. In some instances, the spacing between or pitch of the molecules (e.g., anchor molecules, capture moieties, polymeric analyte, or processed polymeric analyte such as the stacked plurality of modified monomers) may be at most about 10 pm, at most about 5 pm, at most about 1pm, at most about 500 nm, at most about 100 nm, at most about 90 nm, at most about 80 nm, at most about 70 nm, at most about 60 nm, at most about 50 nm, at most about 40 nm, at most about 30 nm, at most about 20 nm, at most about 10 nm, at most about 5 nm, or less. Similarly, the spacing or distance between a polymeric analyte and a polymerizable molecule or capturemoiety may be about 1 nanometer (nm), about 2 nm, about 3 nm, about 4 nm, about 5 nm, about 6 nm, about 8 nm, about 9 nm, about 10 nm, about 20 nm, about 30 nm, about 40 nm, about 50 nm, about 60 nm, about 70 nm, about 80 nm, about 90 nm, about 100 nm, about 500 nm, about 1pm or greater. In some instances, the average spacing between the capture moiety and the polymeric analyte coupled to the substrate may be at most about 1pm, at most about 500 nm, at most about 100 nm, at most about 90 nm, at most about 80 nm, at most about 70 nm, at most about 60 nm, at most about 50 nm, at most about 40 nm, at most about 30 nm, at most about 20 nm, at most about 10 nm, at most about 5 nm, or less. A range of average distances between the polymerizable molecules from one another or from the polymeric analytes may be used, e.g., from about 1 nm to about 40 nm, from about 2 nm to about 10 nm, etc. In some instances, the substrate may be coupled to only a single molecule. In some instances, the molecules may be spaced or distributed unevenly across the substrate (e.g., with varying distances between the molecules).

[0391] The concentration or density of the molecules attached to the substrate may be modulated using one or more suitable approaches, including patterning or random deposition approaches. Examples of methods to control the concentration or density of the molecules attached to the substrate include limited dilution, addition of chaotropes (e.g., guanidine, formamide, urea), using metal organic compounds, etc. The molecules may be attached to the substrate in a patterned fashion, e.g., using self-assembling monolayers, photopatterning, lithography, etching, or a combination thereof, or the molecules may be randomly arranged.

[0392] The substrate may comprise any useful size or dimension (e.g., length, width, height, diameter, radius), surface area, volume, or ratio or combination thereof. The substrate may comprise a bead or particle that may comprise a diameter of about 1 nanometer (nm), about 2 nm, about 3 nm, about 4 nm, about 5 nm, about 6 nm, about 8 nm, about 9 nm, about 10 nm, about 20 nm, about 30 nm, about 40 nm, about 50 nm, about 60 nm, about 70 nm, about 80 nm, about 90 nm, about 100 nm, about 500 nm, about 1 m, about 2 pm, about 3 pm, about 4 pm, about 5 pm, about 6 pm, about 7 pm, about 8 pm, about 9 pm, about 10 pm, about 20 pm, about 30 pm, about 40 pm, about 50pm, about 60pm, about 70 pm, about 80 pm, about 90 pm, about 100 pm, about 200 pm, about 300 pm, about 400 pm, about 500 pm, about 600 pm, about 700 pm, about 800 pm, about 900 pm, about 1 millimeter (mm) or greater. The substrate may comprise a surface area of about 1 square nanometer (nm2), about 10 nm2, about 100 nm2, about 1,000 nm2, about 10,000 nm2, about 100,000 nm2, about 1 pm2, about 10 pm2, about 100 pm2, about 1,000 pm2, about 10,000 pm2, about 100,000 pm2, about 1 mm2, about 10 mm2, about 100mm2, about 1,000 mm2, about 10,000 mm2, about 100,000 mm2, about 1,000,000 mm2or greater.

[0393] The molecules may be coupled to the substrate in an ordered, semi-ordered, or random arrangement. In ordered arrangements, the molecules may be patterned using any conventional approach such as lithography (e.g., soft lithography, photolithography), etching (e.g., ion etching, photo etching), or other patterning approach. In some instances, a linker (e.g., bifunctional linker) may be used to facilitate the coupling of the molecules (e.g., polymeric analytes, polymerizable molecules, capture moieties) to the substrate; such linkers may be patterned using any useful technique such as self-assembling monolayers, photopatteming, lithography, etching. In some instances, the molecules may be coupled to the substrate in a random arrangement. For example, the molecules may be provided at a stoichiometric ratio or controlled concentration to couple the molecules at any useful ratio or density. In some instances, the substrate may comprise topographical or patterned features which may facilitate attachment of linkers to the patterned features.

[0394] In some instances, the methods provided herein may comprise using a plurality of substrates. For instance, the preparation of the modified monomers or stacked plurality of modified monomers may be performed using a substrate. The modified monomers or stacked plurality of modified monomers may then be removed from the substrate and contacted with an additional substrate for coupling and detection. In one such example, a modified monomer or stacked plurality of modified monomers may be contacted with a flow cell comprising one or more attachment or anchor molecules (e.g., anchor nucleic acid molecule). The modified monomer or stacked plurality of modified monomers may be coupled to the flow cell via one of the anchor molecules, linearized (e.g., using flow or electrophoretic force), and attached at another point to another anchor molecule. The linearized molecule may then be detected, e.g., using fluorescently labeled binders.

[0395] Polymerizable Molecules'. The polymerizable molecules described herein may be any useful type of polymerizable molecule. The polymerizable molecules may by naturally occurring, such as biological polymers (e.g., nucleic acid molecules, peptides, polysaccharides, fatty acids), or other naturally occurring polymers, e.g., rubber, cellulose, starches, polyhydroxyalkanoates, chitosan, dextran, structural proteins (e.g., collagen, hyaluronic acid, glycosaminoglycans), agarose, carrageenan, isphagula, acacia, agar, gelatin, shellac, xanthan gum, guar gum, alginate, etc. The polymerizable molecules may be synthetic, e.g., acrylics, nylons, silicones, viscose, rayon, polyesters, polycarboxylic acids, polyvinyl acetate,polyacrylamide, polyacrylate, polyethylene glycol, polyurethane, polylactic acid, silica, polystyrene, polyacrylonitrile, polybutadiene, polycarbonate, polyethylene terephthalate, poly(chlorotrifluoroethylene), poly(ethylene oxide), polyethylene terephthalate), polyethylene, polyisobutylene, poly(methyl methacrylate), poly(oxymethylene), polyformaldehyde, polypropylene, polystyrene, poly(tetrafluoroethylene), poly(vinyl acetate), poly(vinyl alcohol), poly(vinyl chloride), poly(vinylidene dichloride), poly(vinylidene difluoride), poly(vinyl fluoride) and combinations thereof. The polymerizable molecules may comprise one or more reactive moi eties (e.g., radical groups) to initiate polymerization or may be polymerized via contacting of an initiating agent (e.g., ammonium persulfate, peroxide, or other radicalizing agent). The polymerizable molecules may be polymerizable via contacting of an enzyme (e.g., polymerizing enzyme such as polymerases), ribozyme or DNAzyme. Alternatively or in addition to, the polymerizable molecules may be polymerizable via self-assembly. The polymerizable molecules may comprise a single polymer type (e.g., a homopolymer) or more than one polymer type (e.g., a copolymer) and may comprise random or arranged monomers. The polymerizable molecules may be a block polymer, alternating copolymer, periodic copolymer, statistical copolymer, stereoblock copolymer, gradient copolymer, branched copolymer, graft copolymer, etc.

[0396] The same or different types of polymerizable molecules may be used in the methods described herein. For example, the first polymerizable molecule comprised by or coupled to the binding agent may be a nucleic acid molecule, and the second polymerizable molecule may be a peptide. In another example, both the first polymerizable molecule and the second polymerizable molecule are nucleic acid molecules. In such an example, the first polymerizable molecule may be coupled to the second polymerizable molecule via ligation or hybridization. For instance, the first polymerizable molecule may comprise a first nucleic acid sequence and the second polymerizable molecule may comprise a second nucleic acid sequence. The first nucleic acid sequence may be complementary or partially complementary to the second nucleic acid sequence, and the coupling may comprise hybridizing the first nucleic acid sequence or portion thereof to the second nucleic acid sequence or portion thereof. Alternatively, the first nucleic acid sequence and the nucleic acid sequence may be complementary to two sequences of a splint or bridge oligonucleotide, and coupling may be mediated via hybridization to the splint oligo. The first nucleic acid sequence may be ligated to the second nucleic acid sequence, either chemically (e.g., via click chemistry approaches in which the first polymerizable molecule and the second polymerizable molecule comprise one member of a click chemistry pair) or enzymatically (e.g., using a ligase). In another example, both the first polymerizable moleculeand the second polymerizable molecule may be peptides, and the peptides may be coupled to one another, e.g., via enzymatic or chemical ligation.

[0397] The polymerizable molecules may comprise functional portions. For example, the polymerizable molecules may comprise a nucleic acid molecule comprising a functional sequence, such as a primer sequence (e.g., universal priming site), a sequencing sequence, a read sequence, a unique molecular identifier (UMI), a barcode sequence, a cleavage sequence (e.g., a restriction site, a Cas-binding sequence), a transposition sequence (e.g., a mosaic end sequence), or a combination thereof.

[0398] The polymerizable molecules may be any useful size. The polymerizable molecules may be about 1 angstrom, about 2 angstrom, about 3 angstrom, about 4 angstrom, about 5 angstrom, about 6 angstrom, about 7 angstrom, about 8 angstrom, about 9 angstrom, about 10 angstrom, about 20 angstrom, about 30 angstrom, about 40 angstrom, about 50 angstrom, about 60 angstrom, about 70 angstrom, about 80 angstrom, about 90 angstrom, about 100 angstrom, about 200 angstrom, about 300 angstrom, about 400 angstrom, bout 500 angstrom, about 600 angstrom, about 700 angstrom, about 800 angstrom, about 900 angstrom, about 1000 angstrom, about 10,000 angstrom, about 100,000 angstrom or greater in size, length, or another dimension. In some instances, the polymerizable molecule (e.g., the first polymerizable molecule or the second polymerizable molecule) comprises a nucleic acid molecule comprising one or more nucleotide bases. The polymerizable molecule may comprise any useful number of nucleotide bases, e.g., about 1 base, about 2 bases, about 3 bases, about 4 bases, about 5 bases, about 6 bases, about 7 bases, about 8 bases, about 9 bases, about 10 bases, about 20 bases, about 30 bases, about 40 bases, about 50 bases, about 60 bases, about 70 bases, about 80 bases, about 90 bases, about 100 bases, about 200 bases, about 300 bases, about 400 bases, about 500 bases, about 600 bases, about 700 bases, about 800 bases, about 900 bases, about 1000 bases, or a greater number of bases.

[0399] The polymerizable molecules may comprise a nucleic acid molecule. The nucleic acid molecule can be single stranded, double stranded, or partially double-stranded. The nucleic acid molecule may comprise a modified nucleotide or non-canonical base. For instance, the polymerizable molecules may comprise a pseudo-complementary base, a bridged nucleic acid (BNA), a xenonucleic acid (XNA), a locked nucleic acid (LNA), a peptide nucleic acid (PNA), a gamma-PNA molecule, a morpholino, or a combination thereof. In some instances, a polymerizable molecule may comprise a hexitol nucleic acid (HNA) or a cyclohexyl nucleic acid (CeNA), which may be useful in rendering the polymerizable molecule more resistant to acid degradation (e.g., as used in conventional Edman degradation). Alternatively or in additionto, a polymerizable molecule may comprise naturally occurring bases that are more resistant to acid degradation, e.g., be composed of primarily thymine or cytosine. For example, a nucleic acid molecule may comprise at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, or 100% thymines or cytosines, which can render the nucleic acid molecule more acid resistant as compared to a nucleic acid molecule comprising adenines or guanines.

[0400] Sequencing reagents'. One or more operations of the method may be mediated using a compound (e.g., sequencing reagent) provided herein, such as a sequencing agent of Formula (la), (lb), (I-Aa), (LAb), (Ila), (lib), (II-Aa), (Il-Ab), (Illa), (Illb), (III-Aa), (III-Ab), (IVa), (IVb), (IV-Aa), (IV-Ab), (IV-Ba), (IV-Bb), (V), (V-A), (V-B), (VI), (VI-A), (VII), or (VILA). In some instances, the coupling of the monomer to the capture moiety to generate a monomer- capture moiety complex is mediated using a sequencing reagent. The coupling of the sequencing reagent to the monomer or capture moiety may be covalent or noncovalent. In an example, a sequencing reagent may comprise a first reactive group that is able to couple to a monomer of the polymeric analyte (e.g., an amino acid of a peptide) and optionally, cleave the amino acid from a peptide. For example, the first reactive group may be an amino-acid reactive group, e.g., a guanidinyl group, an isothiocyanate (ITC) such as phenyl isothiocyanate (PITC), 3 -pyridyl isothiocyanate (PYITC), 2-piperidinoethyl isothiocyanate (PEITC), 3-(4-morpholino) propyl isothiocyanate (MPITC), 3- (diethylamino)propyl isothiocyanate (DEPTIC) or naphthylisothiocyanate (NITC), fluorescein isothiocyanate (FITC), ammonium thiocyanate, potassium thiocyanate, trimethyl silyl isothiocyanate (TMS-ITC), phenyl phosphoroisothiocyanatidate, acetyl isothiocyanate (AITC), or an aldehyde group, e.g., orthophthalaldehyde (OP A), , 3 -naphthalenedi carboxyaldehyde (NDA), 2-pyridinecarboxyaldehyde, which can react with an N-terminal amino acid (NTAA). The sequencing reagent may additionally comprise a second reactive group that is capable of coupling, either directly or indirectly, to the capture moiety. In an example of direct coupling, the capture moiety may comprise a click chemistry moiety (e.g., alkyne), and the second reactive group of the sequencing reagent may comprise an additional click chemistry moiety (e.g., azide) that can react with the click chemistry moiety of the capture moiety. Alternatively, the sequencing reagent may be coupled indirectly to the capture moiety, e.g., via noncovalent interaction or via an intermediate linking molecule. In some instances, the intermediate linking molecule may comprise a third polymerizable molecule (e.g., a polymer or nucleic acid molecule) that can couple the sequencing reagent to the capture moiety. In one such example, the third polymerizable molecule may comprise (i) a third reactive group that is capable of coupling tothe second reactive group (e.g., via alkyne-azide click chemistry) of the sequencing reagent and (ii) a moiety that can couple to the capture moiety (e.g., another orthogonal click chemistry reaction, avidin-biotin interaction, nucleic acid coupling or hybridization). In some instances, the third polymerizable molecule comprises a nucleic acid molecule that comprises (i) a click chemistry moiety e.g., alkyne) that can conjugate to the first reactive group e.g., azide) of the sequencing reagent and (ii) a nucleic acid sequence that can couple to the capture moiety, e.g., via ligation, splint ligation, or hybridization. In some instances, the sequencing reagent comprises a linking nucleic acid molecule that comprises a self-splinting moiety.

[0401] When applicable, the click chemistry moieties of the sequencing reagent and capture moiety or intermediate linking molecule may comprise any suitable bioorthogonal moieties, as described elsewhere herein, e.g., alkenes, alkynes, azides, epoxides, amines, thiols, nitrones, isonitriles, isocyanides, aziridines, activated esters, and tetrazines, and combinations, variations, or derivatives thereof. The sequencing reagent may be subjected to conditions sufficient to react the first click chemistry moiety to the second click chemistry moiety, e.g., provision of metal catalysts, appropriate solvents, pH, temperature, ionic concentration, or light / energy for any useful duration of time.

[0402] The first reactive group of the sequencing reagent may be an amino acid-reactive moiety. The amino acid- reactive moiety of the sequencing reagent may be any useful moiety that enables the reactive moiety to conjugate to and optionally cleave an amino acid. In some examples, the first reactive moiety can react with a terminal amino acid (e.g., NTAA or CTAA). In such examples, the first reactive ...

Claims

CLAIMSWHAT IS CLAIMED IS:

1. A compound of Formula (la) or Formula (lb):Formula (la) Formula (lb), wherein:LG is a leaving group, optionally substituted with one or more electronwithdrawing groups;L1is a linker, a bond, or absent;R1, R2, and R3are each independently hydrogen, a protecting group, an aminoprotecting group, an electron-withdrawing group, an amino-electron- withdrawing group, or an electron-withdrawing protecting group; andB is a reactive moiety or a polymer.

2. The compound of claim 1, represented by Formula (I-Aa) or Formula (I-Ab):Formula (LAa) Formula (I- Ab), wherein:Ring A is C4-C12 heteroaryl or C4-C12 heterocycle, optionally substituted with one or more electron-withdrawing groups.

3. A compound of Formula (Ila) or Formula (lib):Formula (Ila) Formula (lib), wherein,LG is a leaving group, optionally substituted with one or more electronwithdrawing groups;R1and R2are each independently hydrogen, a protecting group, an aminoprotecting group, an electron-withdrawing group, an amino-electron- withdrawing group, or an electron-withdrawing protecting group;L1is a linker, a bond, or absent;B is a reactive moiety or a polymer; and n is an integer from 0 to 3.

4. The compound of claim 3, represented by Formula (II-Aa) or Formula (Il-Ab):Formula (II-Aa) Formula (II- Ab) wherein:Ring A is C4-C12 heteroaryl or C4-C12 heterocycle, optionally substituted with one or more electron-withdrawing groups.

5. A compound of Formula (Illa), Formula (Illb), Formula (IVa), or Formula (IVb):Formula (Illa), Formula (Illb), Formula (IVa), Formula (IVb), wherein,LG is a leaving group, optionally substituted with one or more electronwithdrawing groups;L1is a linker, a bond, or absent;L2is a hydrogen or a linker;L3and L4are independently linkers, taken together to form heterocycle;B is a reactive moiety or a polymer; andR1and R2are each independently hydrogen, a protecting group, an aminoprotecting group, an electron-withdrawing group, an amino-electron- withdrawing group, or an electron-withdrawing protecting group.

6. The compound of claim 5, represented by Formula (III-Aa), Formula (III-Ab), Formula (IV-Aa), or Formula (IV-Ab):Formula (III-Aa), Formula (III-Ab), Formula (IV-Aa), Formula (IV-Ab) whereinRing A is C4-C12 heteroaryl or C4-C12 heterocycle, optionally substituted with one or more electron-withdrawing groups.

7. The compound of claim 5 or 6, wherein L3and L4are taken together to form Cs-Cs heterocycle.

8. The compound of any one of claims 5-7, represented by Formula (IV-Ba) or Formula (IV-Bb):Formula (I V-B a) Formula (IV-Bb), wherein n is an integer from 0 to 3.

9. A compound of Formula (V):Formula (V), wherein each LG is independently a leaving group, optionally substituted with one or more electron-withdrawing groups; andL1is a linker, a bond, or absent; andB is a reactive moiety or a polymer.

10. The compound of claim 9, represented by Formula (V-A) or Formula (V-B):Formula (V-A) Formula (V-B) wherein each Ring A is independently C4-C12 heteroaryl or C4-C12 heterocycle, optionally substituted with one or more electron-withdrawing groups.

11. A compound of Formula (VI) or Formula (Formula (VI) Formula (VII), wherein each LG is independently a leaving group, optionally substituted with one or more electron-withdrawing groups; andL1is a linker, a bond, or absent; andL2is a hydrogen or a linker;L3and L4are independently linkers, taken together to form heterocycle;B is a reactive moiety or a polymer.

12. The compound of claim 11, represented by Formula (VI-A) or Formula (VII-A):Formula (VI-A) Formula (VII-A), wherein each Ring A is independently C4-C12 heteroaryl or C4-C12 heterocycle, optionally substituted with one or more electron-withdrawing groups.

13. The compound of any one of the preceding claims, wherein the leaving group is electronwithdrawing.

14. The compound of any one of the preceding claims, wherein the leaving group is C4-C12 heteroaryl or C4-C12 heterocycle, each optionally substituted with one or more electronwithdrawing groups.

15. The compound of any one of the preceding claims, wherein the leaving group or Ring A is C4-C6 heteroaryl, optionally substituted with one or more electron-withdrawing groups.

16. The compound of any one of the preceding claims, wherein the leaving group or Ring A is an azole or azine, optionally substituted with one or more electron-withdrawing groups.

17. The compound of any one of the preceding claims, wherein the leaving group or Ring A is diazole, triazole, or tetrazole, optionally substituted with one or more electronwithdrawing groups.

18. The compound of any one of the preceding claims, wherein the leaving group or Ring A is diazole, triazole, or tetrazole, each optionally fused with aryl or heteroaryl.

19. The compound of any one of the preceding claims, wherein the leaving group or Ring A is triazole fused with aryl or heteroaryl.

20. The compound of any one of the preceding claims, wherein the leaving group or Ring A is:wherein, each of X1-X5 is independently selected from N, CH, or C-EWG; each of Y1-Y4 is independently selected from N, CH, or C-EWG; and EWG is an electron-withdrawing group; wherein at least one of X1-X5 is N.

21. The compound of any one of the preceding claims, wherein the leaving group or Ring Aoptionally substituted with one or more electron- withdrawing groups.

22. The compound of any one of the preceding claims, wherein the leaving group or Ring A is C4-C6 heteroaryl fused with C4-C6 aryl, optionally substituted with one or more electron-withdrawing groups.

23. The compound of any one of the preceding claims, wherein the leaving group or Ring A is diazole, triazole, or tetrazole, fused with aryl, optionally substituted with one or more electron-withdrawing groups.

24. The compound of any one of the preceding claims, wherein the leaving group or Ring Agroups.

25. The compound of any one of the preceding claims, wherein the electron-withdrawing group is each independently selected from halogen, haloalkyl, nitro, sulfonate, amino, alkylamino, cyano, or a carbonyl.

26. The compound of claim 25, wherein the carbonyl comprises an acetyl group or a derivative thereof, a carboxylic acid, or an aldehyde.

27. The compound of any one of the preceding claims, wherein at least one electronwithdrawing group comprises Ci-Ce haloalkyl.

28. The compound of any one of the preceding claims, wherein at least one electronwithdrawing group comprises CF3.

29. The compound of any one of the preceding claims, wherein at least one electronwithdrawing group comprises nitro.

30. The compound of any one of the preceding claims, wherein at least one electronwithdrawing group comprises a halogen.

31. The compound of claim 30, wherein the halogen is fluoro or chloro.

33. The compound of any one of the preceding claims, wherein the linker is configured to enable a nucleic acid attraction.

34. The compound of claim 33, wherein the nucleic acid attraction is ionic.

35. The compound of claim 33, wherein the nucleic acid attraction is non-ionic.

36. The compound of any one of the preceding claims, wherein the linker is configured to provide steric space for B, wherein B is a click chemistry moiety.

37. The compound of any one of the preceding claims, wherein the linker is electronwithdrawing.

38. The compound of any one of the preceding claims, wherein a first end of the linker is el ectron- withdrawing .

39. The compound of any one of the preceding claims, wherein both a first end and a second end of the linker are electron-withdrawing.

40. The compound of any one of the preceding claims, wherein each linker is independently selected from optionally substituted alkyl, optionally substituted heteroalkyl, an oligonucleotide, DNA, RNA, or a peptide.

41. The compound of any one of the preceding claims, wherein the linker is or comprises:wherein at least one of x is an integer greater than 0.

42. The compound of any one of the preceding claims, wherein at least one linker is -CH2-.

43. The compound of any one of the preceding claims, wherein at least one of R1, R2, and R3is an amine-protecting group.

44. The compound of any one of the preceding claims, wherein R1, R2, and R3are each independently hydrogen, a protecting group, or an amino-protecting group, wherein the protecting group or amino-protecting group comprises carbobenzyl oxy (CBz) (e.g., benzyl carbamate), acetamide (Ac), trifluoroacetyl (TFAc), phthalimide, benzyl, benzylamine (Bn), benzoyl (Bz), triphenylmethyl (Tr) (e.g., triphenylmethylamine), benzylideneneamine, tosyl (Ts) (e.g., / ?-toluenesulfonamide), Methylsulfonylethoxycarbonyl (Msc), tert-butyloxycarbonyl (Boc) (e.g., / -butyl carbamate), or fluorenylmethyloxy carbonyl (Fmoc) (e.g., 9-fluorenylmethyl carbamate), 2,7-disulfo-9-fluorenylmethoxycarbonyl (Smoc), or acetyl.

45. The compound of any one of the preceding claims, wherein the reactive moiety B comprises a click chemistry moiety.

46. The compound of claim 45, wherein the click chemistry moiety is an azide, a sulfonyl azide, an alkyne, a thio acid, a tetrazine, bicyclo[6.1.0]nonyne (BCN), a diarylcyclooctyne (e.g., DBCO), transcyclooctene (TCO), cyclopropane, norbomene, or a spiroalkene.

47. The compound of any one of the preceding claims, wherein the reactive moiety B comprises a phosphite, phosphine, pentafluorophenyl ester (PFP), tetrafluorophenyl ester (TFP), 4-sulfo-2,3,5,6-tetrafluorophenyl ester (STP), or thio-phthalimide.

48. The compound of any one of claims 1, 2, or 13-47, wherein the compound is represented by the structure:

49. The compound of any one of claims 3, 4, or 13-47, wherein the compound is represented by the structure:

50. The compound of any one of claims 5-8, or 13-47, wherein the compound is represented by the structure:

51. The compound of any one of claims 9, 10, or 13-47, wherein the compound is represented by the structure:

52. The compound of any one of claims 11, 12, or 13-47, wherein the compound is represented by the structure:

53. The compound of any one of claims 1-44, wherein the reactive moiety B comprises the polymer.

54. The compound of claim 53, wherein the polymer comprises a polynucleotide, a polypeptide, a polymeric material, or a combination thereof.

55. The compound of claim 54, wherein the polynucleotide comprises a deoxyribonucleic acid (DNA), a ribonucleic acid (RNA), a locked nucleic acid (LNA), a parallel-stranded DNA (p-DNA), a non-constrained nucleic acid (NNA), or a modified version thereof.

56. The compound of claim 54, wherein the polynucleotide is a deoxyribonucleic acid (DNA).

57. The compound of any one of the preceding claims, wherein the compound is a sequencing reagent.

58. A method of using the sequencing reagent of claim 56, the method comprising:(a) providing a polymeric analyte;(b) contacting the polymeric analyte with the sequencing reagent, wherein the sequencing reagent binds to a monomer of the polymeric analyte thereby generating a sequencing reagent-monomer complex; and(c) cleaving the sequencing reagent-monomer complex from the polymeric analyte, thereby providing a cleaved sequencing reagent-monomer complex.

59. The method of claim 58, further comprising, coupling the sequencing reagent or the sequencing reagent-monomer complex to a capture moiety.

60. The method of claim 59, wherein the capture moiety comprises a DNA molecule.

61. The method of claim 59, wherein the capture moiety or the sequencing reagent comprises a modified DNA molecule.

62. The method of claim 61, wherein the modified DNA molecule comprises a click chemistry moiety-modified DNA molecule.

63. The method of claim 61 or 62, wherein the sequencing reagent is coupled to the modified DNA molecule via click chemistry.

64. The method of any one of claims 58-63, wherein the polymeric analyte comprises a polypeptide.

65. The method of any one of claims 58-64, wherein the monomer comprises a terminal amino acid residue.

66. The method of any one of claims 58-65, further comprising, detecting the cleaved sequencing reagent-monomer complex or derivative thereof, wherein detectingcomprises contacting the sequencing reagent-monomer complex or derivative thereof with a binding agent.

67. The method of claim 66, wherein the binding agent comprises an antibody, nanobody, single chain variable fragment (scFv), or aptamer.

68. The method of claim 66, wherein the binding agent comprises a polymerizable molecule.

69. The method of claim 68, wherein the polymerizable molecule comprises a nucleic acid molecule.

70. The method of any one of claims 58-69, further comprising deprotecting the protecting groups (e.g., R\R2, and / or R3) of the sequencing reagent.

71. The method of claim 70, wherein the deprotecting is performed before cleaving the sequence reagent-monomer complex from the polymeric analyte.

72. The method of claim 70, wherein the deprotecting is performed before coupling the sequencing reagent to the capture moiety.

73. The method of any one of claims 70-72, wherein the deprotecting is performed in the presence of a base.

74. The method of any one of claims 58-73, further comprising:(d) contacting the polymeric analyte with an additional sequencing reagent, wherein the additional sequencing reagent binds to an additional monomer of the polymeric analyte, thereby generating an additional sequencing reagent-monomer complex;(e) coupling the additional sequencing reagent to the cleaved sequencing reagent- monomer complex; and(f) cleaving the additional sequencing reagent-monomer complex from the polymeric analyte, thereby providing a stacked sequencing reagent-monomer complex.

75. The method of any one of claims 58-74, wherein in (a), the polymeric analyte is coupled to a substrate.

76. The method of any one of claims 58-75, wherein cleaving the sequencing reagent- monomer complex from the polymeric analyte is performed chemically or enzymatically.

77. The method of claim 76, wherein cleaving the sequencing reagent-monomer complex from the polymeric analyte is performed chemically in presence of a base.

78. The method of any one of claims 58-77, further comprising, detecting the cleaved sequencing reagent-monomer complex or derivative thereof.

79. The method of claim 78, wherein detecting is performed using a nanopore.

0. The method of claim 79, wherein detecting is performed by measuring a signal from the nanopore as the cleaved sequencing reagent-monomer complex or derivative thereof translocates through the nanopore, thereby generating a measured signal, and using the measured signal to identify the monomer.

Citation Information

Patent Citations

  • Single-molecule protein and peptide sequencing

    US11499979B2

  • Protein sequencing via coupling of polymerizable molecules

    US12259393B2

  • Protein sequencing via coupling of polymerizable molecules

    US12259394B2

  • C-Terminal Peptide Modification

    US20250188117A1

  • Single-molecule peptide sequencing using xanthate amino acid reactive groups

    US20250188538A1