Single molecule peptide sequencing using dithioesters and thiocarbamoyl amino acid reactive groups

By using a sequencing reagent of Formula I to form a covalent bond with the N-terminal amino acid of a polypeptide and utilizing a nanopore sequencing system, the problems of insufficient sensitivity and spatial information in protein sequencing in the prior art are solved, and efficient sequencing of low-copy number proteins is achieved.

CN120604125APending Publication Date: 2025-09-05GLYPHIC BIOTECHNOLOGIES INC
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202380091504.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2022-11-16
Filing Date
2023-11-14
Publication Date
2025-09-05

AI Technical Summary

Technical Problem

Existing protein sequencing methods such as mass spectrometry and Edman sequencing have deficiencies in sensitivity and throughput, cannot effectively quantify low-copy number proteins, and lack single-molecule sensitivity and spatial information.

Method used

A sequencing reagent of formula I is used, which contains a dithioester or thiocarbamoyl reactive group, forms a covalent bond with the N-terminal amino acid of the polypeptide, is partially connected to the polymer via click chemistry, and protein sequencing is performed using a nanopore sequencing system.

Benefits of technology

It achieves high-sensitivity sequencing of low-copy number proteins, provides single-molecule detection capabilities, and can provide spatial information in the cellular environment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120604125A_ABST
    Figure CN120604125A_ABST
Patent Text Reader

Abstract

The present disclosure provides reagents and methods useful for single molecule sequencing of proteins by using unique sequencing reagents. The reagents and methods described herein provide high throughput single molecule peptide and protein sequencing under mild conditions, allowing for high resolution study of complex biological systems.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross-references

[0002] This application claims the benefit of U.S. Provisional Patent Application No. 63 / 384,007, filed on November 16, 2022, which is incorporated herein by reference in its entirety.

[0003] STATEMENT REGARDING FEDERALLY SPONSORED RESEARCH OR DEVELOPMENT

[0004] This invention was made with U.S. Government support under Grant No. HG012563 awarded by the National Institutes of Health. The Government has certain rights in this invention. Background Art

[0005] Proteins play a vital role at the cellular level, carrying out a variety of indispensable functions. Having the technology required to quantify and identify proteins is crucial for understanding their contributions to biological function. Advances in proteomics have lagged behind, while DNA sequencing has rapidly advanced genomic research, largely due to technologies that allow high-throughput sequencing. Currently available methods for studying proteins include mass spectrometry, Edman sequencing, and immunohistochemistry.

[0006] Mass spectrometry (MS) enables protein identification and quantification based on the mass-to-charge ratio of peptide fragments that can be bioinformatically mapped back to genomic databases. However, despite significant progress, MS has not yet quantified a complete set of proteins from biological systems. MS exhibits attomole detection sensitivity for whole proteins and subattomole sensitivity after fractionation. However, functionally important low-copy number proteins, which account for approximately 10% of mammalian protein expression, remain undetected.

[0007] Edman degradation allows for the sequential and selective removal of individual N-terminal amino acids, which are subsequently identified by HPLC (high performance liquid chromatography). Edman protein sequencing uses phenyl isothiocyanate (PITC) conjugated to the N-terminal amino acid, and then upon acid and heat treatment, the PITC-labeled N-terminal amino acid is removed to remove the first N-terminal amino acid for identification. Although Edman sequencing can have an efficiency of 98%, its main disadvantages are its inherent low throughput, the requirement for highly purified single proteins, and its unsuitability for whole systems biology. In addition, Edman degradation has many other disadvantages, including harsh reaction conditions (such as heat and acidic conditions that are not suitable for the use or analysis of nucleic acids), the resulting chiral residues that can hinder detection via stereoisomer-specific binders, and the modification of lysine residues that can hinder further functionalization of the lysine side chain. Both Edman degradation and mass spectrometry can sequence proteins, but lack single-molecule sensitivity and do not provide spatial information about proteins in a cellular environment.

[0008] Regarding spatial information, immunohistochemistry is a protein identification method that allows visualization of the cellular localization of proteins but does not provide sequence information. Immunohistochemistry involves identifying proteins through the use of fluorophore-conjugated antibodies. This method does not include protein sequence information but allows identification of proteins and their respective localization. The main limitation is scalability, as even the perfect construction of antibodies specific for every protein in the proteome would require approximately 25,000 antibodies and approximately 6,250 rounds of four-color imaging. Summary of the Invention

[0009] In view of the current need for improved single-molecule protein sequencing methods, formulas, compounds, and methods are provided herein to meet these needs. A brief overview of various exemplary embodiments is provided. Some simplifications and omissions may be made in the following overview to emphasize and introduce aspects of certain embodiments disclosed herein, but do not limit the scope of the present disclosure. Subsequent sections will describe various embodiments in detail sufficient to enable one of ordinary skill in the art to implement and use the concepts disclosed herein.

[0010] In some aspects of the present disclosure, provided herein is a sequencing reagent of Formula I:

[0011]

[0012] or a stereoisomer, tautomer, or salt thereof. In some embodiments, A comprises a (e.g., first) reactive group configured to form a covalent bond with the N-terminal amino acid of the polypeptide, wherein the reactive group is selected from a dithioester or a thiocarbonyl. In some embodiments, the thiocarbonyl is a thiocarbamoyl. In some embodiments, A comprises a (e.g., first) reactive group configured to form a covalent bond with the N-terminal amino acid of the polypeptide, wherein the first reactive group is a dithioester or a thiocarbamoyl. In some embodiments, B comprises a second reactive group. In some embodiments, L 1 Includes a linker coupled to A and B.

[0013] In some embodiments, the second reactive group is covalently attached or configured to be covalently attached to a polymer. In some embodiments, the second reactive group is covalently attached to a polymer. In some embodiments, the polymer includes polyethylene glycol (PEG). In some embodiments, the polymer includes deoxyribonucleic acid (DNA) or ribonucleic acid (RNA). In some embodiments, the polymer is covalently attached or configured to be covalently attached to a surface. In some embodiments, the polymer is covalently attached to a surface.

[0014] In some embodiments, the second reactive group is covalently attached or configured to be covalently attached to a surface-bound linker. In some embodiments, the second reactive group is covalently attached to a surface-bound linker. In some embodiments, the second reactive group is covalently attached to a surface-bound linker, and wherein the surface-bound linker is attached to the surface. In some embodiments, the surface-bound linker comprises an ethyl group. In some embodiments, the surface-bound group comprises a propyl group. In some embodiments, the surface-bound linker comprises a nucleic acid molecule.

[0015] In some embodiments, the second reactive group comprises:

[0016]

[0017] in Indicates the orientation of the second reactive group relative to the reactive group. In some embodiments, the reactive group includes:

[0018]

[0019] in Indicates the orientation of the reactive group relative to a second reactive group. In some embodiments, the reactive group includes:

[0020]

[0021] in Indicates the orientation of the reactive group relative to a second reactive group. In some embodiments, the reactive group includes: in Indicates the orientation of the reactive group relative to a second reactive group. In some embodiments, the reactive group includes:

[0022]

[0023] in Indicates the orientation of the reactive group relative to a second reactive group. In some embodiments, the reactive group includes: in Indicates the orientation of the first reactive group relative to the second reactive group. In some embodiments, the reactive group includes in Indicates the orientation of the first reactive group relative to the second reactive group. In some embodiments, the reactive group comprises a thioacetyl group. In some embodiments, the reactive group comprises a thiobenzoyl group. In some embodiments, the reactive group comprises a derivative of N-thiobenzoylsuccinimide. In some embodiments, the reactive group comprises a derivative of cyanomethyldithiobenzoate.

[0024] In some embodiments, the second reactive group comprises a click chemistry moiety. In some embodiments, the click chemistry moiety is an azide, an alkyne, a dibenzocyclooctyne (DBCO), a tetrazine, or a trans-cyclooctene (TCO).

[0025] In some embodiments, L1 comprises a cleavable linker. In some embodiments, the cleavable linker comprises a disulfide bond, a hydrazone, a DNA molecule comprising a cleavage site, a peptide cleavable by an enzyme, or a click chemistry moiety. In some embodiments, the cleavable linker comprises a hydrazone. In some embodiments, the cleavable linker comprises o-aminobenzylhydrazone.

[0026] In some embodiments, L1 comprises a non-cleavable linker.

[0027] In some embodiments, the reactive group is covalently linked to the N-terminal amino acid of the polypeptide. In some embodiments, the second reactive group is directly or indirectly linked to the substrate. In some embodiments, the reactive group is covalently linked to the N-terminal amino acid of the polypeptide that is linked to the substrate; and wherein the second reactive group is directly or indirectly linked to the substrate.

[0028] Another aspect of the present disclosure provides a method of using a sequencing reagent of Formula I, the method comprising providing a substrate, a substrate-bound capture moiety, and a polymer analyte, contacting the polymer analyte with the sequencing reagent, wherein the sequencing reagent binds to monomers of the polymer analyte to form a sequencing reagent-monomer complex, tethering the sequencing reagent-monomer complex to the substrate via the substrate-bound capture moiety, cleaving the sequencing reagent-monomer complex from the polymer analyte to provide a detectable complex, and detecting the detectable complex.

[0029] Another aspect of the present disclosure provides a method of using a sequencing reagent, the method comprising (a) L1 comprising a non-cleavable linker; (b) contacting a polymer analyte with a sequencing reagent, wherein the sequencing reagent binds to monomers of the polymer analyte to form a sequencing reagent-monomer complex; (c) tethering the sequencing reagent-monomer complex to a capture moiety; (d) cleaving the sequencing reagent-monomer complex from the polymer analyte to provide a detectable complex; and (e) detecting the detectable complex.

[0030] In some embodiments, the polymer analyte comprises a polypeptide and the method comprises contacting the polypeptide with an alkylating agent before contacting the polypeptide with a sequencing reagent. In some embodiments, the alkylating agent comprises 4-vinylpyridine. In some embodiments, the alkylating agent comprises iodoacetamide. In some embodiments, the polymer analyte comprises a polypeptide. In some embodiments, the monomer comprises a terminal amino acid residue. In some embodiments, the substrate-bound capture portion comprises a DNA primer. In some embodiments, the capture portion comprises a DNA molecule. In some embodiments, detecting the detectable complex comprises contacting the sequencing reagent-monomer complex with a binding agent. In some embodiments, the binding agent comprises an antibody, a nanobody, a single-chain variable fragment (scFv), or an aptamer. In some embodiments, the binding agent comprises a polymerizable molecule. In some embodiments, the polymerizable molecule comprises a nucleic acid. In some embodiments, the method further comprises repeating the method, thereby sequencing the polymer analyte.

[0031] In some embodiments, the method further comprises coupling the polymerizable molecule to a capture moiety or an additional polymerizable molecule, and wherein detecting comprises sequencing the polymerizable molecule.

[0032] In some embodiments, the capture moiety is coupled to the polymer analyte, and further comprising partitioning the sequencing reagent-monomer complex into partitions using a binding agent, and in the partitions, coupling the barcode molecule to the capture moiety.

[0033] In some embodiments, the binding agent comprises a fluorophore, and wherein detecting comprises imaging the fluorophore.

[0034] In some embodiments, the method further comprises repeating (b)-(d).

[0035] In some embodiments, detecting comprises translocating the detectable complex through the nanopore and identifying the monomer.

[0036] In some embodiments, the method further comprises repeating the method, thereby sequencing the polymer analyte.

[0037] Another aspect of the present disclosure provides a non-transitory computer-readable medium comprising machine-executable code that, when executed by one or more computer processors, implements any of the methods described above or elsewhere herein.

[0038] Another aspect of the present disclosure provides a system comprising one or more computer processors and a computer memory coupled thereto, wherein the computer memory comprises machine executable code that, when executed by the one or more computer processors, implements any method described above or elsewhere herein.

[0039] Further aspects and advantages of the present disclosure will become apparent to those skilled in the art from the detailed description that follows, wherein only illustrative embodiments of the present disclosure are shown and described. As will be appreciated, the present disclosure is capable of other and different embodiments, and its several details are capable of modification in various obvious respects, all without departing from the present disclosure. Accordingly, the drawings and description are to be regarded as illustrative in nature, and not restrictive.

[0040] Incorporation by reference

[0041] All publications, patents, and patent applications mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent, or patent application was specifically and individually indicated to be incorporated by reference. To the extent that publications and patents or patent applications incorporated by reference conflict with the disclosure contained in this specification, this specification is intended to supersede and / or take precedence over any such conflicting material. BRIEF DESCRIPTION OF THE DRAWINGS

[0042] The novel features of the present invention are set forth with particularity in the appended claims. A better understanding of the features and advantages of the present invention will be obtained by reference to the following detailed description and accompanying drawings (also referred to herein as "Figures" and "FIG.") which set forth illustrative embodiments in which the principles of the invention are utilized, wherein:

[0043] Figure 1A An example workflow for processing polymeric analyte molecules (eg, peptides) described herein is schematically shown. Figure 1BAnother example workflow for processing polymer analytes in solution is schematically shown. Figure 1C Another example workflow for processing polymer analytes and detecting them is schematically shown. Figure 1D Another example workflow for processing and detecting polymer analytes using a nanopore sequencing system is schematically shown. Figure 1E Another example workflow for processing and detecting polymer analytes using a nanopore sequencing system is schematically shown. Figure 1F Another example workflow for processing and detecting polymer analytes in solution is schematically shown.

[0044] Figure 2 Example linkers for attaching polymerizable molecules to polymeric analytes are shown schematically.

[0045] Figure 3 An example of a sequencing method to analyze or characterize a polymer analyte is schematically shown.

[0046] Figure 4 A computer system programmed or otherwise configured to implement the methods provided herein is schematically illustrated.

[0047] Figure 5 A representative reaction scheme for the removal and stabilization of the N-terminal amino acid of a peptide using a thioacetyl moiety is shown.

[0048] Figure 6 A representative reaction scheme is shown for the removal, stabilization, and tethering of the N-terminal amino acid of a peptide to a substrate using a thioacetyl moiety.

[0049] Figure 7 A representative reaction scheme is shown for the removal, stabilization, and tethering of the N-terminal amino acid of a peptide to a substrate using a thiocarbamyl moiety.

[0050] Figure 8 Another representative reaction scheme using the sequencing reagents described herein is shown.

[0051] Figure 9 Example schematics showing the synthesis of example sequencing reagents described herein.

[0052] Figure 10 Another example schematic diagram illustrating the synthesis of example sequencing reagents described herein.

[0053] Figure 11 Another example schematic diagram illustrating the synthesis of example sequencing reagents described herein.

[0054] Figure 12 Another example schematic diagram illustrating the synthesis of example sequencing reagents described herein.

[0055] Figure 13 Another example schematic diagram illustrating the synthesis of example sequencing reagents described herein. DETAILED DESCRIPTION

[0056] Although various embodiments of the present invention have been shown and described herein, it will be apparent to those skilled in the art that such embodiments are provided by way of example only. Many variations, modifications, and substitutions may occur to those skilled in the art without departing from the present invention. It should be understood that various alternatives to the embodiments of the present invention described herein may be employed.

[0057] definition

[0058] When the term "at least," "greater than," or "greater than or equal to" precedes the first value in a series of two or more values, the term "at least," "greater than," or "greater than or equal to" applies to every value in the series. For example, greater than or equal to 1, 2, or 3 is equivalent to greater than or equal to 1, greater than or equal to 2, or greater than or equal to 3.

[0059] When the term "not more than," "less than," or "less than or equal to" precedes the first value in a series of two or more values, the term "not more than," "less than," or "less than or equal to" applies to every value in the series. For example, less than or equal to 3, 2, or 1 is equivalent to less than or equal to 3, less than or equal to 2, or less than or equal to 1.

[0060] References to "one embodiment," "an embodiment," "example embodiment," "some embodiments," "certain embodiments," "various embodiments," etc., indicate that embodiments of the disclosed technology so described may include a particular feature, structure, or characteristic, but not every embodiment is required to include that particular feature, structure, or characteristic. Furthermore, repeated use of the phrase "in one embodiment" does not necessarily refer to the same embodiment, although it may.

[0061] In this article, ranges can be expressed as from "about" or "approximately" or "substantially" a specific value and / or to "about" or "approximately" or "substantially" another specific value. When such a range is expressed, other exemplary embodiments include from one specific value and / or to another specific value. Further, the term "about" means within an acceptable error range of a specific value as determined by those skilled in the art, which will depend in part on how the value is measured or determined, i.e., the limitations of the measurement system. For example, according to the practice of this art, "about" can mean within an acceptable standard deviation. Alternatively, "about" can mean a range of up to ±20%, preferably up to ±10%, more preferably up to ±5%, and still more preferably up to ±1% of a given value. Alternatively, particularly with respect to biological systems or processes, the term can mean within an order of magnitude, preferably within 2 times of a value. When describing a specific value in the application and claims, unless otherwise indicated, the term "about" is implicit and means within an acceptable error range of the specific value in this context.

[0062] “Comprising” or “containing” or “including” means that at least the specified compound, element, particle or method step is present in the composition or article or method, but does not exclude the presence of other compounds, materials, particles, method steps, even if other such compounds, materials, particles, method steps have the same function as that specified.

[0063] Throughout this specification, various components may be identified with specific values ​​or parameters, however, these items are provided as exemplary embodiments. In fact, the exemplary embodiments do not limit the various aspects and concepts of the present disclosure, as many comparable parameters, sizes, ranges and / or values ​​can be implemented. The terms "first," "second," etc., "primary," "secondary," etc. do not indicate any order, quantity, or importance, but are used to distinguish one element from another.

[0064] As used herein, the term "protein" generally refers to a molecule comprising two or more amino acids connected by a peptide bond. Protein may also be referred to as "polypeptide", "oligopeptide" or "peptide". Protein may be a naturally occurring molecule or a synthetic molecule (e.g., artificial protein, peptide, enzyme). Protein may comprise one or more non-natural amino acids, modified amino acids or non-amino acid linkers. Protein may contain D-amino acid enantiomers, L-amino acid enantiomers or both. The amino acids of a protein may be modified naturally or synthetically, such as by post-translational modification or by chemical modification. In some cases, different proteins may be distinguished from each other based on the different genes, different primary sequence lengths or different primary sequence compositions expressed in an organism. However, proteins expressed from the same gene may be different protein variants (proteoforms), such as those distinguished based on non-identical lengths, non-identical amino acid sequences or non-identical post-translational modifications. Different proteins may be distinguished based on one or both of the gene of origin and the protein variant state.

[0065] As used herein, the term "peptide" can refer to any short single peptide chain. The length of the peptide can be no more than about 100, 95, 90, 85, 80, 75, 70, 65, 60, 55, 50, 45, 40, 35, 30, 25, 20, 15, 10, 5 or less than about 5 amino acids. The peptide can have a known or unknown biological function or activity. The peptide can include natural, synthetic, modified or degraded proteins or peptides, or a combination thereof.

[0066] As used herein, the term "single analyte" can refer to an analyte that is manipulated individually or distinguished from other analytes. A single analyte can include a biomolecule or a synthetic molecule. A single analyte can include a small molecule. A single analyte can be a single molecule (e.g., a single biomolecule, such as a single protein, nucleic acid molecule, affinity reagent, lipid, carbohydrate, etc.), a single complex of two or more molecules (e.g., a multimeric protein with two or more separable subunits, a single protein attached to a nucleic acid molecule, or a single protein attached to an affinity reagent), a single particle, etc. Reference to a "single analyte" herein in the context of a composition, system, or method does not necessarily exclude the application of the composition, system, or method to multiple single analytes that are manipulated individually or distinguished, unless the context or clearly indicates otherwise.

[0067] As used herein, "polypeptide" refers to two or more amino acids linked together by peptide bonds. The term "polypeptide" includes proteins having a C-terminus and an N-terminus as generally known in the art, and can be synthetic or naturally occurring. As used herein, "at least a portion of a polypeptide" refers to two or more amino acids in a polypeptide. A polypeptide may comprise one or more peptides. Optionally, a portion of a polypeptide includes at least: 1, 5, 10, 20, 30 or 50 amino acids, continuous or with gaps in the complete amino acid sequence of the polypeptide, or the full amino acid sequence of the polypeptide.

[0068] As used herein, the term "sample" refers to a collected substance or material that contains or is suspected to contain one or more analytes of interest (e.g., biomolecules, such as polypeptides). The sample can be modified for purposes such as storage or stability. The sample can be naturally occurring or synthetic. The sample can be processed to separate or remove unwanted fractions or impurities from the analyte of interest. The sample can be enriched or purified. For example, the sample can include fractions of a separation process (e.g., chromatography, fractionation, electrophoresis, etc.). Alternatively, the sample may not be processed to separate or remove any unwanted fractions or impurities from the analyte of interest. The sample can be obtained from any suitable source or location, including from an organism, a cell, a tissue, a cell preparation, a cell-free composition, an environment (e.g., air, water, dirt, soil, agriculture, soil, dust, sewage). The sample can be obtained from an organism or a part of an organism, such as from a fluid, a tissue, or a cell. The sample can include biological and / or non-biological components. As used herein, the term "biological sample" or "biological source" refers to a sample derived from a primarily biological system or organism, such as one or more viral particles, cells (e.g., personalized cells), organelles (e.g., personalized organelles), tissues, body fluids, bones, cartilage, and exoskeletons. A biological sample can comprise most biological materials based on mass, excluding the weight of the fluid within the sample. A biological sample can comprise one or more proteins, referred to herein as a protein sample. A biological sample can be obtained from various sources, for example, from clinical patient samples, such as blood, serum, plasma, cerebrospinal fluid (CSF), saliva, mucosal secretions, sputum, urine, lymph, sweat, vaginal fluid, semen, feces, amniotic fluid, sweat, synovial fluid, etc. A biological sample can be processed to purify and retain one or more biomolecules (e.g., proteins, nucleic acids, carbohydrates, lipids, glycoproteins, lipoproteins, metabolites, etc.) from a biological sample. A biological sample (e.g., a protein sample) can be derived from cultured cells, which can be treated or untreated. Biological samples (e.g., protein samples) can also be generated from tissue samples, such as biopsy samples, which can optionally be treated to release the biomolecules (e.g., proteins) they contain. Tissue samples can also be derived from in vivo samples, including fresh, frozen, emergency, and fixed tissues.

[0069] As used herein, the terms "antibody" and "immunoglobulin" may generally refer to proteins that can recognize and bind to specific antigens. Antibodies or immunoglobulins may refer to fragments of antibody isotypes, antibodies, including but not limited to Fab, Fv, scFv and Fd fragments, chimeric antibodies, humanized antibodies, single-chain antibodies, and fusion proteins comprising the antigen-binding portion of an antibody and non-antibody proteins. Antibodies may be detectably labeled, for example, using fluorophores, radioisotopes, enzymes (such as peroxidases) that generate detectable products, fluorescent proteins, nucleic acid barcode sequences, etc. Antibodies may also be conjugated to other moieties, such as members of specific binding pairs, for example biotin (members of biotin-avidin specific binding pairs), etc. The term also includes nanobodies, Fab', Fv, F(ab')2, scFv, and other antibody fragments that retain specific binding to an antigen. Antibodies can exist in a variety of other forms, including, for example, Fv, Fab, and (Fab)2, as well as bifunctional (e.g., bispecific) hybrid antibodies (e.g., Lanzavecchia et al., Eur. J. Immunol. 17, 105 (1987)) and single-chain (e.g., Huston et al., Proc. Natl. Acad. Sci. USA, 85, 5879-5883 (1988) and Bird et al., Science, 242, 423-426 (1988), which are incorporated herein by reference). (See generally Hood et al., Immunology, Benjamin, NY, 2nd ed. (1984) and Hunkapiller and Hood, Nature, 323, 15-16 (1986), which are incorporated herein by reference). Naturally occurring immunoglobulin or antibody classes include immunoglobulin A, immunoglobulin G, immunoglobulin D, immunoglobulin E, immunoglobulin M, or other immunoreactive components.

[0070] As used herein, "binding" or "coupling" generally refers to a covalent or non-covalent interaction between two molecules (referred to herein as "binding partners," e.g., a substrate and an enzyme or an antibody and an epitope). The binding between the binding partners can be specific or non-specific.

[0071] As used herein, "specifically binds" or "binds specifically" generally refers to an interaction between binding partners (e.g., a binding partner and a cognate molecule) such that the binding partner binds to another partner but does not bind to other molecules that may be present in the environment (e.g., in a biological sample, in a tissue, in an in vitro assay) under a set of conditions. Specific binding interactions can involve a binding partner that binds to a cognate molecule. Specific binding interactions can involve a binding partner that binds to its cognate molecule at a significantly or substantially higher level or with greater affinity than the binding of the binding partner to a non-cognate molecule. Specific binding interactions can involve a first binding partner that has greater selectivity for binding to a cognate molecule than to a non-cognate molecule.

[0072] The terms "nucleic acid," "nucleic acid molecule," "oligonucleotide," and "polynucleotide" are used interchangeably herein and generally refer to a polymeric form of naturally occurring or synthetic nucleotides or their analogs of any length. A nucleic acid molecule can comprise one or more deoxynucleotides, deoxynucleotide triphosphates, dideoxynucleotide triphosphates, ribonucleotides, hexitol nucleotides, cyclohexane nucleotides, or analogs or combinations thereof. Nucleic acid molecules can include, for example, DNA, RNA, HNA, CeNA, and modified forms thereof. Nucleic acid molecules can comprise nucleotides linked by phosphodiester bonds. Nucleic acid molecules can have any two-dimensional or three-dimensional structure and can perform any known or unknown function. Nucleic acid molecules can be single-stranded, double-stranded, or partially double-stranded. Non-limiting examples of polynucleotides include genes, gene fragments, exons, introns, messenger RNA (mRNA), transfer RNA, ribosomal RNA, non-coding DNA, small interfering RNA, short hairpin RNA, microRNA, scaRNA, ribozymes, riboswitches, viral RNA, complementary DNA (cDNA), cosmid DNA, mitochondrial DNA, chromosomal or genomic DNA, viral DNA, recombinant polynucleotides, branched polynucleotides, plasmids, vectors, isolated DNA of any sequence, control regions, isolated RNA of any sequence, nucleic acid probes, nucleic acid adapters, and primers. Nucleic acid molecules can be linear, circular, or any other geometric shape. Examples of polynucleotide analogs include, but are not limited to, xenonucleic acid (XNA), bridged nucleic acid (BNA), glycol nucleic acid (GNA), hexitol nucleic acid (HNA), cyclohexane nucleic acid (CeNA), 2'-F-arabinose nucleic acid (2'-F-ANA), peptide nucleic acid (PNA), γPNA, morpholino polynucleotides, locked nucleic acid (LNA), threose nucleic acid (TNA), 2'-O-methyl polynucleotides, 2'-O-alkylribosyl substituted polynucleotides, phosphorothioate polynucleotides, and boronate polynucleotides. Polynucleotide analogs can have purine or pyrimidine analogs, including, for example, 7-deazapurine analogs, 8-halogenated purine analogs, 5-halogenated pyrimidine analogs, or universal base analogs that can pair with any base, including hypoxanthine, nitroazoles, isoquinolone analogs, azole carboxamides, and aromatic triazole analogs, or base analogs with additional functionality, such as a biotin moiety for affinity binding.

[0073] As used herein, the term "amino acid" generally refers to an organic compound that is combined to form a protein or peptide. Amino acid generally comprises an amine group, a carboxylic acid group, and a side chain specific to each amino acid, which serves as a monomeric subunit of a peptide. Amino acid can include 20 standards, naturally occurring or typical amino acids and non-standard amino acids. Standard, naturally occurring or typical amino acids include alanine (A or Ala), cysteine ​​(C or Cys), aspartic acid (D or Asp), glutamic acid (E or Glu), phenylalanine (F or Phe), glycine (G or Gly), histidine (H or His), isoleucine (I or Ile), lysine (K or Lys), leucine (L or Leu), methionine (M or Met), asparagine (N or Asn), proline (P or Pro), glutamine (Q or Gln), arginine (R or Arg), serine (S or Ser), threonine (T or Thr), valine (V or Val), tryptophan (W or Trp), and tyrosine (Y or Tyr). The amino acid can be an L-amino acid or a D-amino acid. Non-standard amino acids can be naturally occurring or chemically synthesized modified amino acids, amino acid analogs, amino acid mimetics, non-standard protein amino acids, or non-protein amino acids. Examples of non-standard amino acids include, but are not limited to, selenocysteine, pyrrolysine, and N-formylmethionine, β-amino acids, homoamino acids, proline and pyruvate derivatives, β-substituted alanine derivatives, glycine derivatives, ring-substituted phenylalanine and tyrosine derivatives, linear core amino acids, and N-methyl amino acids.

[0074] As used herein, the term "amino acid type" generally refers to one of the standard, naturally occurring or typical amino acids, for example, a member of the group consisting of alanine (A or Ala), cysteine ​​(C or Cys), aspartic acid (D or Asp), glutamic acid (E or Glu), phenylalanine (F or Phe), glycine (G or Gly), histidine (H or His), isoleucine (I or Ile), lysine (K or Lys), leucine (L or Leu), methionine (M or Met), asparagine (N or Asn), proline (P or Pro), glutamine (Q or Gln), arginine (R or Arg), serine (S or Ser), threonine (T or Thr), valine (V or Val), tryptophan (W or Trp), tyrosine (Y or Tyr), derivatives thereof, and modified forms of any of the above amino acids. The term "amino acid type" can be used herein to distinguish between multiple amino acids containing different side chain groups, rather than multiple identical amino acids (e.g., amino acids at different positions in a single peptide having the same side chain).

[0075] As used herein, the term "post-translational modification" refers to a modification that occurs on a peptide after translation. A post-translational modification can be a covalent modification or an enzymatic modification. Examples of post-translational modifications include, but are not limited to, acylation, acetylation, alkylation (including methylation), biotinylation, butyrylation, carbamylation, carbonylation, deamidation, deimination, diphtheria amide formation, disulfide bridge formation, elimination, flavin attachment, formylation, γ-carboxylation, glutamylation, glycination, glycosylation, glycosylphosphatidylinositolization, heme C attachment, hydroxylation, hydroxybutyrylation, iodination, isoprene The post-translational modifications include modifications of the amino and / or carboxyl termini of the peptide. Modifications of the terminal amino group include, but are not limited to, deamination, N-low alkyl, N-di-low alkyl, and N-acyl modifications. Modifications of the terminal carboxyl group include, but are not limited to, amides, lower alkyl amides, dialkyl amides, and lower alkyl ester modifications (e.g., wherein the lower alkyl group is a C1-C4 alkyl). Post-translational modifications also include modifications to the amino acids falling between the amino and carboxyl termini, such as, but not limited to, those described above. The term post-translational modifications can also include modifications of peptides that include one or more detectable labels. Post-translational modifications can be naturally occurring or synthetic.

[0076] As used herein, the term "binding agent" refers to a molecule that binds, associates, unites, recognizes or combines another molecule, such as a nucleic acid molecule, peptide, polypeptide, protein, carbohydrate, synthetic molecule or small molecule. A binding agent can be bound to a component or feature of a macromolecule or a macromolecule. A binding agent can form a covalent association or non-covalent association with a molecule, a macromolecule or a component or feature of a macromolecule. A binding agent can also be a chimeric binding agent composed of two or more types of molecules, such as a nucleic acid molecule-peptide chimeric binding agent, a carbohydrate-peptide chimeric binding agent or a lipid-peptide chimeric binding agent. A binding agent can be a naturally occurring, synthetically produced or recombinantly expressed molecule. A binding agent can be bound to a single monomer or subunit (e.g., a single amino acid of a peptide) of a polymer analyte such as a macromolecule or to a plurality of connected subunits (e.g., a dipeptide, a tripeptide or a higher order peptide of a longer peptide, polypeptide or protein molecule) of a macromolecule. A binding agent can be bound to a linear molecule or a molecule with a three-dimensional structure (also referred to as a conformation). For example, the antibody binding agent can bind to a linear peptide, polypeptide or protein, or to a conformational peptide, polypeptide or protein. The binding agent can bind to the N-terminal peptide, C-terminal peptide or intermediate peptide of a peptide, polypeptide or protein molecule. The binding agent can bind to the N-terminal amino acid, C-terminal amino acid or intermediate amino acid of a peptide molecule. The binding agent can preferably bind to chemically modified or labeled amino acids, compared to unmodified or unlabeled amino acids. For example, the binding agent can preferably bind to amino acids that have been modified to have an acetyl moiety, an amidino moiety, a dansyl moiety, a PTC moiety, a DNP moiety, a SNP moiety, etc., compared to amino acids that do not have such moieties. The binding agent can bind to naturally occurring or synthetic post-translational modifications of the peptide molecule. The binding agent can exhibit selective binding to a component or feature of a macromolecule (for example, the binding agent can selectively bind to one of 20 possible natural amino acid residues and bind to the other 19 natural amino acid residues with very low affinity or not bind to them at all). A binding agent may exhibit less selective binding, where the binding agent is able to bind to multiple components or features of a macromolecule (e.g., a binding agent may bind with similar affinity to two or more different amino acid residues). A binding agent may comprise a tag, which may be coupled to the binding agent via a linker.

[0077] As used herein, the term "joint" generally refers to a molecule or part that is involved in joining two or more molecules. A joint can promote the covalent or non-covalent interaction of two or more molecules. A joint can be a cross-linking agent. A joint can be monofunctional, bifunctional, trifunctional, tetrafunctional or multifunctional. A joint can be or comprise a nucleotide, a nucleotide analog, an amino acid, a peptide, a polypeptide or a non-nucleotide chemical moiety, such as an organic or inorganic compound. A joint can comprise a polymer, such as polyethylene glycol (PEG), poly-L-lysine (PLL), poly (DL-lactic acid) (PLA), poly (DL-lactic-co-glycolic acid) copolymer (PLGA), polyornithine, polyarginine, etc. A joint can comprise one or more reactive ends, such as an amine reactive group, a carboxyl reactive group, a sulfhydryl reactive group, a hydroxyl reactive group, etc. In some examples, a joint can be used to join different molecular types, for example, different biomolecule types (such as peptides and nucleic acid molecules, lipids and peptides, carbohydrates and peptides, etc.); non-biological molecule types; or biomolecules to non-biological molecules. For example, a linker can be used to join a binding agent to a tag, a tag to a macromolecule (e.g., a peptide, a nucleic acid molecule), a macromolecule to a solid support, a tag to a solid support, etc. A linker can join two molecules via an enzymatic reaction or a chemical reaction (e.g., click chemistry). A linker can join more than two molecules, for example, via an enzymatic or chemical reaction.

[0078] As used herein, the term "conjugate" generally refers to a covalent or ionic interaction between two entities (eg, molecules, compounds, or a combination thereof).

[0079] As used herein, the term "label" generally refers to a molecule or part that is conjugated to a molecule. Labels can include detectable markers, for example, fluorophores or fluorescent proteins, radioisotopes, enzymes (for example, color production or fluorescent proteins, proteins that can catalyze color production substrates), mass labels, haptens (for example, biotin, digoxin, urushiol, fluorescein), vibration or FTIR labels (for example, alkynyl). Labels can include biomolecules, such as nucleic acid molecules, proteins, lipids, carbohydrates or combinations thereof. Labels can include one or more nucleic acid molecules that can optionally encode information about a label or a label conjugated to a molecule thereon (for example, a binding agent, such as an antibody). For example, labels can include nucleic acid barcode molecules. Labels can include organic compounds or inorganic compounds.

[0080] As used herein, the term "barcode" generally refers to an identifying feature that can be used to distinguish between similar items. 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100, 101, 102, 103, 104, 105, 106, 107, 108, 109 Barcodes can be artificial or naturally occurring sequences, including peptides, polypeptides, peptides, binding agents, binding agent collections from binding cycles, sample molecules, sample collections, molecules within compartments (e.g., droplets, beads, partitions, or separate locations), macromolecules within a compartment collection, fractions of macromolecules, collections of macromolecule fractions, spatial regions or collections of spatial regions, libraries of macromolecules, or libraries of binding agents. In certain embodiments, each barcode within a barcode population is different. In other embodiments, a portion of the barcodes in a barcode population are different, for example, at least about 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 97% or 99% of the barcodes in the barcode population are different. The barcode population can be generated randomly or non-randomly. The barcode population can include error correction barcodes. The barcodes can be used to computationally deconvolute sequence reads derived from separate molecules, samples, libraries, etc. The barcodes can contain, for example, multiplexed information generated by different samples, compartments, separate molecules, etc. The barcodes can also be used to deconvolve a collection of molecules that have been distributed into small compartments for enhanced mapping. For example, instead of mapping peptides back to the proteome, peptides can be mapped back to the protein molecule or protein complex from which they originated.Barcodes can contain any useful structural part or motif, such as a hairpin, loop sequence, or spacer.The barcode can comprise an artificial or modified nucleic acid, e.g., locked nucleic acid (LNA), protein nucleic acid (PNA), hexitol nucleic acid (HNA), cyclohexane nucleic acid (CeNA), or a combination thereof. The barcode can comprise a protein (e.g., a Tal effector, a Cas protein (e.g., Cas9), Argonaut, or coiled-coil) or be generated using a protein.

[0081] As used herein, a "sample barcode," also referred to as a "sample tag," generally refers to a barcode molecule that contains identification information of a sample from which the barcoded molecule is derived.

[0082] As used herein, "spatial barcode" generally refers to a barcode molecule that contains identification information for a 2D or 3D sample (e.g., a tissue section) region from which the molecule originates or is derived. Spatial barcodes can be used for molecular pathology on tissue sections. Spatial barcodes can allow multiplexed sequencing of multiple samples or libraries from tissue sections.

[0083] As used herein, "temporal barcode" generally refers to a barcode molecule that contains time-based information associated with the barcoded molecule. The time-based data types encoded in the temporal barcode can include information such as the lifetime of the barcoded molecule, the time of sample collection, the time or duration since the start of the experiment or induction with a stimulus, information about the age of the cell or tissue, the sequence of interactions between molecules, etc. It is possible to combine different types of barcodes (e.g., spatial, temporal, cell-specific) in a single multiplexed barcode.

[0084] As used herein, the term "nucleic acid sequence" or "oligonucleotide sequence" generally refers to a continuous string of nucleotide bases, and may refer to the specific placement of the nucleotide bases relative to each other when they occur in an oligonucleotide. Similarly, the term "polypeptide sequence" or "amino acid sequence" refers to a continuous string of amino acids, and may refer to the specific placement of the amino acids relative to each other when they occur in a polypeptide.

[0085] According to the present invention, " nucleic acid molecule " can include any polymer or oligomer of nucleotide, such as pyrimidine and purine bases, such as cytosine, thymine and uracil and adenine and guanine, and combination thereof. The nucleotide sequence can include any deoxyribonucleotide, ribonucleotide, hexitol nucleotide, cyclohexane nucleotide, peptide nucleic acid component and any chemical variant thereof, such as methylation, 7-deazapurine analogs, 8-halogenated purine analogs, the hydroxymethylation or glycosylation form of these bases etc. Polymer or oligomer can be heterogeneous or homogeneous in composition, and can be separated from naturally occurring sources or can be artificially or synthetically produced. Nucleic acid molecule can include DNA, RNA, HNA, CeNA or its mixture, and can exist permanently or transitionally in single-stranded or double-stranded form (including homoduplex, heteroduplex and hybridization state).

[0086] The terms "complementary" or "complementarity" refer to polynucleotides (i.e., sequences of nucleotides) that are related by the Watson-Crick base pairing rules. For example, the sequence "5'-AGT-3'" is complementary to the sequence "5'-ACT-3'". Complementarity can be "partial", in which only some of the bases of the nucleic acids match according to the base pairing rules, or there can be "complete" or "perfect" complementarity between nucleic acids. The degree of complementarity between nucleic acid strands can have a significant effect on the efficiency and strength of hybridization between nucleic acid strands under defined conditions.

[0087] As used herein, the term "hybridization" is used in reference to the pairing of complementary nucleic acids. Hybridization and hybridization strength (e.g., the strength of association between nucleic acids) are affected by factors such as the degree of complementarity between the nucleic acids, the stringency of the conditions involved, and the melting temperature of the hybrid formed. Hybridization methods involve annealing one nucleic acid to another complementary nucleic acid, for example, based on Watson-Crick base pairing.

[0088] As used herein, the term "proteomics" generally refers to the quantitative and / or qualitative analysis of a proteome within a sample (such as a biological sample, e.g., from a cell, tissue, or body fluid). Proteomics can include analysis of the spatial distribution of proteins within a sample (e.g., a cell and / or tissue). Proteomics can include studying the dynamic state of a proteome, e.g., how one or more proteins change over time. A proteome can include multiple "-groups," such as a kinase group, a secretome, a receptor group (e.g., a GPCome), an immune proteome, a nutritional proteome, a subset of proteomes defined by post-translational modifications (e.g., phosphorylation, ubiquitination, methylation, acetylation, glycosylation, oxidation, lipidation, and / or nitrosylation), such as a phosphorylated proteome (e.g., a phosphotyrosine proteome, a tyrosine kinase group, and a tyrosine phosphate group), a glycoprotein group, etc., a subset of proteomes associated with a tissue or organ, a developmental stage, or a physiological or pathological condition, a subset of proteomes associated with a cellular process such as a cell cycle, differentiation (or dedifferentiation), cell death, aging, cell migration, transformation, or metastasis, or any combination thereof.

[0089] The terminal amino acid with a free amino group at one end of the peptide chain may be referred to as "N-terminal amino acid" (NTAA) in this article. The terminal amino acid with a free carboxyl group at the other end of the chain may be referred to as "C-terminal amino acid" (CTAA) in this article. The amino acids constituting the peptide can be numbered in sequence, where the length of the peptide is "n" amino acids. As used herein, in some cases, NTAA may be considered to be the nth amino acid (also referred to as "n NTAA" in this article). In such cases, the next amino acid is the n-1th amino acid, followed by the n-2th amino acid, so reducing the length of the peptide from the N-terminal to the C-terminal. Alternatively, CTAA may be considered to be the nth amino acid (also referred to as "n CTAA" in this article). In such cases, the next amino acid is the n-1th amino acid, followed by the n-2th amino acid, so reducing the length of the peptide from the C-terminal to the N-terminal. NTAA, CTAA or both may be modified or labeled with chemical moieties.

[0090] As used herein, the terms "determine," "measure," "evaluate," and "assay" are used interchangeably and include both quantitative and qualitative determinations.

[0091] As used herein, the term "unique molecular identifier" or "UMI" generally refers to a molecular barcode that contains indexing information. A UMI can comprise a length of about 3 to about 150 bases (3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85 , 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100, 105, 110, 115, 120, 125, 130, 135, 140, 145, or 150 bases). A UMI can provide a unique identifier tag for each molecule (e.g., a peptide, a binding agent, a nucleic acid molecule) that comprises or is conjugated to a UMI. A UMI can include a random sequence (e.g., a random N-mer).

[0092] As used herein, " derivative " of nucleic acid molecules refers to nucleic acid molecules derived from origin nucleic acid molecules in general. Derivative can have the nucleotide sequence identical or substantially identical with origin nucleic acid molecules, or derivative can comprise complement or partial complement as origin nucleic acid molecules. Derivative can be nucleic acid (for example, DNA or RNA) of the same type as origin nucleic acid molecules, or derivative can be nucleic acid of different types (for example, cDNA generated from RNA molecules). Nucleic acid molecule derivative can show sequence identity as origin nucleic acid molecules. Derivative nucleic acid molecules can also be carried out additional processing from origin nucleic acid molecules, for example, chemical or enzymatic modification, splicing, connection, polymerization, fragmentation, tag cutting (tagmentation) (for example, using transposase), digestion etc.

[0093] 1. Derivative polypeptides or peptides can be derived from an origin polypeptide (or peptide). The derivative can comprise the same amino acid sequence as the origin polypeptide, or the sequence can be different. The derivative polypeptide can be produced by the origin polypeptide or can be subjected to additional processing, such as chemical or enzymatic modification, to the derivative polypeptide from the origin polypeptide. The derivative polypeptide can include one or more tags, nucleic acid molecules, barcode molecules, labels (e.g., detectable labels), fluorophores, probes, linkers, post-translational modifications, chemical protecting groups or other chemical moieties.

[0094] As used herein, the term "compartment" or "partition" generally refers to a physical area or volume that separates or isolates a subset of molecules from a sample of molecules. For example, a compartment can separate individual cells from other cells, or separate a subset of the proteome of a sample from the rest of the proteome of the sample. A compartment can be an aqueous compartment (e.g., a microfluidic droplet), a solid compartment (e.g., a picotiter well or microtiter well on a plate, tube, vial, gel bead), or an isolated area on a surface. A compartment can contain one or more beads to which a macromolecule can be affixed.

[0095] In some embodiments, the partitions as described herein are droplets. The terms "drop," "droplet," and "microdroplet" are used interchangeably herein to refer to small, generally spherical structures containing at least a first fluid phase (e.g., an aqueous phase (e.g., water)) bounded by a second fluid phase (e.g., oil) that is immiscible with the first fluid phase. In some embodiments, droplets according to the present disclosure may contain a first fluid phase (e.g., oil) bounded by a second immiscible fluid phase (e.g., an aqueous phase fluid (e.g., water)). In some embodiments, the second fluid phase will be an immiscible phase carrier fluid. Thus, droplets according to the present disclosure can be provided as water-in-oil emulsions or oil-in-water emulsions. As described herein for discrete entities, droplets can be sized and / or shaped. For example, the diameter range of droplets according to the present disclosure is generally 1 μm to 1000 μm, inclusive. Droplets according to the present disclosure can be used to encapsulate cells, nucleic acids (e.g., DNA), enzymes, reagents, and various other components. The term droplet may be used to refer to droplets generated in, on, by, and / or flowing from, or applied by, a microfluidic device.

[0096] As used herein, the term "carrier fluid" refers to a fluid configured or selected to contain one or more discrete entities (e.g., droplets), as described herein. The carrier fluid can include one or more substances and can have one or more properties, such as viscosity, that allow it to flow through the microfluidic device or a portion thereof, such as a delivery orifice. In some embodiments, the carrier fluid includes, for example, oil or water, and can be in a liquid or gas phase. Suitable carrier fluids are described in more detail herein.

[0097] As used herein, the terms "solid support," "solid surface," or "solid substrate" or "substrate" refer to any solid material to which molecules can associate directly or indirectly, including porous and non-porous materials. Molecules can associate with a substrate through covalent or non-covalent interactions, or a combination thereof. The substrate can be two-dimensional (e.g., a planar surface) or three-dimensional (e.g., a gel matrix or beads). In non-limiting examples, the solid support can include beads, microbeads, arrays, glass surfaces, silicon surfaces, plastic surfaces, filters, membranes, nylon or other polymers, silicon wafer chips, flow-through chips, flow cells, biochips including signal transduction electronic devices, channels, microtiter wells, ELISA plates, rotating interferometry disks, nitrocellulose membranes, nitrocellulose-based polymer surfaces, polymer matrices, nanoparticles, or microspheres. The material for solid support includes but is not limited to acrylamide, agarose, cellulose, nitrocellulose, glass, gold, quartz, polystyrene, polyethylene vinyl acetate, polypropylene, polymethacrylate, polyethylene, polyethylene oxide, polysilicate, polycarbonate, Teflon, fluorocarbon, nylon, silicone rubber, polyanhydride, polyglycolic acid, polylactic acid, polyorthoester, functionalized silane, polypropyl fumarate, collagen, glycosaminoglycan, polyamino acid, dextran or its any combination.Solid support further includes film, membrane, bottle, dish, fiber, braided fiber, shaped polymer such as pipe, particle, bead, microsphere, particulate or its any combination.For example, when solid surface is bead, bead can include but is not limited to ceramic beads, polystyrene beads, polymer beads, methyl styrene beads, agarose beads, acrylamide beads, solid beads, porous beads, magnetic or paramagnetic beads, glass beads or controlled pore beads.Bead can be spherical or irregularly shaped. The beads can range in size from nanometers (e.g., 100 nm) to millimeters (e.g., 1 mm). In certain embodiments, the beads range in size from about 0.2 microns to about 200 microns or from about 0.5 microns to about 5 microns.In some embodiments, the diameter of the beads can be about 1, 1.5, 2, 2.5, 2.8, 3, 3.5, 4, 4.5, 5, 5.5, 6, 6.5, 7, 7.5, 8, 8.5, 9, 9.5, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65 94, 95, 96, 97, 98, 99, or 100 μm. In certain embodiments, a "bead" solid support can refer to a single bead or a plurality of beads.

[0098] As used herein, " sequencing " generally refers to determining: the nucleotide (base sequence) order in (A) nucleic acid sample (for example, DNA or RNA); Or determining the amino acid sequence in all or part of (B) polymer (such as protein, peptide or other polymer molecules). Many technologies are available, such as Sanger sequencing or high-throughput sequencing technology (HTS). Sanger sequencing can involve sequencing via detection by (capillary) electrophoresis, in which a run can carry out sequence analysis to up to 384 capillaries. High-throughput sequencing involves the parallel sequencing of thousands or millions or more sequences. HTS can be defined as next generation sequencing (NGS), that is, based on solid phase pyrophosphate sequencing technology or based on the next generation-next generation sequencing technology of single nucleotide real-time sequencing (SMRT). HTS technology is available, such as provided by Roche, Illumina and Applied Biosystems (Life Technologies). Additional high-throughput sequencing technologies are described by and / or available from: Helicos, Pacific Biosciences, Complete Genomics, Ion Torrent Systems, Oxford Nanopore Technologies, Nabsys, ZS Genetics, GnuBio.

[0099] As used herein, " next generation sequencing " refers to a high-throughput sequencing method that allows parallel sequencing of millions to billions of molecules. The example of next generation sequencing method includes synthetic sequencing, connection sequencing, hybridization sequencing, polymerase colony sequencing (polony sequencing), ion semiconductor sequencing, nanopore sequencing and pyrophosphate sequencing. By attaching primers to a solid substrate and attaching complementary sequences to nucleic acid molecules, nucleic acid molecules can hybridize with the solid substrate via primers, and then multiple copies can be generated in discrete areas on the solid substrate by using polymerase to amplify (these groupings are sometimes referred to as polymerase colonies or polymerase colonies). Therefore, in the sequencing process, the nucleotides at a specific position can be sequenced multiple times (for example, hundreds or thousands of times)--this coverage depth is referred to as "deep sequencing". Examples of high-throughput nucleic acid sequencing technologies include platforms offered by Illumina, BGI, Qiagen, ThermoFisher, and Roche, including formats such as parallel bead arrays, sequencing by synthesis, sequencing by ligation, capillary electrophoresis, electronic microchips, "biochips," microarrays, parallel microchips, and single-molecule arrays, as reviewed by Service (Science 311: 1544-1546, 2006).

[0100] As used herein, "analyzing" macromolecules refers to quantifying, characterizing, distinguishing, or combining all or part of the components of a molecule (e.g., a macromolecule, a biomolecule such as a protein, an amino acid, a nucleic acid molecule, etc.). For example, analyzing a peptide, a polypeptide, or a protein can include determining all or part of the amino acid sequence (continuous or discontinuous) of the peptide. Analyzing macromolecules can include partially identifying the components of the macromolecule. For example, the partial identification of amino acids in a protein sequence can identify the amino acids in the protein as belonging to a possible subset of amino acids. The analysis can be performed sequentially, for example, starting with the analysis of n NTAAs, and then proceeding to the next amino acid of the peptide (i.e., n-1, n-2, n-3, etc.). In such cases, sequencing can be performed by cutting n NTAAs, thereby converting the n-1 amino acid of the peptide into an N-terminal amino acid (referred to herein as "n-1 NTAA"). Similarly, analyzing peptides can start from the C-terminus to the N-terminus, and each round of cutting from the C-terminus produces a new CTAA. The cutting of n CTAA converts the n-1 amino acid of the peptide into a C-terminal amino acid, which is referred to herein as "n-1 CTAA." Analyzing the peptides can also include determining the presence and frequency of post-translational modifications on the peptides, which may or may not include information about the sequential order of the post-translational modifications on the peptides. Analyzing the peptides can also include determining the presence and frequency of epitopes in the peptides, which may or may not include information about the sequential order or position of the epitopes within the peptides. Analyzing the peptides can include combining different types of analyses, for example, to obtain epitope information, amino acid sequence information, post-translational modification information, or any combination thereof.

[0101] As used herein, the term "analyte" generally refers to a substance of interest for further identification, characterization, or measurement. In a non-limiting example, an analyte can be an ion, chemical, compound, small molecule, element, particle, metal, biomolecule, macromolecule, metabolite, lipid, carbohydrate, peptide or protein, nucleic acid molecule, organelle, or cell. The analyte can be naturally occurring or synthetic. The analyte can be a solid, semi-solid, liquid, semi-liquid, gas, or plasma. The analyte can be characterized qualitatively or quantitatively. A portion of the analyte can be analyzed. For example, the analyte can be a peptide, and the constituent amino acids can be analyzed. The analyte can include a polymer, also referred to herein as a "polymer analyte," which generally refers to an analyte of interest comprising one or more monomers. In a non-limiting example, a polymer analyte can be a group of ions, chemicals, compounds, small molecules, elements, particles, metals, or biomolecules, macromolecules, metabolites, lipids, carbohydrates, peptides or proteins, nucleic acid molecules, organelles, or cells.

[0102] As used herein, the term "array" generally refers to a population of molecules attached to one or more solid supports so that molecules at one address can be distinguished from molecules at other addresses. An array can include different molecules at different addresses, each located on a solid support. Alternatively, an array can include separate solid supports, each acting as an address with different molecules, wherein the different molecules can be identified based on the position of the solid support on the surface to which the solid support is attached or based on the position of the solid support in a liquid such as a fluid stream. The molecules of the array can be, for example, nucleic acids (such as SNAPs), polypeptides, proteins, peptides, oligopeptides, enzymes, ligands, or receptors (such as antibodies, functional fragments of antibodies, or aptamers). The addresses of the array can optionally be optically observable, and in some configurations, adjacent addresses can be optically distinguishable when detected using the methods or devices set forth herein.

[0103] As used herein, the term "functionalized" refers to any material or substance that has been modified to include a functional group. Functionalized materials or substances can be natural or synthetic functionalized. For example, polypeptides can be naturally functionalized with phosphate groups, oligosaccharides (e.g., glycosyl, glycosylphosphatidylinositol, or phosphoglycosyl), nitrosyl, methyl, acetyl, lipids (e.g., glycosylphosphatidylinositol, myristoyl, or isoprenyl), ubiquitin, or other naturally occurring post-translational modifications. Functionalized materials or substances can be functionalized for any given purpose, including changing chemical properties (e.g., changing hydrophobicity or changing surface charge density) or changing reactivity (e.g., being able to react with a moiety or reagent to form a covalent bond with the moiety or reagent).

[0104] As used herein, the term "click reaction", "click chemistry" or "bioorthogonal reaction" refers to a single-step, thermodynamically favorable conjugation reaction using biocompatible reagents. The click reaction can not use toxic or bioincompatible reagents (e.g., acids, bases, heavy metals) or does not generate toxic or bioincompatible byproducts. The click reaction can use an aqueous solvent or buffer (e.g., phosphate buffered solution, Tris buffer, saline buffer, MOPS, etc.). The click reaction can be thermodynamically favorable if it has a negative reaction Gibbs free energy, for example, less than about -5 kilojoules / mole (kJ / mol), -10 kJ / mol, -25 kJ / mol, -50 kJ / mol, -100 kJ / mol, -200 kJ / mol, -300 kJ / mol, -400 kJ / mol or less than -500 kJ / mol. Exemplary bioorthogonal and click reactions are described in detail in WO 2019 / 195633A1, which is incorporated herein by reference in its entirety. Exemplary click reactions can include metal-catalyzed azide-alkyne cycloadditions, strain-promoted azide-alkyne cycloadditions, strain-promoted azide-nitrone cycloadditions, strained ene reactions, thiol-ene reactions, Diels-Alder reactions, inverse electron demand Diels-Alder reactions, [3+2] cycloadditions, [4+1] cycloadditions, nucleophilic substitutions, dihydroxylation reactions, thiol-alkyne reactions, photoclick reactions (photoclick), nitrone dipolar cycloadditions, norbornene cycloadditions, oxa-norbornadiene cycloadditions, tetrazine connections, and tetrazole photoclick reactions.Exemplary functional groups or reactive handles for performing click reactions can include alkenes (e.g., linear alkenes or cyclic alkenes such as trans-cyclooctene (TCO)), alkynes (e.g., linear alkynes or cyclic alkynes (e.g., cyclooctynes ​​or derivatives thereof such as aza-dimethoxycyclooctyne (DIMAC), symmetrical pyrrolocyclooctyne (SYPCO), pyrrolocyclooctyne (PYRROC), difluorocyclooctyne (DIFO), α,α-bis(trifluoromethyl)pyrrolocyclooctyne (TRIPCO), bicyclo[6.1.0]nonyne (BCN), dibenzocyclooctyne (DIBO), difluorinated cyclooctyne (DIFO), difluorobenzocyclooctyne (DIFBO), dibenzoazacyclooctyne (DBCO), difluoro-azacyclooctyne (DIFO), difluorobenzocyclooctyne (DIFBO), difluoro-azacyclooctyne (DBCO), difluoro-azacyclooctyne (DIFO), difluoro-aza ... Hetero-dibenzocyclooctyne (F2-DIBAC), biaryl-azacyclooctyne ketone (BARAC), difluorodimethoxydibenzocyclooctyne alcohol (FMDIBO), difluorodimethoxydibenzocyclooctyne ketone (keto-FMDIBO) and 3,3,6,6-tetramethylthioheptyne (TMTH)), TMTH-sulfoximine (TMTHSI), azides, epoxides, amines, thiols, nitrones, isonitriles, isocyanides, aziridines, activated esters and tetrazines, triazoles, and combinations, variants or derivatives thereof. The click chemistry moieties can be subjected to conditions sufficient to react the first click chemistry moiety with the second click chemistry moiety, for example, providing a metal catalyst, a suitable solvent, pH, temperature, ion concentration or light / energy for any useful duration.

[0105] As used herein, the terms "group" and "moiety" are intended to be synonymous when used to refer to the structure of a molecule. The terms refer to a component or portion of a molecule. Unless otherwise indicated, the terms do not necessarily indicate the relative size of the component or portion compared to the molecule. Unless otherwise indicated, the terms do not necessarily indicate the relative size of the component or portion compared to any other component or portion of the molecule. A group or moiety may contain one or more atoms.

[0106] As used herein, "primer" generally refers to a nucleic acid molecule that can trigger the synthesis of a nucleic acid molecule (e.g., DNA or RNA). A primer can be single-stranded. A primer can include one or more recognition sites for binding of a protein (e.g., polymerase, restriction enzyme, cutting enzyme, nuclease, etc.) to a primer or to a primer hybridized with a template strand. A primer can include DNA, RNA, or other nucleic acid analogs or atypical bases (e.g., spacer portion, uracil, abasic site). A primer can optionally include any number of functional sequences, such as sequencing primer sequences (e.g., P5 or P7 sequences), sequencing primer binding sequences, read sequences (e.g., R1 or R2 sequences), restriction sites, nuclease recognition sites, abasic sites, cutting sites, transposition sites, barcode sequences, unique molecular identifiers (UMIs), etc.

[0107] "Amplify" or "amplify" generally refers to a polynucleotide amplification reaction, i.e., a population of polynucleotides replicated from one or more starting sequences. Amplification can refer to various amplification reactions, including but not limited to polymerase chain reaction (PCR), linear polymerase reaction, nucleic acid sequence-based amplification, rolling circle amplification, and the like. An amplification reaction can generate an amplicon.

[0108] As mentioned herein, "adapter" generally refers to a short nucleic acid molecule (e.g., a length of about 10 to about 100 base pairs). An adaptor can include a short double-stranded DNA molecule. An adaptor can be attached to the end of a DNA fragment or an amplicon, for example, via polymerization or connection. An adaptor can include a synthetic oligonucleotide, for example, an oligonucleotide having a nucleotide sequence that is at least partially complementary to each other. An adaptor can have a blunt end, can have a staggered end (also referred to herein as a 3' or 5' "overhang sequence" or "sticky end"), or a blunt end and a staggered end. An adaptor can be attached to a fragment (e.g., via connection) to provide a fragment to which the adaptor is connected; the fragment to which the adaptor is connected can serve as a starting point for subsequent manipulation, for example, for amplification or sequencing. An adaptor can be functionalized, for example, conjugated to a label, a probe, a detectable label, an affinity capture reagent (e.g., biotin or streptavidin).

[0109] As used herein, the term "capturing moiety" generally refers to a molecule configured to be coupled to another part or molecule. The capturing moiety can be a biomolecule, such as a lipid, carbohydrate, sugar, amino acid, peptide or protein, nucleotide, nucleic acid molecule, metabolite or a combination thereof (e.g., glycoprotein, lipoprotein, glycosaminoglycan, etc.). The capturing moiety can be a small molecule, an organic compound, an inorganic compound, a metal, a polymer, an ion or other molecule or molecular compound. The capturing moiety can include a macromolecule. The capturing moiety can include an enzyme, an antibody, an antibody fragment, a nanobody, an aptamer, biotin, streptavidin, avidin, neutravidin (neutravidin) or its analogs or derivatives. The capturing moiety can include more than one molecule, such as a dimer, trimer, tetramer, pentamer, hexamer, heptamer, octamer, etc. The capturing moiety can be a solid substrate or a part of a solid substrate, or the capturing moiety can be separated from the substrate, such as in a fluid medium (e.g., air, in a liquid solution). The capturing moiety can be specific for a binding partner or multiple binding partners. The capture moiety may be capable of binding to one molecule or moiety (monovalent) or multiple molecules or moieties (multivalent).

[0110] As used herein, the abbreviations for the natural 1-enantiomer amino acids are conventional and may be as follows: alanine (A, Ala); arginine (R, Arg); asparagine (N, Asn); aspartic acid (D, Asp); cysteine ​​(C, Cys); glutamic acid (E, Glu); glutamine (Q, Gln); glycine (G, Gly); histidine (H, His); isoleucine (I, Ile); leucine (L, Leu); lysine (K, Lys); methionine (M, Met); phenylalanine (F, Phe); proline (P, Pro); serine (S, Ser); threonine (T, Thr); tryptophan (W, Trp); tyrosine (Y, Tyr); valine (V, Val). Unless otherwise indicated, X may represent any amino acid. In some aspects, X can be asparagine (N), glutamine (Q), histidine (H), lysine (K), or arginine (R). These amino acids are also referred to as "[amino acid] [residue / residue]" (e.g., lysine residue, lysine residues, leucine residue, leucine residues, etc.).

[0111] "Amino" refers to a -NH2 group.

[0112] "Cyano" refers to a -CN group.

[0113] "Nitro" refers to a -NO2 group.

[0114] "Oxo" refers to a =0 group.

[0115] "Hydroxy" refers to an -OH group.

[0116] "Alkyl" generally refers to an acyclic (e.g., straight or branched chain) or cyclic hydrocarbon (e.g., chain) group consisting solely of carbon and hydrogen atoms, such as a 1 to 15 carbon atom (e.g., C1-C 15 Unless otherwise indicated, an alkyl group is saturated or unsaturated (e.g., an alkenyl group, which contains at least one carbon-carbon double bond). Unless otherwise indicated, the disclosure provided herein of "alkyl" is intended to include the independent representation of a saturated "alkyl." The alkyl groups described herein are typically monovalent, but can also be divalent (which may also be described herein as "alkylene" or "alkylenyl"). In certain embodiments, an alkyl group contains 1 to 13 carbon atoms (e.g., C1-C 13In certain embodiments, the alkyl group comprises 1 to 8 carbon atoms (e.g., C1-C8 alkyl). In other embodiments, the alkyl group comprises 1 to 5 carbon atoms (e.g., C1-C5 alkyl). In other embodiments, the alkyl group comprises 1 to 4 carbon atoms (e.g., C1-C4 alkyl). In other embodiments, the alkyl group comprises 1 to 3 carbon atoms (e.g., C1-C3 alkyl). In other embodiments, the alkyl group comprises 1 to 2 carbon atoms (e.g., C1-C2 alkyl). In other embodiments, the alkyl group comprises 1 carbon atom (e.g., C1 alkyl). In other embodiments, the alkyl group comprises 5 to 15 carbon atoms (e.g., C5-C 15 In other embodiments, the alkyl group comprises 5 to 8 carbon atoms (e.g., C5-C8 alkyl). In other embodiments, the alkyl group comprises 2 to 5 carbon atoms (e.g., C2-C5 alkyl). In other embodiments, the alkyl group comprises 3 to 5 carbon atoms (e.g., C3-C5 alkyl). In other embodiments, the alkyl group is selected from methyl, ethyl, 1-propyl (n-propyl), 1-methylethyl (isopropyl), 1-butyl (n-butyl), 1-methylpropyl (sec-butyl), 2-methylpropyl (isobutyl), 1,1-dimethylethyl (tert-butyl), 1-pentyl (n-pentyl). The alkyl group is attached to the rest of the molecule by a single bond. Typically, the alkyl group is independently substituted or unsubstituted. Unless otherwise noted, each representation of "alkyl" provided herein includes a specific and clear representation of unsaturated "alkyl". Similarly, unless otherwise indicated, alkyl groups are optionally substituted with one or more of the following substituents: halo, cyano, nitro, oxo, thio, imino, oxime, trimethylsilyl, -OR a 、-SR a 、-OC(O)-R a 、-N(R a )2、-C(O)R a 、-C(O)OR a 、-C(O)N(R a )2、-N(R a )C(O)OR a 、-OC(O)-N(R a )2、-N(R a )C(O)R a 、-N(R a )S(O) t R a (where t is 1 or 2), -S(O) t OR a (where t is 1 or 2), -S(O) t R a (where t is 1 or 2) and -S(O) t N(Ra )2 (where t is 1 or 2), where each R a and R is independently hydrogen, alkyl (optionally substituted with halogen, hydroxy, methoxy or trifluoromethyl), fluoroalkyl, carbocyclyl (optionally substituted with halogen, hydroxy, methoxy or trifluoromethyl), carbocyclylalkyl (optionally substituted with halogen, hydroxy, methoxy or trifluoromethyl), aryl (optionally substituted with halogen, hydroxy, methoxy or trifluoromethyl), aralkyl (optionally substituted with halogen, hydroxy, methoxy or trifluoromethyl), heterocyclyl (optionally substituted with halogen, hydroxy, methoxy or trifluoromethyl), heterocyclylalkyl (optionally substituted with halogen, hydroxy, methoxy or trifluoromethyl), heteroaryl (optionally substituted with halogen, hydroxy, methoxy or trifluoromethyl) or heteroarylalkyl (optionally substituted with halogen, hydroxy, methoxy or trifluoromethyl).

[0117] "Aryl" refers to a radical derived from an aromatic monocyclic or polycyclic hydrocarbon ring system by removing hydrogen atoms from ring carbon atoms. The aromatic monocyclic or polycyclic hydrocarbon ring system contains only hydrogen and carbons from 5 to 18 carbon atoms, wherein at least one ring in the ring system is fully unsaturated, i.e., it contains a cyclic, delocalized (4n+2) π-electron system according to Hückel's theory. Ring systems from which aryl groups are derived include, but are not limited to, radicals such as benzene, fluorene, indane, indene, tetrahydronaphthalene, and naphthalene. Unless stated otherwise specifically in the specification, the term "aryl" or the prefix "ar-" (such as in "aralkyl") is meant to include aryl groups optionally substituted by one or more substituents independently selected from the group consisting of alkyl, alkenyl, alkynyl, halo, fluoroalkyl, cyano, nitro, optionally substituted aryl, optionally substituted aralkyl, optionally substituted aralkenyl, optionally substituted aralkynyl, optionally substituted carbocyclyl, optionally substituted carbocyclylalkyl, optionally substituted heterocyclyl, optionally substituted heterocyclylalkyl, optionally substituted heteroaryl, optionally substituted heteroarylalkyl, -R b -OR a 、-R b -OC(O)-R a 、-R b -OC(O)-OR a 、-R b -OC(O)-N(R a )2. -R b -N(R a )2. -R b -C(O)R a 、-R b -C(O)OR a 、-R b -C(O)N(R a )2. -R b -OR c -C(O)N(Ra )2. -R b -N(R a )C(O)OR a 、-R b -N(R a )C(O)R a 、-R b -N(R a )S(O) t R a (where t is 1 or 2), -R b -S(O) t R a (where t is 1 or 2), -R b -S(O) t OR a (where t is 1 or 2) and -R b -S(O) t N(R a )2 (where t is 1 or 2), where each R a R is independently hydrogen, alkyl (optionally substituted with halogen, hydroxy, methoxy or trifluoromethyl), fluoroalkyl, cycloalkyl (optionally substituted with halogen, hydroxy, methoxy or trifluoromethyl), cycloalkylalkyl (optionally substituted with halogen, hydroxy, methoxy or trifluoromethyl), aryl (optionally substituted with halogen, hydroxy, methoxy or trifluoromethyl), aralkyl (optionally substituted with halogen, hydroxy, methoxy or trifluoromethyl), heterocyclyl (optionally substituted with halogen, hydroxy, methoxy or trifluoromethyl), heterocyclylalkyl (optionally substituted with halogen, hydroxy, methoxy or trifluoromethyl), heteroaryl (optionally substituted with halogen, hydroxy, methoxy or trifluoromethyl) or heteroarylalkyl (optionally substituted with halogen, hydroxy, methoxy or trifluoromethyl), each R b are independently a direct bond or a straight or branched alkylene or alkenylene chain, and R c is a straight or branched alkylene or alkenylene chain, and wherein each of the above substituents is unsubstituted unless otherwise indicated.

[0118] "Carbocyclyl" or "cycloalkyl" refers to a stable non-aromatic monocyclic or polycyclic hydrocarbon radical consisting only of carbon and hydrogen atoms, including fused or bridged ring systems having 3 to 15 carbon atoms. In certain embodiments, the carbocyclyl comprises 3 to 10 carbon atoms. In other embodiments, the carbocyclyl comprises 5 to 7 carbon atoms. The carbocyclyl is attached to the remainder of the molecule by a single bond. The carbocyclyl or cycloalkyl is saturated (i.e., containing only a single C-C bond) or unsaturated (i.e., containing one or more double or triple bonds). The example of a saturated cycloalkyl includes, for example, cyclopropyl, cyclobutyl, cyclopentyl, cyclohexyl, cycloheptyl and cyclooctyl. Unsaturated carbocyclyl is also referred to as "cycloalkenyl". The example of a monocyclic cycloalkenyl includes, for example, cyclopentenyl, cyclohexenyl, cycloheptenyl and cyclooctenyl. Polycyclic carbocyclyls include, for example, adamantyl, norbornyl (i.e., bicyclo[2.2.1]heptyl), norbornenyl, decalinyl, 7,7-dimethyl-bicyclo[2.2.1]heptyl, etc. Unless otherwise specifically stated in the specification, the term "carbocyclyl" is intended to include carbocyclyls that are optionally substituted with one or more substituents independently selected from the group consisting of alkyl, alkenyl, alkynyl, halo, fluoroalkyl, oxo, thio, cyano, nitro, optionally substituted aryl, optionally substituted aralkyl, optionally substituted aralkenyl, optionally substituted aralkynyl, optionally substituted carbocyclyl, optionally substituted carbocyclylalkyl, optionally substituted heterocyclyl, optionally substituted heterocyclylalkyl, optionally substituted heteroaryl, optionally substituted heteroarylalkyl, -R b -OR a 、-R b -OC(O)-R a 、-R b -OC(O)-OR a 、-R b -OC(O)-N(R a )2. -R b -N(R a )2. -R b -C(O)R a 、-R b -C(O)OR a 、-R b -C(O)N(R a )2. -R b -OR c -C(O)N(R a )2. -R b -N(R a )C(O)OR a 、-R b -N(R a )C(O)R a 、-R b -N(R a )S(O)t R a (where t is 1 or 2), -R b -S(O) t R a (where t is 1 or 2), -R b -S(O) t OR a (where t is 1 or 2) and -R b -S(O) t N(R a )2 (where t is 1 or 2), where each R a R is independently hydrogen, alkyl (optionally substituted with halogen, hydroxy, methoxy or trifluoromethyl), fluoroalkyl, cycloalkyl (optionally substituted with halogen, hydroxy, methoxy or trifluoromethyl), cycloalkylalkyl (optionally substituted with halogen, hydroxy, methoxy or trifluoromethyl), aryl (optionally substituted with halogen, hydroxy, methoxy or trifluoromethyl), aralkyl (optionally substituted with halogen, hydroxy, methoxy or trifluoromethyl), heterocyclyl (optionally substituted with halogen, hydroxy, methoxy or trifluoromethyl), heterocyclylalkyl (optionally substituted with halogen, hydroxy, methoxy or trifluoromethyl), heteroaryl (optionally substituted with halogen, hydroxy, methoxy or trifluoromethyl) or heteroarylalkyl (optionally substituted with halogen, hydroxy, methoxy or trifluoromethyl), each R b are independently a direct bond or a straight or branched alkylene or alkenylene chain, and R c is a straight or branched alkylene or alkenylene chain, and wherein each of the above substituents is unsubstituted unless otherwise indicated.

[0119] "Polycycloalkyl" refers to a fused or bridged ring system having 3 to 15 carbon atoms and optionally 1-6 heteroatoms.

[0120] "Carbocyclylalkyl" refers to a group of the formula -R c -carbocyclic group, wherein R c is an alkylene chain as defined above. The alkylene chain and the carbocyclyl group are optionally substituted as defined above.

[0121] "Halo" or "halogen" refers to a fluoro, bromo, chloro, or iodo substituent.

[0122] "Haloalkyl" refers to an alkyl group as defined above substituted with one or more halo groups as defined above, for example, trihalomethyl, dihalomethyl, halomethyl, etc. In some embodiments, the haloalkyl group is a fluoroalkyl group, for example, trifluoromethyl, difluoromethyl, fluoromethyl, 2,2,2-trifluoroethyl, 1-fluoromethyl-2-fluoroethyl, etc. In some embodiments, the alkyl portion of the fluoroalkyl group is optionally substituted as defined above for an alkyl group.

[0123] The term "heteroalkyl" refers to an alkyl group as defined above in which one or more of the alkyl's backbone carbon atoms are replaced by heteroatoms (with an appropriate number of substituents or valences, e.g., -CH2- can be replaced by -NH- or -O-). For example, each substituted carbon atom is independently substituted by a heteroatom, such as where the carbon is replaced by nitrogen, oxygen, sulfur, or other suitable heteroatoms. In some cases, each substituted carbon atom is independently substituted by oxygen, nitrogen (e.g., -NH-, -N(alkyl)- or -N(aryl)- or with another substituent contemplated herein) or sulfur (e.g., -S-, -S(=O)- or -S(=O)2-). In some embodiments, heteroalkyl is attached to the rest of the molecule at a carbon atom of the heteroalkyl. In some embodiments, heteroalkyl is attached to the rest of the molecule at a heteroatom of the heteroalkyl. In some embodiments, heteroalkyl is C1-C 18 In some embodiments, the heteroalkyl group is C1-C 12 In some embodiments, heteroalkyl is C1-C6 heteroalkyl. In some embodiments, heteroalkyl is C1-C4 heteroalkyl. In some embodiments, heteroalkyl includes alkylamino, alkylaminoalkyl, aminoalkyl, heterocycloalkyl, heterocycloalkyl, heterocyclyl, and heterocycloalkylalkyl as defined herein. Unless otherwise specifically stated in the specification, heteroalkyl does not include alkoxy as defined herein. Unless otherwise specifically stated in the specification, heteroalkyl is optionally substituted as defined above for alkyl.

[0124] "Heteroalkylene" refers to a divalent heteroalkyl group as defined above that links one part of the molecule to another part of the molecule. Unless otherwise specifically stated, heteroalkylene is optionally substituted as defined above for an alkyl group.

[0125] "Heterocyclic radical" refers to a stable 3 to 18-membered non-aromatic ring group containing 2 to 12 carbon atoms and 1 to 6 heteroatoms selected from nitrogen, oxygen and sulfur. Unless otherwise specifically stated in the specification, heterocyclic radical is a monocyclic, bicyclic, tricyclic or tetracyclic ring system, which optionally includes a fused or bridged ring system. The heteroatoms in the heterocyclic radical are optionally oxidized. One or more nitrogen atoms (if present) are optionally quaternized. The heterocyclic radical is partially or completely saturated. The heterocyclic radical is saturated (i.e., only containing a single C-C bond) or unsaturated (e.g., containing one or more double bonds or triple bonds in the ring system). In some cases, the heterocyclic radical is saturated. In some cases, the heterocyclic radical is saturated and substituted. In some cases, the heterocyclic radical is unsaturated. Examples of such heterocyclic groups include, but are not limited to, dioxolanyl, thienyl[1,3]dithianyl, decahydroisoquinolinyl, imidazolinyl, imidazolidinyl, isothiazolidinyl, isoxazolidinyl, morpholinyl, octahydroindolyl, octahydroisoindolyl, 2-oxopiperazinyl, 2-oxopiperidinyl, 2-oxopyrrolidinyl, oxazolidinyl, piperidinyl, piperazinyl, 4-piperidonyl, pyrrolidinyl, pyrazolidinyl, quinuclidinyl, thiazolidinyl, tetrahydrofuranyl, trithianyl, tetrahydropyranyl, thiomorpholinyl, thiamorpholinyl, 1-oxo-thiomorpholinyl, and 1,1-dioxo-thiomorpholinyl. Unless otherwise specifically stated in the specification, the term "heterocyclyl" is intended to include heterocyclyl groups as defined above that are optionally substituted by one or more substituents selected from the group consisting of alkyl, alkenyl, alkynyl, halo, fluoroalkyl, oxo, thio, cyano, nitro, optionally substituted aryl, optionally substituted aralkyl, optionally substituted aralkenyl, optionally substituted aralkynyl, optionally substituted carbocyclyl, optionally substituted carbocyclylalkyl, optionally substituted heterocyclyl, optionally substituted heterocyclylalkyl, optionally substituted heteroaryl, optionally substituted heteroarylalkyl, -R b -OR a 、-R b -OC(O)-R a 、-R b -OC(O)-OR a 、-R b -OC(O)-N(R a )2. -R b -N(R a )2. -R b -C(O)R a 、-R b -C(O)OR a 、-R b -C(O)N(R a )2. -R b -OR c -C(O)N(R a )2. -R b -N(Ra )C(O)OR a 、-R b -N(R a )C(O)R a 、-R b -N(R a )S(O) t R a (where t is 1 or 2), -R b -S(O) t R a (where t is 1 or 2), -R b -S(O) t OR a (where t is 1 or 2) and -R b -S(O) t N(R a )2 (where t is 1 or 2), where each R a R is independently hydrogen, alkyl (optionally substituted with halogen, hydroxy, methoxy or trifluoromethyl), fluoroalkyl, cycloalkyl (optionally substituted with halogen, hydroxy, methoxy or trifluoromethyl), cycloalkylalkyl (optionally substituted with halogen, hydroxy, methoxy or trifluoromethyl), aryl (optionally substituted with halogen, hydroxy, methoxy or trifluoromethyl), aralkyl (optionally substituted with halogen, hydroxy, methoxy or trifluoromethyl), heterocyclyl (optionally substituted with halogen, hydroxy, methoxy or trifluoromethyl), heterocyclylalkyl (optionally substituted with halogen, hydroxy, methoxy or trifluoromethyl), heteroaryl (optionally substituted with halogen, hydroxy, methoxy or trifluoromethyl) or heteroarylalkyl (optionally substituted with halogen, hydroxy, methoxy or trifluoromethyl), each R b are independently a direct bond or a straight or branched alkylene or alkenylene chain, and R c is a straight or branched alkylene or alkenylene chain, and wherein each of the above substituents is unsubstituted unless otherwise indicated.

[0126] "N-heterocyclyl" or "N-attached heterocyclyl" refers to a heterocyclyl as defined above that contains at least one nitrogen atom, and wherein the point of attachment of the heterocyclyl to the rest of the molecule is through the nitrogen atom in the heterocyclyl. N-heterocyclyl groups are optionally substituted as described above for heterocyclyl groups. Examples of such N-heterocyclyl groups include, but are not limited to, 1-morpholinyl, 1-piperidinyl, 1-piperazinyl, 1-pyrrolidinyl, pyrazolidinyl, imidazolinyl, and imidazolidinyl.

[0127] "C-heterocyclyl" or "C-attached heterocyclyl" refers to a heterocyclyl as defined above that contains at least one heteroatom, and wherein the point of attachment of the heterocyclyl to the rest of the molecule is through a carbon atom in the heterocyclyl. C-heterocyclyl groups are optionally substituted as described above for heterocyclyl groups. Examples of such C-heterocyclyl groups include, but are not limited to, 2-morpholinyl, 2-, 3-, or 4-piperidinyl, 2-piperazinyl, 2- or 3-pyrrolidinyl, and the like.

[0128] "Heterocyclylalkyl" refers to a group of the formula -R c -heterocyclic group, wherein R c is an alkylene chain as defined above. If the heterocyclyl group is a nitrogen-containing heterocyclyl group, the heterocyclyl group is optionally attached to the alkyl group at the nitrogen atom. The alkylene chain of the heterocyclylalkyl group is optionally substituted as defined above for an alkylene chain. The heterocyclyl portion of the heterocyclylalkyl group is optionally substituted as defined above for a heterocyclyl group.

[0129] "Heterocyclylalkoxy" refers to a group of the formula -OR c - a group bonded to the oxygen atom of a heterocyclic group, wherein R c is an alkylene chain as defined above. If the heterocyclyl group is a nitrogen-containing heterocyclyl group, the heterocyclyl group is optionally attached to the alkyl group at the nitrogen atom. The alkylene chain of the heterocyclylalkoxy group is optionally substituted as defined above for an alkylene chain. The heterocyclyl portion of the heterocyclylalkoxy group is optionally substituted as defined above for a heterocyclyl group.

[0130] "Heteroaryl" refers to a group derived from a 3 to 18-membered aromatic ring group containing 2 to 17 carbon atoms and 1 to 6 heteroatoms selected from nitrogen, oxygen and sulfur. As used herein, heteroaryl is a monocyclic, bicyclic, tricyclic or tetracyclic ring system in which at least one ring is fully unsaturated, i.e., it contains a cyclic, delocalized (4n+2)π-electron system according to Huckel theory. Heteroaryl includes fused or bridged ring systems. The heteroatoms in the heteroaryl are optionally oxidized. One or more nitrogen atoms, if present, are optionally quaternized. The heteroaryl is attached to the rest of the molecule via any atom of the ring. Examples of heteroaryl include, but are not limited to, aza-. benzo[b][1,4]dioxolyl, benzofuranyl, benzoxazolyl, benzo[d]thiazolyl, benzothiadiazolyl, benzo[b][1,4]dioxolyl, benzo[b][1,4]oxazinyl, 1,4-benzodioxanyl, benzonaphthofuranyl, benzoxazolyl, benzodioxolyl, benzodioxinyl, benzopyranyl, benzopyranonyl, benzofuranyl, benzofuranonyl, benzothiophenyl (benzophenylthio), benzothieno[3,2-d]pyrimidinyl, benzotriazolyl, benzo[4,6]imidazo[1,2-a]pyridinyl, carbazolyl, cinnolinyl, cyclopenta[d]pyrimidinyl, 6,7-dihydro-5H-cyclopenta[4,5]thieno[2,3-d]pyrimidinyl, 5,6-dihydrobenzo[h]quinazolinyl, 5,6-dihydrobenzo[h]cinnolinyl, 6,7-dihydro-5H -benzo[6,7]cyclohepta[1,2-c]pyridazinyl, dibenzofuranyl, dibenzothiophenyl, furanyl, furanonyl, furano[3,2-c]pyridinyl, 5,6,7,8,9,10-hexahydrocyclooctano[d]pyrimidinyl, 5,6,7,8,9,10-hexahydrocyclooctano[d]pyridazinyl, 5,6,7,8,9,10-hexahydrocyclooctano[d]pyridinyl, isothiazolyl, imidazolyl, indazolyl, indolyl, indazolyl, isoindolyl, indolyl, isoindolyl, isoquinolinyl, indolizinyl, isoxazolyl, 5,8-methano-5,6,7,8-tetrahydroquinazolinyl, naphthyridinyl, 1,6-naphthyridonyl, oxadiazolyl, 2-oxoazepine yl, oxazolyl, oxiranyl, 5,6,6a,7,8,9,10,10a-octahydrobenzo[h]quinazolinyl, 1-phenyl-1H-pyrrolyl, phenazinyl, phenothiazinyl, phenoxazinyl, phthalazinyl, pteridinyl, purinyl, pyrrolyl, pyrazolyl, pyrazolo[3,4-d]pyrimidinyl, pyrido[3,2-d]pyrimidinyl, pyrido[3,4-d]pyrimidinyl, pyrazinyl, pyrimidinyl, pyridazinyl, pyrrolyl, quinazolinyl, quinoxalinyl, quinolinyl, isoquinolinyl, tetrahydroquinolinyl, 5-[[(4-d]-1-methyl-2-oxazol-1-yl)-1H-pyrrolyl] ... ,6,7,8-tetrahydroquinazolinyl, 5,6,7,8-tetrahydrobenzo[4,5]thieno[2,3-d]pyrimidinyl, 6,7,8,9-tetrahydro-5H-cyclohepta[4,5]thieno[2,3-d]pyrimidinyl, 5,6,7,8-tetrahydropyrido[4,5-c]pyridazinyl, thiazolyl, thiadiazolyl, triazolyl, tetrazolyl, triazinyl, thieno[2,3-d]pyrimidinyl, thieno[3,2-d]pyrimidinyl, thieno[2,3-c]pyridinyl and phenylthio (i.e., thienyl). Unless otherwise specifically stated in the specification, the term "heteroaryl" is intended to include heteroaryl groups as defined above that are optionally substituted by one or more substituents selected from the group consisting of alkyl, alkenyl, alkynyl, halo, fluoroalkyl, haloalkenyl, haloalkynyl, oxo, thio, cyano, nitro, optionally substituted aryl, optionally substituted aralkyl, optionally substituted aralkenyl, optionally substituted aralkynyl, optionally substituted carbocyclyl, optionally substituted carbocyclylalkyl, optionally substituted heterocyclyl, optionally substituted heterocyclylalkyl, optionally substituted heteroaryl, optionally substituted heteroarylalkyl, -R b -OR a 、-R b -OC(O)-R a 、-R b -OC(O)-OR a 、-R b -OC(O)-N(R a )2. -R b -N(R a )2. -R b -C(O)R a 、-R b -C(O)OR a 、-R b -C(O)N(R a )2. -R b -OR c -C(O)N(R a )2. -R b -N(R a )C(O)OR a 、-R b -N(R a )C(O)R a、-R b -N(R a )S(O) t R a (where t is 1 or 2), -R b -S(O) t R a (where t is 1 or 2), -R b -S(O) t OR a (where t is 1 or 2) and -R b -S(O) t N(R a )2 (where t is 1 or 2), where each R a R is independently hydrogen, alkyl (optionally substituted with halogen, hydroxy, methoxy or trifluoromethyl), fluoroalkyl, cycloalkyl (optionally substituted with halogen, hydroxy, methoxy or trifluoromethyl), cycloalkylalkyl (optionally substituted with halogen, hydroxy, methoxy or trifluoromethyl), aryl (optionally substituted with halogen, hydroxy, methoxy or trifluoromethyl), aralkyl (optionally substituted with halogen, hydroxy, methoxy or trifluoromethyl), heterocyclyl (optionally substituted with halogen, hydroxy, methoxy or trifluoromethyl), heterocyclylalkyl (optionally substituted with halogen, hydroxy, methoxy or trifluoromethyl), heteroaryl (optionally substituted with halogen, hydroxy, methoxy or trifluoromethyl) or heteroarylalkyl (optionally substituted with halogen, hydroxy, methoxy or trifluoromethyl), each R b are independently a direct bond or a straight or branched alkylene or alkenylene chain, and R c is a straight or branched alkylene or alkenylene chain, and wherein each of the above substituents is unsubstituted unless otherwise indicated.

[0131] "N-heteroaryl" refers to a heteroaryl group as defined above that contains at least one nitrogen and wherein the point of attachment of the heteroaryl group to the rest of the molecule is through the nitrogen atom in the heteroaryl group. N-heteroaryl groups are optionally substituted as described above for heteroaryl groups.

[0132] "C-heteroaryl" refers to a heteroaryl group as defined above, and wherein the point of attachment of the heteroaryl group to the rest of the molecule is through a carbon atom in the heteroaryl group. A C-heteroaryl group is optionally substituted as described above for a heteroaryl group.

[0133] "Heteroarylalkyl" refers to a group of the formula -R c -heteroaryl groups, wherein R cis an alkylene chain as defined above. If the heteroaryl group is a nitrogen-containing heteroaryl group, the heteroaryl group is optionally attached to the alkyl group at the nitrogen atom. The alkylene chain of the heteroarylalkyl group is optionally substituted as defined above for an alkylene chain. The heteroaryl portion of the heteroarylalkyl group is optionally substituted as defined above for a heteroaryl group.

[0134] "Heteroarylalkoxy" refers to a group of the formula -OR c - a group bonded to the oxygen atom of a heteroaryl group, wherein R c is an alkylene chain as defined above. If the heteroaryl group is a nitrogen-containing heteroaryl group, the heteroaryl group is optionally attached to the alkyl group at the nitrogen atom. The alkylene chain of the heteroarylalkoxy group is optionally substituted as defined above for an alkylene chain. The heteroaryl portion of the heteroarylalkoxy group is optionally substituted as defined above for a heteroaryl group.

[0135] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure belongs. Suitable methods and materials are described below, although methods and materials similar or equivalent to those described herein can be used in the practice or testing of the present disclosure.

[0136] Single-molecule sequencing of polymeric analytes using thiocarbamoyl and dithioester reactive groups

[0137] The present disclosure provides a novel method for sequencing polymer analytes such as peptides, wherein a sequencing reagent of Formula I is used to separate monomers (e.g., terminal amino acids from peptides) in the following manner: contacting the sequencing reagent with the monomers of the polymer analyte, tethering the monomers (e.g., terminal amino acids) to a capture portion using a sequencing reagent, cutting the monomers from the polymer analyte (e.g., peptide), and detecting the monomers or their derivatives. In some cases, detection includes using a specific binding agent of a monomer, a sequencing reagent-monomer complex (e.g., a sequencing reagent-amino acid complex), or a derivative thereof. Separating monomers from the polymer analyte in the methods provided herein can avoid local environmental problems, where adjacent monomers may affect the binding (e.g., selectivity or affinity) of the binding agent. In some cases, detection is performed using a direct readout method, e.g., using a nanopore, which can identify monomers, sequencing reagent-monomer complexes, or their derivatives.

[0138] In some cases, the polymer analytes described herein include peptides, and the methods provided herein may include performing an Edman degradation reaction or an Edman-like degradation reaction to cleave the N-terminal amino acid from the peptide. In typical peptide sequencing using Edman degradation, a phenyl isothiocyanate reactive group (PITC) is used to react with the N-terminal amino acid to produce a phenylthiocarbamoyl intermediate, which is cleaved to produce a cyclic compound. The sequential removal of amino acid residues allows peptide sequencing to be performed without damaging the peptide or protein itself. However, there are many problems with Edman degradation and the use of PITC as a reactive group. These problems include the need for harsh reaction conditions (e.g., high heat and acid). These harsh reaction conditions are not suitable for nucleic acid molecules (e.g., DNA used for barcoding of monomers or coupling monomers to substrates, as described elsewhere herein and Figure 1A 、 1B , 1D, 1E, and 1F). Edman degradation can also lead to racemization of the cleaved N-terminal amino acid from the peptide, which can limit recognition and detection by the binder (for example, if the binder recognizes only one stereoisomer). In addition, in amino acids with amine side chain functionality (such as lysine), the PITC-reactive group can react with the side chain to generate a PITC-conjugated amino acid, which may require additional processing for detection or require that the detection method be able to detect the PITC-conjugated amino acid.

[0139] In some embodiments, a sequencing reagent is provided herein that comprises a reactive group containing a dithioester. In some embodiments, the dithioester includes thiobenzoyl, thioacetyl, dithiocarboxylate (carbodithioic ester), cyanomethyl dithiobenzoate or N-thiobenzoyl succinimide. Compared with PITC reactive groups and Edman degradation, the alternative reactive groups comprising dithioesters provided herein provide many advantages for peptide sequencing. The dithioester reactive groups provided herein (e.g., thioacetyl or thiobenzoyl) may require milder reaction conditions than the high heat and acidic conditions required for Edman degradation, thereby making the sequencing operations described elsewhere herein (e.g., using nucleic acid molecule linkers) more feasible. In addition, in some embodiments, the dithioester reactive groups provided herein (e.g., thioacetyl or thiobenzoyl) can result in non-chiral cleavage products, which can eliminate the non-binding or limited binding of the binding agent that recognizes a specific stereoisomer. Furthermore, the dithioester reactive groups provided herein can avoid byproduct formation of lysine residues associated with Edman degradation, thereby providing additional functionalizable (amine) groups and / or more directly detecting monomers containing primary amine groups.

[0140] In some embodiments, the reactive group may include a dithioester. The dithioester reactive group may include a thiobenzoyl group. The thiobenzoyl reactive group may avoid side reactions associated with Edman degradation and may lead to more accurate and clear sequence determination, as described in Barrett et al. 1975. Febs Letters 57, 1, 19-21. The thiobenzoyl group may react with amino acids and produce achiral side chain residues after cleavage, which may improve detectability. For example, the binding agents described herein may only recognize specific stereoisomers of achiral side chains or amino acid side chains; therefore, preventing racemization of the side chains of the cleaved amino acids may allow the binding agent to recognize and specifically bind to the side chain (or a complex comprising the side chain). In some embodiments, the reactive group may include a thioacetyl group. In addition to the ability to couple under mild conditions as described in Stolowitz et al. 1989. Analytical Biochemistry 181, 113-119, the thioacetyl reactive group may also have the advantage of avoiding interference that may occur during the entire coupling reaction. In some embodiments, the reactive group can include dithiocarboxylate. Compared with the traditional Edman degradation (for example, high temperature) as described in Previero et al. 1974.Febs Letters51,1,68-72, the dithiocarboxylate reactive group is similar to other reactive groups proposed herein, and the required reaction conditions are not so extreme, more gentle. In some embodiments, the reactive group can include cyanomethyl dithiobenzoate. The cyanomethyl dithiobenzoate reactive group has advantages over traditional Edman degradation and PITC reactive groups, because the derivative cyclization reaction can occur under less extreme conditions (for example, high heat), and the hydrolysis of the non-terminal peptide bond with problems in amino acid sequencing can be avoided, as described in Previero et al. 1970.Biochemical and Biophysical Research Communications.40,3,549-556. In some embodiments, the reactive group can include N-thiobenzoylsuccinimide. The use of N-thiobenzoylsuccinimide reactive groups offers several advantages over PITC reactive groups, including a short cyclization step, a thiazolinone as the final product of the cyclization, and the thiazolinone can be analyzed after a derivatization step or released as a free amino acid after acid hydrolysis, as described in Cavadore et al. 1978. Analytical Biochemistry. 91, 1, 236-240. In some embodiments, the reactive group can include a thiocarbamoyl group.Thiocarbamoyl groups may have the advantage of milder reaction conditions relative to PITC reactive groups, or the use of thiocarbamoyl reactive groups may result in achiral residues.

[0141] This article provides a sequencing reagent of formula I:

[0142]

[0143] or a stereoisomer, tautomer or salt thereof.

[0144] In some embodiments, A comprises a first reactive group configured to form a covalent bond with a monomer of a polymer analyte (e.g., the N-terminal amino acid of a polypeptide), wherein the reactive group is selected from a dithioester or a thiocarbonyl group. In some embodiments, A comprises a reactive group configured to form a covalent bond with the N-terminal amino acid of a polypeptide, wherein the reactive group is a dithioester. In some embodiments, A comprises a reactive group configured to form a covalent bond with the N-terminal amino acid of a polypeptide, wherein the reactive group is a thiocarbonyl group. In some cases, the thiocarbonyl group is a thiocarbamoyl group. In some embodiments, B comprises a second reactive group. The second reactive group can be used to couple a sequencing reagent to another molecule, such as a capture moiety. In some embodiments, B comprises a click chemistry moiety. In some embodiments, L 1 Includes a linker coupled to A and B.

[0145] In some cases, "reactive group" as used herein may refer interchangeably to "first reactive group."

[0146] In some embodiments, the second reactive group is covalently attached to a polymer. For example, a sequencing reagent may include a first reactive group, a polymer, and a second reactive group. Alternatively or additionally, a sequencing reagent may include a first reactive group and a first click chemistry group (e.g., an azide), which may react with a linker molecule (such as a polymer) comprising a second click chemistry group (e.g., an alkyne, such as DBCO). The reaction of the first click chemistry group and the second click chemistry group may thus produce a sequencing reagent covalently attached to the linker molecule. The linker molecule may additionally include other reactive groups. In such examples, the reaction of the first click chemistry group and the second click chemistry group may produce a sequencing reagent comprising (i) a first reactive group, (ii) a linker molecule, and (iii) a second reactive group.

[0147] The connecting molecule can include a polymer. The polymer can be a synthetic polymer, such as polyalkylene glycol (PAG) (for example, polyethylene glycol (PEG) or polypropylene glycol (PPG)), or a naturally occurring polymer, which optionally can be synthetic, such as deoxyribonucleic acid (DNA) or ribonucleic acid (RNA). In some embodiments, the polymer includes (PEG). In some embodiments, the polymer includes deoxyribonucleic acid (DNA) or ribonucleic acid (RNA). In some embodiments, the polymer includes deoxyribonucleic acid (DNA). In some embodiments, the polymer includes ribonucleic acid (RNA). In some embodiments, the polymer is covalently attached to the surface.

[0148] In some embodiments, the second reactive group is covalently attached or is configured to be covalently attached to a surface-bound joint (for example, a surface-bound capture moiety). In some embodiments, the second reactive group is covalently attached to a surface-bound joint, and the surface-bound joint is connected to the surface. In some embodiments, the surface-bound joint comprises an ethyl or propyl group. In some embodiments, the surface-bound joint comprises an ethyl group. In some embodiments, the surface-bound joint comprises a propyl group. In some embodiments, the surface-bound joint comprises an ethyl group. In some embodiments, the surface-bound joint comprises a nucleic acid molecule. In some cases, the second reactive group comprises a click chemistry moiety (for example, an azide, an alkyne), and the surface or the surface-bound joint comprises another click chemistry moiety, which can react with the click chemistry moiety of the second reactive group. In one such example, the second reactive group can comprise an azide that can react with an alkyne (for example, DBCO) surface or a surface-bound joint.

[0149] In some embodiments, the second reactive group comprises:

[0150]

[0151] in Indicates the orientation of the second reactive group relative to the reactive group.

[0152] In some embodiments, the second reactive group comprises a click chemistry moiety, e.g., as described elsewhere herein. In some embodiments, the first reactive group comprises a thioacetyl group. In some embodiments, the reactive group comprises:

[0153]

[0154] in Indicates the orientation of the reactive group relative to the second reactive group. In some embodiments, the first reactive group comprises a thiobenzoyl group. In some embodiments, the reactive group comprises:

[0155]

[0156] in Indicates the orientation of the reactive group relative to the second reactive group.

[0157] In some embodiments, the reactive group comprises

[0158]

[0159] in Indicates the orientation of the reactive group relative to the substrate tethering moiety. In some embodiments, the reactive group comprises a thioacetyl group. In some embodiments, the reactive group comprises a thiobenzoyl group. In some embodiments, the reactive group comprises a derivative of N-thiobenzoylsuccinimide. In some embodiments, the reactive group comprises a derivative of cyanomethyldithiobenzoate. In some embodiments, the first reactive group comprises In some embodiments, the (e.g., first) reactive group includes In some embodiments, the (e.g., first) reactive group includes In some embodiments, the (e.g., first) reactive group includes In some embodiments, the (e.g., first) reactive group includes In some embodiments, the (e.g., first) reactive group includes In some embodiments, the (eg, first) reactive group is linked to the N-terminal amino acid of the polypeptide via a covalent bond.

[0160] In some embodiments, L1 includes a cleavable linker. Non-limiting examples of cleavable linkers include disulfide bonds, hydrazones, DNA, peptides, and click chemistry moieties. The cleavable linker can be cut using any suitable mechanism, such as by applying a stimulus. The stimulus can be, for example, a chemical stimulus, a biological stimulus, a thermal stimulus (e.g., applying heat), a light stimulus, a physical or mechanical stimulus, or a combination of other types of stimulus or stimulus. In some cases, the stimulus is a chemical stimulus, for example, a change in pH, a solubilizing agent, an initiator, a free radical generator, the addition of a reducing agent, etc. For example, a cleavable linker can be included in a disulfide bond that can be cut when a reducing agent (e.g., DTT or TCEP) is applied. Alternatively or additionally, the linker can include a conjugate receptor that can be clicked and clicked, for example, as described in Diehl et al. Nature Chemistry.2016.8,968-973, which is incorporated herein by reference in its entirety. In another example, the linker can comprise a reversible hydrazone bond, for example, as described in Nisal et al. Organic and Biomolecular Chemistry. 2018. Iss. 23, which is incorporated herein by reference in its entirety. The hydrazone bond can be o-aminobenzylhydrazone.

[0161] In some cases, the stimulus is a biological stimulus, for example, an enzyme that can cut a joint or catalyze the cutting of a joint (e.g., nuclease, uracil DNA glycosylase). In one such example, the joint can comprise a nucleic acid molecule (e.g., DNA) comprising a cleavage site (e.g., restriction site, uracil, abasic site) that can be cut using an enzyme (e.g., restriction enzyme, uracil DNA glycosylase, nuclease, etc.). Alternatively or additionally, the nucleic acid molecule can comprise a partially hybridized portion or toehold region that can be dehybridized by strand displacement (e.g., using a strand displacement enzyme such as Phi29 or a competitive nucleic acid molecule). In another example, the joint can comprise a peptide sequence that can be cut using a protease (e.g., exopeptidase, aminopeptidase, proteinase K, lysC, etc.).

[0162] In some embodiments, L1 comprises a hydrophilic polymer.

[0163] In some embodiments, L1 comprises a polyalkylene glycol linker. The linker can be, for example, a linear or branched polyalkylene glycol (PAG). For example, the linker can be a branched PEG linker. The branched PEG linker can have three arms or four arms. In some embodiments, the linker has the general structure PEG CCO-CCO-CCO or PPG CCCO-CCCO-CCCO. In some embodiments, the linker is a poly(propylene oxide) linker, such as poly(propylene glycol) (PPG), which has the structure H[OCH(CH3)CH2].n OH, where n is an integer equal to or greater than 1. The linker can include a combination of polyalkylene oxides, such as poly(ethylene glycol)-poly(propylene glycol)-poly(ethylene glycol) diacrylate. For example, the linker can have the following structure: wherein at least one of x, y, and z is an integer greater than 0.

[0164] In some embodiments, L1 comprises a non-cleavable linker. In some cases, L1 comprises an alkyl or hydrocarbon chain.

[0165] In some embodiments, the second reactive moiety is directly or indirectly attached to the substrate, as described elsewhere herein. In some embodiments, the first reactive group is covalently linked to the N-terminal amino acid of the polypeptide attached to the substrate; and wherein the second reactive moiety is directly or indirectly attached to the substrate.

[0166] In some embodiments, provided herein is a sequencing reagent of Formula II:

[0167]

[0168] In some embodiments, R1 is a leaving group, an aryl group, or an electronegative group. In some embodiments, S-R1 is a leaving group, an aryl group, or an electronegative group. In some embodiments, the leaving group includes an electrophilic group. The electrophilic group may include, for example, S, SH, SO3CF3, SO3H or NHTf. In some embodiments, the electrophilic group includes SR*, wherein R* includes H, R', OH, OR', NH2 or NHR', wherein R' is C1-C6 alkyl optionally substituted with one or more members selected from halo, C1-C3 alkyl, C1-C3 alkoxy, C1-C3 haloalkyl, phenyl, 5-membered heteroaryl and 6-membered heteroaryl, wherein phenyl, 5-membered heteroaryl and 6-membered heteroaryl are optionally substituted with one or two members selected from halo, OH, C1-C3 alkyl, C1-C3 alkoxy, C1-C3 haloalkyl, NO2, CN , COOR" and CON(R")2, wherein each R" is independently H or C1-3 alkyl. In some embodiments, the leaving group comprises a C1-C6 alkyl or aromatic group. In some embodiments, R2 is a linker (e.g., L1). In some embodiments, L1 is a cleavable linker. In some embodiments, R3 is a click chemistry moiety (e.g., azide, alkyne, DBCO, tetrazine, trans-cyclooctene (TCO)). In some embodiments, R3 is azide. In some embodiments, R3 is alkyne.

[0169] In some embodiments, provided herein is a sequencing reagent of Formula III:

[0170]

[0171] In some embodiments, R4 is a linker (e.g., L1). In some embodiments, R4 is a cleavable linker. In some embodiments, R5 is a linker (e.g., L1). In some embodiments, R5 is a cleavable linker. In some embodiments, R6 is a click chemistry moiety (e.g., azide, alkyne). In some embodiments, R6 is azide. In some embodiments, R6 is alkyne. In some embodiments, R7 is azide. In some embodiments, R7 is alkyne. In specific embodiments, R6 or R7 is a click chemistry moiety (e.g., azide, alkyne) and R4 or R5 is a linker (e.g., L1). In specific embodiments, at least one of R4-R6 or R5-R7 is a linker-click chemistry moiety. In some embodiments, R8 is a leaving group. In some embodiments, a leaving group includes an electrophilic group. An electrophilic group can include, for example, S, SO3CF3, SO3H or NHTf. In some embodiments, the electrophilic group comprises SR*, wherein R* comprises H, R', OH, OR', NH2 or NHR', wherein R' is C1-C6 alkyl optionally substituted with one or more members selected from halo, C1-C3 alkyl, C1-C3 alkoxy, C1-C3 haloalkyl, phenyl, 5-membered heteroaryl and 6-membered heteroaryl, wherein phenyl, 5-membered heteroaryl and 6-membered heteroaryl are optionally substituted with one or two members selected from halo, OH, C1-C3 alkyl, C1-C3 alkoxy, C1-C3 haloalkyl, NO2, CN, COOR" and CON(R")2, wherein each R" is independently H or C1-3 alkyl. In some embodiments, the leaving group comprises C1-C6 alkyl or an aromatic group. In some embodiments, R8 is halogen or halide. In some embodiments, R8 comprises diimidazole.

[0172] Analysis of polymeric analytes by local tethering

[0173] The present disclosure also provides a method for processing and analyzing polymer analytes (e.g., peptides, polymers, nucleic acid molecules, etc.) using sequencing reagents as described herein. The method for processing polymer analytes can achieve a highly parallelized and accurate mode for sequencing and identifying polymer analytes. The systems and methods of the present disclosure can include using sequencing reagents to couple the monomers of the polymer analyte to a capture portion, cutting the monomers from the polymer analyte to generate a detectable complex, and detecting the detectable complex. In some embodiments, the method includes providing a capture portion, a polymer analyte and a sequencing reagent, and contacting the polymer analyte with a sequencing reagent. The sequencing reagent can be combined with the monomer of the polymer analyte to form a sequencing reagent-monomer complex. The sequencing reagent-monomer complex can be coupled to the capture portion and then cut from the polymer analyte to provide a detectable complex. One or more operations can be iterated to sequentially process, analyze, or characterize a single monomer of a polymer analyte.

[0174] In some embodiments, detection includes using a monomer-specific binding agent to identify and bind to the cut monomer. In some embodiments, the monomer-specific binding agent is used for direct or indirect detection; for example, the binding agent can include a detectable label (e.g., a fluorophore, a mass label, a radioisotope) that can be directly detected, or the binding agent can include a polymerizable molecule with encoded information, which can be transferred by coupling or copying the encoded information to a capture portion or another polymerizable molecule. In some cases, another polymerizable molecule or capture portion is located near the polymer analyte. Alternatively or additionally, the binding agent can be used to sort the cut monomer, for example, sorting into separate partitions or compartments for downstream labeling, for example, with an identifying barcode molecule. In some embodiments, the detectable complex is detected without performing another labeling operation or contacting with a binding agent.

[0175] One or more operations can be iterated or repeated any number of times to obtain information about all monomers or monomer subsets of polymer analytes, and optionally, information about the order of monomers relative to polymer analytes. Information can be read from polymerizable molecules using, for example, conventional next generation sequencing or nanopore sequencing methods. Beneficially, by cutting monomers from polymer analytes, monomers can be removed from adjacent monomers, which can eliminate the local environmental effects caused by adjacent monomers. In some cases, cutting monomers from polymer analytes enables binding agents to specifically bind to single monomers without being affected by adjacent surrounding monomers. In some cases, cutting monomers from polymer analytes can result in intramolecular expansion, wherein the distance between monomers increases, thereby improving downstream readout and identification (for example, using nanopore sequencers). Therefore, methods disclosed herein can achieve more accurate molecular identification and polymer sequencing, which has applications in diagnosing diseases, monitoring protein dynamics or protein interactions, single cell proteomics, developing or characterizing therapeutic agents.

[0176] One or more methods of the present disclosure may employ sequencing reagents as described herein. The sequencing reagent may be capable of coupling to (i) a monomer of a polymer analyte and (ii) a capture portion, which may be used to locally tether the monomer to the vicinity of the polymer analyte once it is cut. The methods and systems disclosed herein may additionally include a substrate for local tethering; the substrate may, for example, be coupled to the polymer analyte, the capture portion, and another polymerizable molecule. In some cases, the other polymerizable molecule is encoded with information of the polymerizable molecule from the binding agent. In other cases, the binding agent includes a detectable label that can indicate a binding event between the binding agent and the monomer. Alternatively or additionally, one or more processes described herein may be performed in solution or in the absence of a substrate.

[0177] The method disclosed herein for processing a polymer analyte comprising a plurality of monomers may include cutting a monomer of the polymer analyte and coupling the monomer to a capture portion (e.g., coupled to a substrate or in solution) for subsequent processing or analysis. In one example, the method disclosed herein may include: providing (i) a polymer analyte comprising a plurality of monomers and (ii) a capture portion; coupling a monomer of the plurality of monomers to a capture portion to generate a monomer-capture portion complex; cutting the monomer; contacting the cut monomer-capture portion complex with a binding agent; and coupling a first polymerizable molecule to a second polymerizable molecule or a capture portion. In some cases, the coupling of the monomer to the capture portion is mediated using a sequencing reagent. For example, a sequencing reagent can be coupled to a monomer to form a sequencing reagent-monomer complex. The sequencing reagent-monomer complex can then be coupled to a capture portion. In some cases, the first polymerizable molecule is coupled to a binding agent and contains information about the binding agent, such as the identity of the binding agent or its homologous molecule, which can be transferred to a second polymerizable molecule or a capture portion. In some cases, the capture portion can be a copy or identical molecule of the second polymerizable molecule.

[0178] Alternatively or additionally, a binding agent can be used to sort the mixture of cleaved monomers by identity or type, and after sorting, an identification tag or barcode identifying the monomer type can be coupled to the capture moiety. Example methods and systems of such processing methods and systems are described in U.S. Patent No. 11,499,979, International Patent Application No. PCT / US2023 / 017954, and U.S. Provisional Patent Application No. 63 / 507,558, filed June 12, 2023, each of which is incorporated herein by reference in its entirety.

[0179] Figure 1A An example workflow for analyzing a polymeric analyte is shown. In workflow 100, a polymeric analyte 103 is provided and sequentially depolymerized into single monomers by contact with a sequencing reagent and coupling to a capture moiety, cleavage of the monomers, and indirect detection using a binding agent comprising a polymerizable molecule that identifies or encodes the binding agent. The polymerizable molecule of the binding agent is coupled to another polymerizable molecule, and the monomers are cleaved or blocked from the capture moiety to prevent downstream recognition from another binding agent, thereby allowing the process to be iterated to analyze all or a subset of the single monomers contained in the polymeric analyte. Figure 1AIn sub-figure A, substrate 101 is coupled to a polymer analyte 103 (e.g., a peptide to be sequenced), a capture portion 105 (e.g., a nucleic acid molecule), and an additional polymerizable molecule 107 (e.g., an additional nucleic acid molecule). In some cases, the capture portion 105 and the additional polymerizable molecule 107 are the same molecule (e.g., comprising the same sequence). The polymer analyte can be contacted with a sequencing reagent 109 (e.g., a sequencing reagent of Formula I-III) comprising (i) a first reactive group that is a terminal monomer-coupling group (e.g., an amino acid-reactive group comprising a dithioester or thiocarbamoyl group) and (ii) a second reactive group, such as a capture-binding portion comprising a click chemistry portion (e.g., an azide). In some cases, sequencing reagent 109 is coupled to a terminal monomer (e.g., a terminal amino acid, such as an N-terminal amino acid (NTAA)). The coupling results in the formation of a sequencing reagent-monomer complex. In Figure 1A In panel B, a linker nucleic acid molecule 111 comprising a click chemistry moiety (e.g., alkyne) reacts with and covalently links to a second reactive group of a sequencing reagent 109 within a sequencing reagent-monomer complex. In some cases, pre-coupled linker nucleic acid molecules 111 and sequencing reagents 109 are provided (see, e.g., Figure 2 ). In such cases, the linked nucleic acid molecule 111 may constitute the capture-binding portion of the sequencing reagent. Figure 1A In sub-figure C, the linking nucleic acid molecule 111 is coupled to the capture moiety 105, thereby generating a monomer-capture moiety complex comprising a sequencing reagent-monomer complex coupled to the capture moiety 105. The coupling can be mediated by hybridization of the linking nucleic acid molecule 111 to the capture moiety 105 (hybridization not shown) or by using a splint oligonucleotide 113 comprising a sequence complementary to the sequences of the linking nucleic acid molecule 111 and the capture moiety 105. In some cases (not shown), the linking nucleic acid molecule 111 comprises a self-splinting sequence such that the linking nucleic acid molecule 111 can be coupled to the capture moiety 105 in the absence of a separate splint molecule. A ligase can be used to covalently link the linking nucleic acid molecule 111 to the capture moiety 105. Alternatively, the linking nucleic acid molecule 111 can comprise a first reactive portion (e.g., a click chemistry portion, not shown) that can react with a second reactive portion (not shown) of the capture moiety 105. In Figure 1AIn sub-figure D, the monomer-capture moiety complex is subjected to conditions sufficient to cleave the terminal monomer (e.g., amino acid) from the polymer analyte 103 (e.g., peptide), thereby generating a detectable complex comprising the cleaved monomer. The conditions may include performing an Edman or Edman-like degradation reaction. Cleavage of the monomer from the polymer analyte produces a cleaved monomer-capture moiety complex comprising the cleaved monomer, sequencing reagent 109, linker nucleic acid molecule 111, and capture moiety 105. Figure 1A In sub-figure E, a binding agent 115 (e.g., an antibody) comprising another polymerizable molecule 117 (e.g., a nucleic acid molecule) is provided and contacted with a detectable complex. The binding agent 115 can be specific for a monomer (e.g., for an amino acid type) or for a sequencing reagent-monomer complex (e.g., a sequencing reagent-amino acid complex). In some cases, when the monomer is still attached to the polymer analyte, the binding agent 115 recognizes and binds to the sequencing reagent-monomer complex, rather than the monomer. The polymerizable molecule 117 of the binding agent can contain information about the identity of the binding agent or the specific monomer (e.g., amino acid type) to which the binding agent binds. The polymerizable molecule 117 of the binding agent can contain additional sequences, such as barcode sequences, UMIs, restriction sites, transposition sites, sequences to indicate the number of cycles or iterations, or other functional sequences. The polymerizable molecule 117 of the binding agent can be coupled to another polymerizable molecule 107 coupled to the substrate 101. In some cases, an extension reaction can be performed (e.g., using a polymerase) to copy the sequence of the polymerizable molecule 117 of the binding agent to the additional polymerizable molecule 107 coupled to the substrate 101. Alternatively, the polymerizable molecule 117 of the binding agent can be chemically (e.g., by complementary click chemistry) or enzymatically (e.g., using a ligase, ribozyme, or deoxyribozyme) linked to the additional polymerizable molecule 107. Optionally, the polymerizable molecule 117 of the binding agent can be cleaved from the binding agent 115 (not shown). Figure 1A In sub-graph F, the monomer of the detectable complex can be decoupled from the capture portion 105 (e.g., removed or cut). For example, the monomer, sequencing reagent 109, and all or part of the connected nucleic acid molecules 111 can be cut (depicted as a star). Cutting can be performed chemically, mechanically, or enzymatically. In the example of enzymatic cutting, the connected nucleic acid molecules 111 can include restriction sites or other cutting sites (e.g., uracil), and the cutting occurs by introducing a restriction enzyme or a cutting enzyme (e.g., uracil DNA glycosylase, nickase) to cut the restriction / cutting site. Alternatively, the cut monomer-capture portion complex (not shown) can be blocked with a blocking agent. The workflow 100 can then be iterated or repeated to sequence all or part of the polymer analyte 103.

[0180] Figure 1BAnother example of processing and characterizing polymer analyte 103 (for example, peptide) is schematically shown.Polymer analyte 103 can be labeled (or provided as pre-labeled) with capture portion 105 (for example, polymerizable molecule, such as nucleic acid molecule).Capture portion 105 can, for example, be attached to the end (for example, as shown at C-terminal end) or internal residue of peptide.In some cases, capture portion 105 can include barcode sequence, for example, the barcode of identification polymer analyte, sample source, partition or compartment etc.In some cases, capture portion 105 can not be attached to peptide, but can be associated with peptide (for example, by indirect interaction).In other examples (not shown), peptide (directly or indirectly) is coupled to non-nucleic acid molecule (for example, polymerizable molecule, such as other peptide, or other detectable label, such as mass label, fluorophore, radioisotope etc.).Provide and comprise the sequencing reagent 109 of connection nucleic acid molecule 111, for example, the sequencing reagent of formula I-III. The linker nucleic acid molecule 111 can comprise any useful sequence, such as a primer sequence, a barcode sequence, a UMI, a restriction site, and the like, and can comprise a capture-binding moiety. In some cases, the linker nucleic acid molecule 111 comprises a cleavable moiety, such as a restriction site, an abasic site, uracil, a transposition site, and the like. In process 110, the sequencing reagent 109 is coupled to a monomer (e.g., a terminal amino acid) of a polymer analyte to generate a sequencing reagent-monomer complex. Before, during, or after the coupling of the sequencing reagent to the monomer, the linker nucleic acid molecule 111 can be coupled to the capture moiety 105 via the capture-binding moiety, such as by hybridization (not shown), ligation (shown as process 112), or splint ligation (not shown), thereby generating a monomer-capture moiety complex. In process 113, the monomer is cleaved (e.g., chemically or enzymatically) from the polymer analyte 103, thereby providing a detectable complex comprising the cleaved monomer-capture moiety complex or a portion thereof. The cleaved monomer-capture moiety complex comprises a cleaved monomer coupled to a sequencing reagent 109, a linker nucleic acid molecule 111, and a capture moiety 105. In some cases, the cleaved monomer-capture moiety complex remains coupled to the polymeric analyte 103 via the linker nucleic acid molecule 111 and the capture moiety 105. The detectable complex comprising the cleaved monomer-capture moiety complex (or a portion thereof) can be detected directly or further processed for downstream analysis.

[0181] Downstream processing and analysis may include sorting, detection, or both. In some cases, a plurality of binding agents 115 are provided; the binding agent 115 can recognize and bind to different monomer types (e.g., different amino acid types). In some cases, the plurality of binding agents 115 recognize and bind to the sequencing reagent-monomer complex (e.g., recognize and bind to different monomer types contained by the sequencing reagent-monomer complex). In some cases, when the monomer is still attached to the polymer analyte, the plurality of binding agents 115 recognize and bind to the sequencing reagent-monomer complex, rather than the monomer. The binding agent 115 can be contacted with a plurality of cleaved monomer-capture portion complexes containing different monomer types (e.g., different amino acids) and bind to their respective targets. In some cases (not shown), the binding agent 115 comprises a polymerizable molecule, such as a nucleic acid molecule, that identifies the binding agent or its homologous molecule; the polymerizable molecule can be coupled or transferred to the capture portion 105 (not shown), for example, by nucleic acid extension, connection, transposition, etc. Alternatively or additionally, in process 119, the binding agents 115 can be separated or sorted into separate compartments (not shown), for example, using complementary nucleic acid sequences of the binding agent's nucleic acid molecules. Alternatively or additionally, the binding agents can comprise sorting tags, such as reporter molecules, mass tags, fluorophores, or fluorescent proteins, which can enable sorting of different binding agent types. In one such example, a first binding agent for a first monomer type (e.g., an amino acid residue) can comprise a GFP tag and a second binding agent for a second monomer type can comprise an RFP tag that can be sorted by fluorescence (e.g., using FACS) or affinity sorting (e.g., using beads with anti-GFP and anti-RFP antibodies).

[0182] In some cases, after process 119, sequencing reagent 109 or linked nucleic acid molecule 111 or a portion thereof can be removed from the sequencing reagent-monomer complex or cleaved monomer-sequencing reagent complex, for example, by restriction digestion or cleavage of the uracil of linked nucleic acid molecule 111 (e.g., using UDG or USER enzymes).

[0183] In some cases, barcoding of the cleaved monomer-capture moiety complex can be performed. For example, as described above, the binding agent 115 can contain a polymerizable barcode molecule, such as a nucleic acid barcode molecule, that identifies the binding agent or its homologous molecule; the polymerizable barcode molecule can be coupled or transferred to the capture moiety 105 (not shown), thereby barcoding the capture moiety. Alternatively or additionally, after sorting in process 119, the sorted cleaved monomer-capture moiety complex can be barcoded. For example, after sorting, the cleaved monomer-capture moiety complex can be divided by its corresponding monomer type (e.g., amino acid type). Since each compartment contains a known monomer type (e.g., amino acid type) based on the binding spectrum of the binding agent (e.g., specificity for a particular monomer type), the cleaved monomer-capture moiety complex within the compartment can be labeled with an identifying polymerizable molecule 117 (e.g., a nucleic acid barcode molecule) that contains the identity of the particular monomer type of the cleaved monomer-capture moiety complex, thereby generating a barcoded capture moiety.

[0184] In some cases, the polymerizable molecules 117 or the connecting nucleic acid molecules 111 can include time information (e.g., round number or cycle number). After any useful round number or iteration number, the other polymerizable molecules of barcoding can be removed (or the other polymerizable molecules of multiple barcoding) and sequenced, for example, using NGS methods to output the identity of each monomer type processed and the order or position in which they appear in the polymer analyte based on time information.

[0185] Figure 1C Another example workflow for sequencing polymer analytes is schematically shown. Multiple polymer analytes (e.g., peptides) can be coupled to substrate 101 along with multiple capture moieties. For illustrative purposes, further sequencing workflow operations for a single polymer analyte 103 are shown; however, it should be understood that Figures 1A-1FThe workflow operation can be carried out in parallel on a variety of polymer analytes. In process 110, a sequencing reagent 109 (for example, a sequencing reagent of Formula I-III) is provided and coupled to a monomer (for example, a terminal amino acid) to generate a sequencing reagent-monomer complex. Sequencing reagent 109 can be coupled to a monomer (for example, a terminal amino acid, such as an N-terminal amino acid) and (ii) capture portion 105 of (i) polymer units, which can be used to locally tether the monomer to the vicinity of the polymer analyte. In one example, sequencing reagent 109 can include an amino acid reactive group, such as a dithioester or a thiocarbamoyl group, which enables sequencing reagent to be coupled to the N-terminal amino acid. Sequencing reagent can additionally include a second reactive group, for example, which can be coupled to the capture portion and can be used as a capture-binding portion of a substrate-tethered portion. For example, the second reactive portion can include a click chemistry portion, which can be coupled to the click chemistry capture portion, for example, by an azide-alkyne or azide-cycloalkyne reaction. In other examples, the second reactive group can react with a third reactive group coupled to a nucleic acid molecule, which can also be coupled to a nucleic acid capture moiety (e.g., such as a nucleic acid capture moiety) by, for example, hybridization, ligation, or both. Figure 1A As shown). In process 112, the capture-binding portion of the sequencing reagent can be coupled to the capture portion. In process 113, a monomer can be cut from the polymer analyte to generate a detectable complex. In some embodiments, the cutting of the monomer can be mediated by stimulation, such as a chemical reaction or pH change (such as adding an acid). After cutting, a binding agent 115 is provided. The binding agent, such as an antibody or antibody fragment, can be specific to a specific monomer in a variety of monomers, such as a specific amino acid type or its derivative. In some cases, when the monomer is still attached to the polymer analyte, the binding agent 115 recognizes and binds to the sequencing reagent-monomer complex, rather than the monomer. The binding agent can include a detectable label (e.g., a fluorophore, a radioisotope, a mass label, etc.). The detectable label can be detected (e.g., using microscopy or imaging). The identity of the cut terminal monomer (e.g., amino acid) is determined. Subsequently, the detectable complex or a portion thereof, such as the sequencing reagent or sequencing reagent-monomer complex, can be removed or cut, and the process can be repeated or iterated to sequence the remaining monomers of the polymer analyte.

[0186] Figure 1DAnother workflow for sequencing a polymer analyte is shown. A polymer analyte 103 and a capture portion 105 are provided, which are optionally coupled to a substrate 101. The capture portion 105 may comprise a first nucleic acid molecule (e.g., DNA). In process 106, a sequencing reagent 109 comprising a polymerizable molecule (e.g., a connection nucleic acid molecule 111) is provided. In some cases, the sequencing reagent 109 is pre-tethered to the polymerizable molecule (depicted as a connection nucleic acid molecule 111); alternatively, the sequencing reagent 109 and the polymerizable molecule may be provided separately. In process 106, the sequencing reagent 109 may be coupled to a monomer, such as an amino acid (e.g., NTAA) of the polymer analyte 103 (e.g., a peptide), to generate a sequencing reagent-monomer complex. In process 112, the sequencing reagent-monomer complex may be coupled to the capture portion 105, thereby generating a monomer-capture portion complex. The coupling of the sequencing reagent-monomer complex to the capture portion 105 may be mediated by a polymerizable molecule (e.g., a connection nucleic acid molecule 111). Optionally, the sequencing reagent-monomer complex and the capture moiety 105 can be covalently linked together using chemical (e.g., click chemistry) or enzymatic (e.g., ligase) methods (e.g., linker nucleic acid molecule 111 can be covalently linked to the capture moiety 105). Alternatively or additionally, the polymerizable molecule can comprise a complementary first sequence and can hybridize to a second sequence of the capture moiety 105 (not shown), or the polymerizable molecule can be linked to the capture moiety 105 via a splint or bridging molecule that can comprise a sequence complementary to the first sequence of the polymerizable molecule and the second sequence of the capture moiety 105 (not shown). In process 113, the monomer is cleaved from the polymer analyte 103, thereby providing a detectable complex comprising a cleaved monomer-capture moiety complex comprising the cleaved monomer coupled to the sequencing reagent 109, the polymerizable molecule (shown as linker nucleic acid molecule 111), and the capture moiety 105 or a portion thereof. In process 114, a binding agent 115 (e.g., an antibody, a binding protein, etc.) can be contacted with the detectable complex. The binding agent can be configured to recognize all or part of the cleaved monomer-capture moiety complex. For example, the binding agent can recognize a monomer, a sequencing reagent-monomer complex, or an entire monomer-capture moiety complex. In one example, the sequencing reagent can include a dithioester or thiocarbamoyl moiety, and the binding agent can recognize a dithioester or thiocarbamoyl group that reacts with an amino acid or a derivative thereof. In some cases, the binding agent 115 can include a detectable portion (not shown), or can be contacted with another binding agent (e.g., a second antibody) that can optionally include a detectable portion (not shown).

[0187] Any of the processes (e.g., 106, 112, 113, or 114) can be iterated and repeated any number of times ("rounds") using additional sequencing reagents 109 and polymerizable molecules (optionally containing cycle / round information) and tethering the additional polymerizable molecules together (e.g., tethering the additional polymerizable molecules to the polymerizable molecules of the monomer-capture moiety complex). Multiple rounds can be performed until all monomers or a subset of monomers in the polymer analyte 103 are cleaved and tethered together. In some cases, the processes 106, 112, and 113 can be iterated to generate a stacked polymerizable molecule 123 comprising a set of cleaved monomers, such as a set of tandem sequencing reagent-monomer-polymerizable molecule complexes. The stacked polymerizable molecules can then be contacted with a library of binders that can bind to their respective monomer targets (e.g., amino acid types).

[0188] Further downstream analysis can be performed, for example, using a nanopore or nanogap system. In one such example, stacked polymerizable molecules 123, optionally coupled to a binding agent, can be prepared and translocated through a nanopore sequencing system, which can output the identity of the polymerizable molecule (e.g., nucleic acid sequence), the monomer type, and the individual binding agents (if present).

[0189] Alternatively, in some cases, no binding agent may be required to sequence the polymerizable molecules. Figure 1E An example workflow for sequencing polymer analytes is schematically shown. In such an example workflow, Figure 1D , omitting process 114. This workflow can be iterated multiple times to generate stacked polymerizable molecules 123, which can be further processed, for example, analyzed using a nanopore or nanogap sequencing system, to output the identities of the polymerizable molecules and individual monomers.

[0190] Similarly, Figure 1FAnother example workflow for sequencing a polymer analyte in the absence of a binding agent is schematically shown, and it can be prepared in the absence of a substrate. In such an example, a polymer analyte 103 and a capture portion 105 are provided. The capture portion 105 can include a first nucleic acid molecule (e.g., a DNA molecule) and can include identification information of the polymer analyte 103, such as an identification barcode sequence. The capture portion 105 can additionally include a releasable or cleavable portion. In process 106, a sequencing reagent 109 and a polymerizable molecule are provided, such as a connecting nucleic acid molecule 111. In some cases, the sequencing reagent 109 is pre-tethered to the polymerizable molecule (connecting nucleic acid molecule 111); alternatively, the sequencing reagent 109 and the polymerizable molecule (connecting nucleic acid molecule 111) can be provided separately. The polymerizable molecule can include identification time information, for example, the cycle or round in which the polymerizable molecule is provided. In process 106, the sequencing reagent 109 can be coupled to a monomer, such as an amino acid (e.g., NTAA) of the polymer analyte 103, to generate a sequencing reagent-monomer complex. In process 112, the sequencing reagent-monomer complex can be coupled to the capture portion 105. The coupling of the sequencing reagent-monomer complex to the capture portion 105 can be mediated by a polymerizable molecule and optionally an additional polymerizable molecule 116. Optionally, the sequencing reagent-monomer complex and the capture portion can be covalently linked together (e.g., by ligation). Alternatively or additionally, the polymerizable molecule can comprise a complementary first sequence and can hybridize to a second sequence of the capture portion 105 (not shown), or the polymerizable molecule can be linked to the capture portion 105 by a splint or bridging molecule that can comprise a sequence complementary to the first sequence of the polymerizable molecule and the second sequence of the capture portion 105 (not shown). In process 113, a monomer can be cleaved from the polymer analyte 103 to generate a monomer-capture portion complex comprising the cleaved monomer, the sequencing reagent 109, the polymerizable molecule (e.g., the ligated nucleic acid molecule 111), and the capture portion 105. Processes 106, 112, and 113 can be iterated and repeated any number of times ("rounds") using additional sequencing reagents 109 and polymerizable molecules and tethering the additional polymerizable molecules together (e.g., tethering the additional polymerizable molecules to the monomer-capture moiety complex). Multiple rounds can be continued until all monomers or a subset of monomers of the polymer analyte 103 are tethered together. For example, the process can be iterated to produce a stacked polymerizable molecule 123 comprising a set of cleaved monomers, such as a set of tandem monomer-sequencing reagent-polymerizable complexes. After any useful number of rounds, the stacked polymerizable molecule 123 can be cleaved from or at the capture moiety 105, for example, using a cleavable moiety. The cleaved product can then be sequenced using a nanopore or nanogap method.

[0191] In some cases, a polymerizable molecule, such as linker nucleic acid molecule 111, contains temporal information about the cycle in which the polymerizable molecule was provided; thus, this temporal information can be used for quality control. For example, if a missing cycle number is missing, it can be inferred that a certain amino acid is missing or absent in the peptide, that cleavage of an amino acid did not occur, or that some other error occurred. Alternatively or additionally, a temporal barcode can be provided separately and attached to the polymerizable molecule at any useful or convenient step.

[0192] Iteration: In some cases, one or more of the operations described herein may be iterated or repeated. Iterative operations may allow for the sequential processing, analysis, or identification of individual monomers of a polymer analyte, which may allow for the reconstruction of the entire polymer analyte. For example, referring to Figure 1A , the operations of workflow 100 can be performed to encode the identity of the terminal amino acid (e.g., NTAA) onto the additional polymerizable molecule 107 (e.g., via polymerizable molecule 117). The operations of workflow 100 can then be repeated to encode the identity of the n-1 terminal amino acid, the n-2 terminal amino acid, the n-3 terminal amino acid, and so on, until all or part of the peptide has been processed. Encoding can occur on the same (additional) polymerizable molecule 107, for example, to generate a stacked polymerizable molecule comprising multiple polymerizable molecules from multiple binding agents, or encoding can occur on additional polymerizable molecules (not shown) present on the substrate. In the former case, in some cases, the polymerizable molecule of the second (or third, fourth, fifth, ... nth) cycle can be configured to couple only to the first (or second, third, fourth, ... n-1th) polymerizable molecule. For example, the first cycle binding agent polymerizable molecule can contain a unique binding sequence that is not present on the additional polymerizable (or capture) molecule of the substrate, and the second cycle binding agent polymerizable molecule can bind to this binding sequence. Thus, the second cycle binder polymerizable molecules can bind only to the first cycle binder polymerizable molecules and not to any additional polymerizable (or capture) molecules of the substrate. In the event that no binding occurs (a "null" event), a bridging polymerizable molecule can be provided that encodes the null binding event but contains a unique binding sequence so that subsequent rounds can continue even if the binder does not bind to the cleaved monomer.

[0193] Similarly, Figures 1B-1F Any of the operations depicted in can be iterated to sequentially analyze all or a subset of monomers of a polymer analyte. For example, ref. Figure 1B, operations of this workflow can be performed to encode the identity of a terminal monomer (e.g., NTAA) onto another polymerizable molecule (not shown) or onto a barcoded capture portion (e.g., via polymerizable molecule 117). These operations can be repeated to encode the identity of the n-1 terminal amino acid, the n-2 terminal amino acid, the n-3 terminal amino acid, and so on, until all or part of the peptide has been processed. Encoding can occur on the same capture portion 105, for example, to generate a stacked polymerizable molecule comprising multiple barcoded polymerizable molecules 117, or encoding can occur on another polymerizable molecule (not shown). In the former case, in some cases, the polymerizable molecules of the second (or third, fourth, fifth, ... nth) cycle can be configured to couple only to the first (or second, third, fourth, ... n-1th) polymerizable molecule, as described above.

[0194] Where one or more additional polymerizable molecules are used, the polymerizable molecule 117 (e.g., coupled to a binding agent or provided separately after sorting) can additionally encode temporal information, such as a cycle number or iteration number, such that the order of individual monomers can be determined. For example, for a given peptide, the terminal amino acid can be coupled to a capture moiety and cleaved, and then contacted with a binding agent comprising a barcode sequence that identifies (i) the identity of the amino acid (e.g., any of the twenty proteinogenic amino acids) and (ii) the cycle number (e.g., cycle 1) (not shown). The information encoded by the barcode sequence can be coupled to an adjacent (additional) polymerizable molecule (not shown) or a capture moiety 105 or a barcoded capture moiety. After a monomer has been cleaved from the capture moiety (e.g., as Figure 1B ), the workflow can be repeated for the n-1 terminal amino acid, which can again be coupled to a capture moiety, cleaved, and contacted with a binder comprising an additional barcode sequence that identifies (i) the identity of the amino acid (e.g., any of the twenty proteinogenic amino acids) and (ii) the cycle number (e.g., cycle 2). The information encoded by the additional barcode sequence can be transferred to the same capture moiety 105 or barcoded capture moiety, or to an additional polymerizable molecule (not shown). In the former case, the polymerizable molecule can therefore contain information about: (i) the identity of the terminal amino acid, (ii) the cycle number of the terminal amino acid (cycle 1), (iii) the identity of the n-1 terminal amino acid, and (iv) the cycle number of the n-1 terminal amino acid (cycle 2), and so on. Alternatively or additionally, temporal information can be provided about another molecule, such as a capture moiety 105, a linking nucleic acid molecule 111, an additional polymerizable molecule 116, or the like.

[0195] Alternatively or additionally, the binding agent can contain a sorting tag and can provide a barcode sequence after sorting. For example, the binding agent can be used to sort different cleaved amino acid-sequencing reagent complexes (as shown in process 119), and a polymerizable molecule 117 containing the identity of the amino acid type can be provided for each sorted cleaved amino acid-sequencing reagent complex. The polymerizable molecule 117 can also contain time information, for example, the round or cycle in which the molecule is provided.

[0196] In some cases, the time information can be provided separately. For example, when the polymerizable molecule 117 containing the barcode information is coupled to the other polymerizable molecule 107 ( Figure 1A ) or capture part ( Figure 1B ) before, during or after, a time barcode can be provided, which can be coupled to the polymerizable molecule 117 ( Figures 1A-1B ), another polymerizable molecule 107 ( Figure 1A ), capture portion 105 ( Figure 1B 、 1D -1F) or a combination thereof. The time barcode may include any useful agent, including nucleic acid molecules, peptides, lipids, carbohydrates, enzymes (e.g., chromogenic enzymes or luciferases) or ribozymes or deoxyribozymes, fluorophores, dyes, intercalators, dideoxynucleotides, fluorescent nucleic acid molecules or nucleotides, radioisotopes, mass tags or other detectable markers that can indicate the time or cycle (or iteration) number it provides. In some cases, the time barcode includes a cycle-specific nucleic acid barcode molecule that can be coupled to a polymerizable molecule 117 (containing the identity of a monomer) or an end polymerizable molecule of a stacked polymerizable molecule that includes polymerizable molecules from multiple rounds or iterations. The time barcode may include any additional useful functional sequences, such as primer sites, sequencing sites, restriction sites, abasic or cleavable sites, etc. In some cases, the time barcode may include an amplification site that allows the time barcode and optionally coupled polymerizable molecules to be bridge-amplified with other capture or polymerizable molecules.

[0197] Polymer analytes: Polymer analytes can be biomolecules, macromolecules, or synthetic molecules. Polymer analytes can be biomolecules or other biomolecules comprising one or more monomers. Non-limiting examples of polymer biomolecules include nucleic acid molecules (e.g., DNA molecules, RNA molecules, DNA:RNA hybrids, aptamers), peptides and proteins, polysaccharides, lipid polymers (e.g., diglycerides, triglycerides, and other fatty acids). Polymer analytes can be synthetic molecules, such as peptidomimetics or synthetic polymers, or peptide mimetics (e.g., peptidomimetics, β-peptides, D-peptide peptidomimetics). Non-limiting examples of synthetic polymers include acrylic acid, nylon, silicone, viscose, rayon, polyester, polycarboxylic acid, polyvinyl acetate, polyacrylamide, polyacrylate, polyethylene glycol, polyurethane, polylactic acid, silicon dioxide, polystyrene, polyacrylonitrile, polybutadiene, polycarbonate, polyethylene terephthalate, poly(chlorotrifluoroethylene), poly(ethylene oxide), poly(vinyl terephthalate), polyethylene, polyisobutylene, poly(methyl methacrylate), poly(oxymethylene), polyoxymethylene, polypropylene, polystyrene, poly(tetrafluoroethylene), poly(vinyl acetate), poly(vinyl alcohol), poly(vinyl chloride), poly(vinylidene chloride), poly(vinylidene fluoride), poly(vinyl fluoride), or combinations thereof. Polymeric analytes can include a single polymer type (e.g., a homopolymer) or more than one polymer type (e.g., a copolymer) and can include random or arranged monomers. The polymer analyte can be a block polymer, an alternating copolymer, a periodic copolymer, a statistical copolymer, a stereoregular block copolymer, a gradient copolymer, a branched copolymer, a graft copolymer, and the like.

[0198] Polymer analytes can be in any size or include a series of sizes.The size of polymer analytes can be about 1 nanometer (nm), about 5nm, about 10nm, about 20nm, about 30nm, about 40nm, about 50nm, about 60nm, about 70nm, about 80nm, about 90nm, about 100nm, about 200nm, about 300nm, about 400nm, about 500nm, about 600nm, about 700nm, about 800nm, about 900nm, about 1 micron (μm), about 10 μm, about 100 μm, about 1 millimeter (mm) or larger.Multiple polymer analytes can include the polymer analytes in similar size or size range, for example, about 10nm to about 100nm, about 50nm to about 1 μm.Similarly, polymer analytes can have any molecular weight or molecular weight range. The polymeric analyte can be about 10 Daltons (Da), 100 Da, 500 Da, 1 kiloDalton (kDa), 10 kDa, 100 kDa, 1,000 kDa, 10,000 kDa, 100,000 kDa, or greater.The polymeric analyte can include polymeric analytes of similar molecular weight or within a range of molecular weights.

[0199] The monomer of polymer analyte can include any size or size range that is less than the size of whole polymer analyte.The size of monomer can be about 0.1 nanometer (nm), about 0.5nm, about 1nm, about 5nm, about 10nm, about 20nm, about 30nm, about 40nm, about 50nm, about 60nm, about 70nm, about 80nm, about 90nm, about 100nm, about 200nm, about 300nm, about 400nm, about 500nm, about 600nm, about 700nm, about 800nm, about 900nm, about 1 micron (μm), about 10 μm, about 100 μm, about 1 millimeter (mm) or larger.Monomer can have any molecular weight or molecular weight range. Monomers can be about 1 Dalton (Da), 10 Da, 100 Da, 500 Da, 1 kilodalton (kDa), 10 kDa, 100 kDa, 1,000 kDa, 10,000 kDa, 100,000 kDa, or greater. Monomer or polymeric analytes can range in molecular weight; for example, polymeric analytes can include peptides comprising amino acid monomers, the molecular weight of which can vary from 75 Da (glycine) to 204 Da (tryptophan).

[0200] The polymer analyte can include any number of monomers. The polymer analyte can contain about 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1,000, 2,000, 3,000, 4,000, 5,000, 6,000, 7,000, 8,000, 9,000, 10,000, 50,000, 100,000 or more monomers. The polymeric analyte may comprise at least about 2, at least about 3, at least about 4, at least about 5, at least about 6, at least about 7, at least about 8, at least about 9, at least about 10, at least about 20, at least about 30, at least about 40, at least about 50, at least about 60, at least about 70, at least about 80, at least about 90, at least about 100, at least about 500, at least about 1,000, at least about 5,000, at least about 10,000, at least about 50,000, at least about 100,000, or more monomers. Alternatively, the polymeric analyte may comprise at most about 100,000, at most about 50,000, at most about 10,000, at most about 5,000, at most about 1,000, at most about 500, at most about 100, at most about 50, at most about 10, at most about 5 or fewer monomers. A polymeric analyte may comprise a range of monomers; for example, a polymeric analyte may comprise about 5 monomers, while another polymeric analyte may comprise about 500 monomers.

[0201] In some cases, polymer analytes include peptides comprising amino acid monomer units. The peptide can be naturally occurring or synthetic. The peptide can comprise any number of amino acids. The amino acid can be one of the 20 proteinogenic amino acids and can comprise any number of post-translational modifications. The peptide or any of its constituent amino acids can be processed, for example, by contacting with a protecting group, alkylating, beta-eliminating a phosphate group, etc., as described elsewhere herein. In some cases, the peptide is derived from a larger peptide or protein and is fragmented.

[0202] Substrate: One or more operations described herein can be performed using a substrate. For example, one or more molecules described herein (e.g., polymer analytes, capture moieties, polymerizable molecules) can be coupled to a substrate. In some cases, a polymer analyte, a capture moiety, and one or more polymerizable molecules (e.g., a first polymerizable molecule or a second polymerizable molecule) coupled to one or more substrates, or a combination thereof, can be provided. In one example, a polymer analyte, a capture moiety, and a second polymerizable molecule are coupled to a substrate. In some cases, more than one substrate can be used. In such cases, the substrates can comprise the same material or different materials.

[0203] The substrate can be made of any suitable material, e.g., glass, silicon, gel, polymer, etc., as described elsewhere herein. In some cases, the substrate can be beads or gel beads (e.g., polyacrylamide, agarose, or Beads). The substrate can be functionalized. One or more molecules such as capture moieties and polymeric analytes (e.g., peptides) can be coupled to the substrate via covalent or non-covalent interactions. The capture moiety and polymer analyte (e.g., peptide) can be coupled to the substrate using any suitable chemistry, such as a click chemistry moiety (e.g., alkyne-azide coupling), a photoreactive group (e.g., benzophenone), 1-ethyl-3-(3-dimethylaminopropyl)carbodiimide hydrochloride (EDC) (e.g., to couple amino-oligonucleotides or peptides), N-hydroxysulfosuccinimide (NHS), sulfo-NHS or NHS-ester (e.g., to couple thiol oligonucleotides), maleimide, hydrazine, hydroxylamine, thiol, biotin-streptavidin interaction, cystamine, glutaraldehyde, formaldehyde, 4-(N-maleimidomethyl)cyclohexane-1-carboxylic acid succinimidyl ester (SMCC), sulfo-SMCC, 4-(4,6-dimethoxy-1,3,5-triazin-2-yl)-4-methylmorpholinium chloride (DMTMM), silanes (e.g., aminosilanes), combinations thereof, and the like. In some cases, the substrate can be functionalized to include a coupling chemistry to couple a polymer analyte or capture moiety. In one non-limiting example, a substrate (e.g., a bead or surface) can include an alkyne, such as dibenzocyclooctyne (DBCO), which can be configured to react with an amine (e.g., DBCO-alcohol, DBCO-Boc, DBCO-NHS), a carboxyl group (carboyxl) or a carbonyl group (e.g., DBCO, DBCO-silane), a sulfhydryl group, etc. An azide-functionalized nucleic acid or protein can react with DBCO to link the nucleic acid or protein to the DBCO substrate. In other examples, linkers such as bifunctional linkers can be used to attach molecules to substrates; such bifunctional linkers can include the same reactive moiety at both ends or different moieties at each end (e.g., heterobifunctional linkers).

[0204] In some cases, a molecule (e.g., a polymeric analyte, a capture moiety, a polymerizable molecule) can be coupled to a substrate using an enzymatic method, e.g., as described elsewhere herein. For example, an enzyme can be used to attach a chemical linker or moiety (such as a click chemistry moiety) to a polymeric analyte (e.g., a peptide). The chemical linker or moiety can be capable of reacting with another chemical linker or moiety (e.g., a click chemistry moiety) of a substrate, a capture moiety, or a polymerizable molecule.

[0205] The substrate can be coupled to any useful number of molecules (e.g., polymeric analytes, capturing moieties, polymerizable molecules). In some cases, the substrate can include multiple polymeric analytes, multiple capturing moieties, and / or multiple polymerizable molecules, which can be provided in any useful ratio or density. For example, the ratio of polymeric analytes to capturing moieties or polymerizable molecules can be about 1:1, 1:5, 1:10, 1:20, 1:100, 1:1000, 1:10,000, 1:100,000, 1:1,000,000, or less. In some cases, the ratio of polymer analyte to capture moiety or polymerizable molecule can be at most about 1:1, at most about 1:5, at most about 1:10, at most about 1:20, at most about 1:100, at most about 1:1000, at most about 1:10,000, at most about 1:100,000, at most about 1:1,000,000 or less.

[0206] Similarly, molecules (eg, polymeric analytes, capture moieties, or polymerizable molecules) can be coupled to the substrate at any useful density, such as about 1 molecule per square micrometer (μm). 2 ), about 10 molecules / μm 2 , about 100 molecules / μm 2 , about 1,000 molecules / μm 2 , about 10,000 molecules / μm 2 , about 100,000 molecules / μm 2 , about 1,000,000 molecules / μm 2 , about 10,000,00 molecules / μm 2 , about 100,000,000 molecules / μm 2 , about 1,000,000,000 molecules / μm 2 , about 10,000,000,000 molecules / μm 2 , about 100,000,000,000 molecules / μm 2 or greater. The polymer analyte, capture moiety, and polymerizable molecule can be coupled to the substrate at a range of densities, for example, from about 100 to about 10,000 molecules / μm 2 or about 10 to about 1,000 molecules / μm 2 The densities of the polymeric analyte, capture moiety, and polymerizable molecule can be the same or different. For example, the density of the polymerizable molecule can be 1 / 1, 1 / 2, 1 / 3, 1 / 4, 1 / 5, 1 / 6, 1 / 7, 1 / 8, 1 / 9, 1 / 10, 1 / 100, 1 / 1000, 1 / 10,000, 1 / 100,000, 1 / 1,000,000, or less than the density of the polymeric analyte.

[0207] In some cases, the molecule that is coupled to substrate can be spaced apart with specified or controlled distance.For example, the average interval or the distance between the polymerizable molecules that are coupled to substrate can be about 1 nanometer (nm), about 2nm, about 3nm, about 4nm, about 5nm, about 6nm, about 8nm, about 9nm, about 10nm, about 20nm, about 30nm, about 40nm, about 50nm, about 60nm, about 70nm, about 80nm, about 90nm, about 100nm, about 500nm, about 1 μm or larger.In some cases, the interval between the polymerizable molecules that are coupled to substrate can be about 1 μm at most, about 500nm at most, about 100nm at most, about 90nm at most, about 80nm at most, about 70nm at most, about 60nm at most, about 50nm at most, about 40nm at most, about 30nm at most, about 20nm at most, about 10nm at most, about 5nm at most or less. Similarly, the separation or distance between the polymeric analyte and the polymerizable molecule or capture moiety can be about 1 nanometer (nm), about 2 nm, about 3 nm, about 4 nm, about 5 nm, about 6 nm, about 8 nm, about 9 nm, about 10 nm, about 20 nm, about 30 nm, about 40 nm, about 50 nm, about 60 nm, about 70 nm, about 80 nm, about 90 nm, about 100 nm, about 500 nm, about 1 μm or more. In some cases, the average separation between the polymerizable molecule and the polymeric analyte coupled to the substrate can be at most about 1 μm, at most about 500 nm, at most about 100 nm, at most about 90 nm, at most about 80 nm, at most about 70 nm, at most about 60 nm, at most about 50 nm, at most about 40 nm, at most about 30 nm, at most about 20 nm, at most about 10 nm, at most about 5 nm or less. Average distances between polymerizable molecules and each other or to the polymer analyte can range from, for example, about 1 nm to about 40 nm, about 2 nm to about 10 nm, and the like.

[0208] The concentration or density of molecules attached to the substrate can be adjusted using one or more suitable methods, including patterned or random deposition methods. Examples of methods for controlling the concentration or density of molecules attached to the substrate include limiting dilution, adding chaotropic agents (e.g., guanidine, formamide, urea), using metal organic compounds, etc. The molecules can be attached to the substrate in a patterned manner, such as using a self-assembled monolayer, photopatterning, lithography, etching, or a combination thereof, or the molecules can be randomly arranged.

[0209] The substrate can comprise any useful size or dimension (e.g., length, width, height, diameter, radius), surface area, volume, or ratios or combinations thereof. The substrate can comprise beads or particles having a diameter of about 1 nanometer (nm), about 2 nm, about 3 nm, about 4 nm, about 5 nm, about 6 nm, about 8 nm, about 9 nm, about 10 nm, about 20 nm, about 30 nm, about 40 nm, about 50 nm, about 60 nm, about 70 nm, about 80 nm, about 90 nm, about 100 nm, about 500 nm, about 1 μm, about 2 μm, about 3 μm, about 4 μm, about 5 μm, or about 6 μm. The surface area of ​​the substrate can be about 1 mm, about 2 mm, about 3 mm, about 4 mm, about 5 mm, about 6 mm, about 7 mm, about 8 mm, about 9 mm, about 10 mm, about 20 mm, about 30 mm, about 40 mm, about 50 mm, about 60 mm, about 70 mm, about 80 mm, about 90 mm, about 100 mm, about 200 mm, about 300 mm, about 400 mm, about 500 mm, about 600 mm, about 700 mm, about 800 mm, about 900 mm, about 1 mm or more. The surface area of ​​the substrate can include about 1 square nanometer (nm) or more. 2 ), about 10nm 2 , about 100nm 2 , about 1,000nm 2 , about 10,000nm 2 , about 100,000nm 2 , about 1μm 2 , about 10μm 2 , about 100μm 2 , about 1,000μm 2 , about 10,000μm 2 , about 100,000μm 2 , about 1mm 2 , about 10mm 2 , about 100mm 2 , about 1,000mm 2 , about 10,000mm 2 , about 100,000mm 2 , about 1,000,000mm 2 or larger.

[0210] In some cases, molecules can be coupled to substrates in an ordered or random arrangement. In an ordered arrangement, any conventional method can be used for patterning molecules, such as lithography (e.g., soft lithography, photolithography), etching (e.g., ion etching, photoetching) or other patterning methods. In some cases, joints (e.g., bifunctional joints) can be used to promote the coupling of molecules (e.g., polymer analytes, polymerizable molecules, capture moieties) to substrates; such joints can be patterned using any useful technology (e.g., self-assembled monolayers, photopatterning, lithography, etching). In some cases, molecules can be coupled to substrates in a random arrangement. For example, molecules can be provided in a stoichiometric ratio or controlled concentration to couple molecules with any useful ratio or density.

[0211] Polymerizable molecules: The polymerizable molecules described herein can be any useful type of polymerizable molecule. The polymerizable molecules can be naturally occurring, such as biopolymers (e.g., nucleic acid molecules, peptides, polysaccharides, fatty acids) or other naturally occurring polymers, for example, rubber, cellulose, starch, polyhydroxyalkanoates, chitosan, dextran, structural proteins (e.g., collagen, hyaluronic acid, glycosaminoglycans), agarose, carrageenan, isphagula, gum arabic, agar, gelatin, shellac, xanthan gum, guar gum, alginates, and the like. The polymerizable molecules can be synthetic, for example, acrylic, nylon, silicone, viscose, rayon, polyester, polycarboxylic acid, polyvinyl acetate, polyacrylamide, polyacrylate, polyethylene glycol, polyurethane, polylactic acid, silica, polystyrene, polyacrylonitrile, polybutadiene, polycarbonate, polyethylene terephthalate, poly(chlorotrifluoroethylene), poly(ethylene oxide), poly(vinyl terephthalate), polyethylene, polyisobutylene, poly(methyl methacrylate), poly(oxymethylene), polyoxymethylene, polypropylene, polystyrene, poly(tetrafluoroethylene), poly(vinyl acetate), poly(vinyl alcohol), poly(vinyl chloride), poly(vinylidene chloride), poly(vinylidene fluoride), poly(vinyl fluoride), and combinations thereof. The polymerizable molecules can contain one or more reactive moieties (e.g., free radicals) to initiate polymerization or can be polymerizable via contact with an initiator (e.g., ammonium persulfate, peroxide, or other radicalizing agent). The polymerizable molecule can be polymerizable via a catalase (e.g., a polymerizing enzyme such as a polymerase), a ribozyme, or a deoxyribozyme. Alternatively or additionally, the polymerizable molecule can be polymerizable via self-assembly. The polymerizable molecule can include a single polymer type (e.g., a homopolymer) or more than one polymer type (e.g., a copolymer) and can include random or arranged monomers. The polymerizable molecule can be a block polymer, an alternating copolymer, a periodic copolymer, a statistical copolymer, a stereoregular block copolymer, a gradient copolymer, a branched copolymer, a graft copolymer, etc.

[0212] The same or different types of polymerizable molecules can be used for methods described herein. For example, the first polymerizable molecule contained in or coupled to the binding agent can be a nucleic acid molecule, and the second polymerizable molecule can be a peptide. In another example, both the first polymerizable molecule and the second polymerizable molecule are nucleic acid molecules. In such examples, the first polymerizable molecule can be coupled to the second polymerizable molecule via connection or hybridization. For example, the first polymerizable molecule can include a first nucleic acid sequence, and the second polymerizable molecule can include a second nucleic acid sequence. The first nucleic acid sequence can be complementary or partially complementary to the second nucleic acid sequence, and the coupling can include hybridizing the first nucleic acid sequence or its portion with the second nucleic acid sequence or its portion. Alternatively, the first nucleic acid sequence and the nucleic acid sequence can be complementary to the two sequences of the splint or bridging oligonucleotide, and the coupling can be mediated via hybridization with the splint oligonucleotide. The first nucleic acid sequence can be chemically (for example, via click chemistry, wherein the first polymerizable molecule and the second polymerizable molecule include a member of click chemistry pairs) or enzymatically (for example, using a ligase) connected to the second nucleic acid sequence.

[0213] The polymerizable molecule may comprise a functional moiety. For example, the polymerizable molecule may comprise a nucleic acid molecule comprising a functional sequence such as a primer sequence (e.g., a universal priming site), a sequencing sequence, a read sequence, a unique molecular identifier (UMI), a barcode sequence, a cleavage sequence (e.g., a restriction site, a Cas binding sequence), a transposition sequence (e.g., a chimeric end sequence), or a combination thereof.

[0214] In some cases, polymerizable molecule (for example, the first polymerizable molecule or the second polymerizable molecule) includes a nucleic acid molecule comprising one or more nucleotide bases. The polymerizable molecule can include any useful number of nucleotide bases, for example, about 1 base, about 2 bases, about 3 bases, about 4 bases, about 5 bases, about 6 bases, about 7 bases, about 8 bases, about 9 bases, about 10 bases, about 20 bases, about 30 bases, about 40 bases, about 50 bases, about 60 bases, about 70 bases, about 80 bases, about 90 bases, about 100 bases, about 200 bases, about 300 bases, about 400 bases, about 500 bases, about 600 bases, about 700 bases, about 800 bases, about 900 bases, about 1000 bases, or a greater number of bases.

[0215] Polymerizable molecules can include nucleic acid molecules. Nucleic acid molecules can be single-stranded, double-stranded or partially double-stranded. Nucleic acid molecules can include modified nucleotides or atypical bases. For example, polymerizable molecules can include pseudo-complementary bases, bridged nucleic acids (BNA), heterologous nucleic acids (XNA), locked nucleic acids (LNA), peptide nucleic acids (PNA), γ-PNA molecules, morpholinos or combinations thereof. In some cases, polymerizable molecules can include hexitol nucleic acids (HNA) or cyclohexane nucleic acids (CeNA), which can be used to make polymerizable molecules more resistant to acid degradation (for example, as used in traditional Edman degradation). Alternatively or additionally, polymerizable molecules can include naturally occurring bases that are more resistant to acid degradation, for example, mainly consisting of thymine or cytosine. For example, the nucleic acid molecule can comprise at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, or 100% thymine or cytosine, which can render the nucleic acid molecule more acid-resistant compared to nucleic acid molecules comprising adenine or guanine.

[0216] Linker: One or more operations of the method can be mediated using a linker. In some cases, the coupling of the monomer to the capture portion to generate a monomer-capture portion complex is mediated using a sequencing reagent that is or comprises a linker. In some cases, the sequencing reagent is a bifunctional, trifunctional, or multifunctional linker. The coupling of the linker to the monomer or capture portion can be covalent or non-covalent. In an example, the linker can include a first reactive group that is capable of coupling to a monomer of a polymer analyte (e.g., an amino acid of a peptide) and optionally cleaving an amino acid from the peptide. For example, the first reactive group can be an amino acid reactive group, for example, an isothiocyanate (ITC) such as phenyl isothiocyanate (PITC), 3-pyridyl isothiocyanate (PY1TC), 2-piperidylethyl isothiocyanate (PEITC), 3-(4-morpholino)propyl isothiocyanate (MPITC), 3-(diethylamino)propyl isothiocyanate (DEPTIC) or naphthyl isothiocyanate (NITC), fluorescein isothiocyanate (FITC), ammonium thiocyanate, potassium thiocyanate, trimethylsilyl isothiocyanate (TMS-ITC), phenyl phosphoroisothiocyanatidate, acetyl isothiocyanate (AITC) or an aldehyde group, for example, o-phthalaldehyde (OPA), 2,3-naphthalenedicarboxylic acid (NDA), 2-pyridinecarboxaldehyde, which can react with the N-terminal amino acid (NTAA). In some embodiments, the first reactive group includes dithioester or thiocarbamyl, as described elsewhere herein.Joint can additionally include the second reactive group that can be directly or indirectly coupled to capture portion.In the example of direct coupling, capture portion can include click chemistry moiety (for example, alkyne), and the second reactive group of joint can include other click chemistry moieties (for example, azide) that can react with the click chemistry moiety of capture portion.Alternately, joint can be indirectly coupled to capture portion, for example, via non-covalent interaction or via intermediate connecting molecule.In some cases, intermediate connecting molecule can include the 3rd polymerizable molecule (for example, polymer or nucleic acid molecule) that joint can be coupled to capture portion.In one such example, the 3rd polymerizable molecule can include the 3rd reactive group (for example, via alkyne-azide click chemistry) that (i) can be coupled to the second reactive group of joint and (ii) can be coupled to capture portion (for example, another orthogonal click chemistry reaction, avidin-biotin interaction, nucleic acid coupling or hybridization). In some cases, the third polymerizable molecule comprises a nucleic acid molecule comprising (i) a click chemistry moiety (e.g., alkyne) that can be conjugated to a first reactive group (e.g., azide) of a linker and (ii) a nucleic acid sequence that can be coupled to a capture moiety, e.g., via ligation, splinting, or hybridization.In some cases, the linker comprises a linker nucleic acid molecule comprising a self-splinting moiety.

[0217] Where applicable, the click chemistry portion of the linker and capture moiety or intermediate linker molecule can include any suitable bioorthogonal moiety, as described elsewhere herein, such as alkenes, alkynes, azides, epoxides, amines, thiols, nitrones, isonitriles, isocyanides, aziridines, activated esters, and tetrazines, as well as combinations, variants, or derivatives thereof. The linker can be subjected to conditions sufficient to react the first click chemistry moiety with the second click chemistry moiety, for example, providing a metal catalyst, a suitable solvent, pH, temperature, ion concentration, or light / energy for any useful duration.

[0218] The first reactive group of the linker can be an amino acid reactive portion. The amino acid reactive portion of the linker can be any useful portion that enables the reactive portion to be conjugated to an amino acid and optionally cleave an amino acid. In some examples, the first reactive portion can react with a terminal amino acid (e.g., NTAA or CTAA). In such examples, the first reactive moiety can comprise any primary amine or carboxyl reactive group, including but not limited to isocyanate, acyl azide, NHS ester, sulfonyl chloride, aldehyde, glyoxal, epoxide, ethylene oxide, carbonate, aryl halide, imido ester, carbodiimide, anhydride, phenyl ester, isothiocyanate (e.g., phenyl isothiocyanate, sodium isothiocyanate, ammonium isothiocyanate (e.g., tetrabutylammonium isothiocyanate, tetrabutylammonium isothiocyanate), diphenylphosphoryl isothiocyanate), acetyl chloride, cyanogen bromide, carboxypeptidase, azide, alkyne, DBCO, maleimide, succinimide, thiol-thiol disulfide bond, tetrazine, TCO, vinyl, methylcyclopropene, allyloyl, allyl, and the like. Additional examples of amino acid reactive groups are provided in U.S. Patent Publication No. 2020 / 0217853, U.S. Provisional Patent Application No. 63 / 384,007 filed on November 16, 2022, and U.S. Provisional Patent Application No. 63 / 481,932 filed on January 27, 2023, which are incorporated herein by reference in their entirety.

[0219] The joint can include any other useful part.For example, the joint can include a releasable or cleavable part, which can promote the removal of monomers from polymer analytes or their parts or from substrates. Such a releasable or cleavable part can include, for example, a disulfide bond, which can be releasable by contact with a reducing agent (for example, DTT, TCEP). In some examples, the joint can be coupled to a third polymerizable molecule via a releasable or cleavable part alternatively or additionally via click chemistry. Therefore, the coupling between polymerizable molecule and the joint can be reversible. Additionally or alternatively, the joint can include any number of spacers, for example, polymers (for example, PEG, PVA, polyacrylamide), aminocaproic acid, nucleic acids, alkyl chains, etc. Such spacers can increase the distance between any other part of the joint (for example, amino acid reactive groups and polymerizable molecule reactive groups). The joint can include or be coupled to a detectable part, for example, a fluorophore, a radioisotope, a mass tag, a nucleic acid molecule (also as a releasable or cleavable part) or other detectable parts. In some examples, the linker comprises a fluorophore, which allows visualization of the location of the linker using single molecule imaging. In another example, the monomer can be labeled with a first fluorophore and the linker can comprise a second fluorophore to enable visualization of the location of the linker and the monomer (e.g., using dual-channel imaging or FRET).

[0220] The use of a joint comprising two reactive groups can allow the joint to be coupled to a monomer and (ii) an intermediate connecting molecule (e.g., a third polymerizable molecule) or (iii) a capture portion of (i) a polymer analyte. In some cases, when using an intermediate connecting molecule, the joint can be pre-coupled to the intermediate connecting molecule. For example, a precursor joint can include a monomer binding group (e.g., PITC) and a click chemistry moiety (e.g., an azide), which can react with a polymerizable molecule (e.g., an oligonucleotide) comprising a complementary click chemistry moiety (e.g., an alkyne) to generate a joint that can be coupled to a monomer and a capture portion (e.g., another oligonucleotide). In some cases, a joint pre-coupled to an intermediate connecting molecule can be provided.

[0221] Figure 2 Example linkers that can be used to sequence polymeric analytes such as peptides are schematically shown. Figure 2Sub-figure A shows a bifunctional linker 203 (e.g., 1-(but-3-yn-1-yl)-4-phenylisothiocyanate) comprising an amino acid reactive moiety (e.g., PITC) and an alkyne click chemistry moiety that can react with a polymerizable molecule 201 (e.g., a linker nucleic acid molecule) comprising a complementary azide click chemistry moiety. The bifunctional linker can also comprise a spacer moiety, such as an alkyl chain of any length (ethyl is depicted), a polymer of any length (e.g., PEG), etc. The spacer moiety can be positioned between the amino acid reactive moiety and the click chemistry moiety. Figure 2 Sub-figure B shows the product of a click chemistry cycloaddition reaction between an azide and an alkyne group to generate a linker molecule comprising a polymerizable molecule and an amino acid reactive portion. The conjugation of the polymerizable molecule 201 with the bifunctional linker 203 can occur at any useful or convenient step. In an alternative example (not shown), the bifunctional linker 203 can comprise an azide group, such as 1-(2-azidoethyl)-4-phenylisothiocyanate, which can react with the polymerizable molecule 201 comprising the alkyne portion.

[0222] In some cases where polymer analytes include peptides comprising amino acid monomers, the coupling of a joint and an amino acid (e.g., NTAA or CTAA) changes amino acid whose chemical structure is changed. For example, if a joint comprising an isothiocyanate moiety is used, during or after contact with the isothiocyanate moiety, amino acid can be derived into a thiocarbamoyl group (e.g., under mild alkaline conditions). One or more further derivatizations can be carried out. For example, amino acid or amino acid derivatives (e.g., thiocarbamoyl derived amino acid) can be further derived into a thiazolone group (e.g., under acidic conditions), a thiohydantoin group or other chemical moieties. Similarly, a thiazolone group or a thiohydantoin group can be further derived into a thiocarbamoyl group.

[0223] Capture moiety: The capture moiety can be coupled to the monomer of the polymer analyte via any suitable mechanism. The coupling of the monomer to the capture moiety can include covalent interactions or non-covalent interactions. Coupling can occur through the interaction of binding pairs, for example, biotin and avidin (or streptavidin), cyclodextrin and hydrophobic small molecules (e.g., alkanes, benzene, polycyclic compounds), cucurbitane and adamantane or trimethylaminomethylferrocene, diphenylene cycloalkanes (e.g., calixarenes, cavitands, pillararens, tetralactams), etc.

[0224] In some cases, the capture portion includes other polymerizable molecules (e.g., nucleic acid molecules). In such cases, the monomer can first be coupled to the complementary polymerizable molecule (e.g., to generate a peptide-oligonucleotide conjugate) and, for example, directly or via a splint molecule via complementary base pairing, tethered to the capture portion. Alternatively, as described above, the monomer can be coupled to the capture portion via a joint. For example, the joint can include a monomer coupling group (e.g., a dithioester or thiocarbamoyl group, which can be coupled or reacted with the amino acids of the peptide) and a nucleic acid molecule. The capture portion can include other nucleic acid molecules, which can be coupled to the nucleic acid molecules of the joint via hybridization, connection, or both.

[0225] The capture moiety can comprise a nucleic acid molecule that can comprise any naturally occurring, non-naturally occurring, or engineered nucleotide base. For example, the nucleic acid molecule can comprise pseudo-complementary bases, bridged nucleic acids, heterologous nucleic acids, locked nucleic acids, peptide nucleic acids (PNA), gamma-PNA, morpholinos, and the like, as described elsewhere herein.

[0226] The capture portion may comprise one or more functional sequences, including but not limited to a priming sequence, a sequencing sequence, a sequencing read sequence, a chimeric end sequence, a transposase recognition sequence, a cleavage site (e.g., a restriction site), a UMI, a blocking group, a spacer sequence, a barcode sequence, or other functional sequences. In some cases, the capture portion includes a cleavable or releasable portion, such as a restriction enzyme recognition site, an abasic site, a Or uracil cleaved by uracil DNA glycosylase, a disulfide bond that can be released upon addition of a reducing agent.

[0227] In some cases, a capture moiety and a polymerizable molecule coupled to a substrate are provided. In one example, a substrate includes a first nucleic acid molecule, a second nucleic acid molecule, and a capture moiety, which may be a third nucleic acid molecule, coupled thereto. In some cases, a substrate can include identical nucleic acid molecules across the substrate; these identical nucleic acid molecules can serve as both a capture moiety and a polymerizable molecule, to which an additional polymerizable molecule (e.g., coupled to a binding agent) can be coupled. Alternatively or additionally, the capture moiety can be coupled to a polymer analyte (see, e.g., Figure 1F ).

[0228] The capture portion can include any useful part or functional group. The capture portion can have a monomer-capture group, a substrate-tethered group or a joint, or any other functional group or part, for example, for coupling or tethering to other molecules or for detection. In some examples, the capture portion includes a nucleic acid molecule, which includes a substrate-tethered group (for example, biotin, a click chemistry moiety such as an azide), which can be coupled to a substrate (for example, comprising streptavidin or complementary click chemistry). The capture portion can additionally include a binding sequence, and another nucleic acid molecule (for example, a connection nucleic acid molecule, a joint-monomer complex or a binding agent nucleic acid barcode molecule) can be coupled to the binding sequence, for example, via hybridization, connection or both. In some cases, the capture portion includes a single-stranded region in which a single-stranded oligonucleotide or a complementary oligonucleotide can hybridize. The complementary oligonucleotide can include a detectable label (for example, a fluorophore) that allows detection of the capture portion.

[0229] Cleavage: Cleavage of a monomer from a polymeric analyte can be achieved using any suitable mechanism, such as via application of a stimulus. The stimulus can be, for example, a chemical stimulus, a biological stimulus, a thermal stimulus (e.g., application of heat), a light stimulus, a physical or mechanical stimulus, or other types of stimuli or combinations of stimuli. In some cases, the stimulus can be a chemical stimulus, such as a change in pH, addition of a solubilizing agent, an initiator, a free radical generator, a reducing agent, or the like. In some cases, the stimulus can be a biological stimulus, such as an enzyme (e.g., Edmanase, a protease, an endonuclease, an artificial protease, such as an artificial peptidase) that can cleave or catalyze cleavage of a monomer from a polymeric analyte.

[0230] In some examples, the polymer analyte comprises a peptide and the monomer comprises an amino acid (e.g., NTAA, CTAA, or an internal amino acid). The method may include using a sequencing reagent comprising an amino acid reactive group (e.g., a dithioester or a thiocarbamoyl group) by coupling the amino acid reactive group of the sequencing reagent to an amino acid, and using a stimulus (e.g., a change in pH, temperature) to cut the amino acid from the peptide. In one example, the sequencing reagent can be coupled to NTAA under mild alkaline conditions to generate a sequencing reagent-monomer (amino acid) complex, and the Edman degradation reaction or a similar process can be used to achieve cutting NTAA from the peptide. In some examples, cutting can be performed using a change in pH, such as applying an acid (e.g., trifluoroacetic acid) or a base (e.g., sodium hydroxide). As described elsewhere herein, the sequencing reagent can include a portion or molecule (e.g., a nucleic acid molecule or a polymerizable molecule), which can also be coupled to a capture portion such that after cutting, the cut amino acid can be coupled to the capture portion.

[0231] In some cases, once can cut more than one monomer from polymer analyte.Cutting can comprise cutting 2 monomers, 3 monomers, 4 monomers, 5 monomers, 6 monomers, 7 monomers, 8 monomers, 9 monomers, 10 monomers or more monomers.For example, polymer analyte can comprise the peptide comprising multiple amino acid monomers, and can cut single amino acid, dipeptide, tripeptide, tetrapeptide or larger peptide with method described herein.In some cases, can cut at most about 10 monomers, at most about 9 monomers, at most about 8 monomers, at most about 7 monomers, at most about 6 monomers, at most about 5 monomers, at most about 4 monomers, at most about 3 monomers or monomer less in given cutting event.In some cases, can use can identify or cut more than single amino acid enzyme (for example, Edmanase, protease) to mediate the cutting of more than one monomer (for example, amino acid).

[0232] Cleavage of the monomer (or monomers) can be performed using a biological stimulus such as an enzyme. The enzyme can be any useful cleavage enzyme, for example, a protease such as Edmanase, cruzain, x protein (e.g., ClpS, ClpX), proteinase K, exopeptidase, aminopeptidase, diaminopeptidase, serine protease, cysteine ​​protease, threonine protease, aspartic protease, aspartic protease, glutamic protease, metalloprotease, asparagine peptide cleavage enzyme, pepsin, trypsin, pancreatin, Lys-C, Glu-C, Asp-N, chymotrypsin, carboxypeptidase (e.g., carboxypeptidase A, carboxypeptidase B, carboxypeptidase Y), SUMO protease, elastase, papain, endoprotease, protease, Bromelain, collagenase, hyaluronidase, thermolysin, ficin, keratinase, tryptase, fibroblast activation, enterokinase, chymotrypsinogen, chymase, clostripain, calpain, alpha-lytic protease, proline-specific endopeptidase, furin, thrombin, subtilisin, genenase, PCSK9, cathepsin, aminoacylproline dipeptidase, methionine aminopeptidase, cathepsin C, 1-cyclohexen-1-yl-boronic acid pinacol ester, pyroglutamate aminopeptidase, renin, kininogen, kallikrein, DPPIV / CD26, thimet oligopeptidase, prolyl oligopeptidase, leucine aminopeptidase, dipeptidyl peptidase, or other enzymes or proteases, or combinations or variants thereof (e.g., engineered mutants or variants). In some cases, the cleavage enzyme or ribozyme or deoxyribozyme can be configured or engineered to cleave a terminal monomer or multiple monomers; alternatively, the cleavage enzyme or ribozyme or deoxyribozyme can be configured or engineered to perform exo-site cleavage at a non-terminal position of the polymer analyte, for example, at an internal monomer within the polymer analyte, at positions n-1, n-2, n-3, n-4, n-5, n-6, n-7, n-8, n-9, n-10, etc. (where n is the number of monomers in the polymer analyte).

[0233] In the case of enzymatic cleavage, additional reagents can be provided to catalyze or induce cleavage. For example, metalloproteinases, aminopeptidases or exopeptidases can promote the cleavage of an amino acid or multiple amino acids in the presence of a catalyst (e.g., a metal or metal ion (e.g., cobalt)). Thus, a catalyst can be provided to promote the binding of the enzyme to the amino acid or the subsequent cleavage of the amino acid from the peptide. In some examples, cleavage can be mediated by an apoenzyme that is inactive in the absence of a metal catalyst of the cofactor, and cleavage can be controlled by the addition of a metal or metal ion.

[0234] Other examples of cleavage stimuli can include: optical stimulation (e.g., application of UV, X-rays, gamma rays, or other wavelengths of light), mechanical stimulation (e.g., sonication, high voltage), thermal stimulation (e.g., application of heat), or chemical stimulation. In some cases, the polymer analyte can contain or be altered to contain a cleavable or labile bond that can be cleaved upon application of an appropriate stimulus, such as a disulfide bond (e.g., cleavable upon application of a chemical stimulus such as a reducing agent), an ester bond (e.g., cleavable with a change in pH), a vicinal diol bond (e.g., cleavable with sodium periodate), a Diels-Alder bond (e.g., cleavable upon application of heat), a sulfone bond (e.g., cleavable via a base), a silyl ether bond (e.g., cleavable via an acid), a glycosidic bond (e.g., cleavable via an amylase), a peptide bond (e.g., cleavable via a protease), or a phosphodiester bond (e.g., cleavable via a nuclease (e.g., DNAse)).

[0235] Modification of Monomers: In some cases, one or more monomers of a polymeric analyte can be modified. The modification can be naturally occurring (e.g., post-translational modification) or non-naturally occurring, such as via labeling or tagging, for example, with an amino acid or amine reactant such as an isothiocyanate (e.g., PITC, NITC), 1-fluoro-2,-4-dinitrobenzene (DNFB), dansyl chloride, 4-sulfonyl-2-nitrofluorobenzene (SNFB), an acetylating agent, an acylating agent, an alkylating agent, a guanidating agent, a thioacetylating agent, a thioacylating agent, a thenoylating agent, or a derivative or combination thereof. Alternatively or additionally, one or more monomers can be modified to include any useful moiety, such as an adduct (e.g., a polymer such as PEG, a polymerizable molecule such as a nucleic acid molecule, a nanoparticle or nanotube, a peptide or protein), a lipid, a carbohydrate, a metabolite, a fluorophore, a hapten, a quencher, a tag (e.g., a fluorescent tag, a magnetic tag, a radioactive tag), a barcode, or other moiety. In some cases, the monomers of the polymer analyte can be modified to promote the recruitment of an enzyme to recognize or cut a terminal monomer (e.g., NTAA or CTAA of a peptide, a 5' or 3' nucleotide of a nucleic acid molecule, or the first or last monomer of a polymer) or a set of monomers. For example, the terminal amino acid of a peptide analyte can be modified with a sugar to recruit a lectin or a lectin-binding protease. In another example, one or more monomers of a polymer analyte can comprise or be coupled to a nucleic acid molecule having a first sequence that is complementary to a second sequence comprised by an oligonucleotide-binding protease. Hybridization of the first sequence with the second sequence can promote local recruitment of the protease to the monomer to be cut. In yet another example, the peptide analyte can be modified with PITC, which can allow recruitment and cutting by Edmanase. In some examples, the modification of the monomers of the polymer analyte can include an epitope tag that can promote the binding of a binding agent (e.g., after cutting a monomer from the polymer analyte). Examples of such epitope tags include fluorophores, nucleic acid molecules, peptides, haptens, polymers, chemical moieties, or other adduct molecules. Other examples of polymer analyte modifications are described elsewhere herein.

[0236] Polymer analytes can include one or more modified monomers. The modification of the monomer can be naturally occurring or synthetic. Synthetic modification can be carried out before, during or after cutting the monomer from the polymer analyte, and can be beneficial to maintaining the identity of the monomer. For example, during the standard Edman degradation reaction of cutting the terminal amino acid (monomer) from the peptide, some amino acid residues may be changed by the reaction conditions or become undetectable. In one example, the conditions of Edman degradation can cause the oxidation of cysteine ​​residues, the dehydration or destruction of serine or threonine in the form of phenylthiohydantoin (PTH), react with lysine residues and modify lysine residues or make some post-translational modifications undetectable. Therefore, modifying the peptide before analysis, for example protecting some amino acid residues or post-translational modifications, can be useful in more accurately identifying each amino acid residue. In one example of modifications that can be made prior to cleavage, the peptide or portion thereof can be alkylated, e.g., alkylating a cysteine ​​residue (e.g., using 4-vinylpyridine, iodoacetamide, iodoacetate, chloroacetate); acetylated, e.g., reacting a serine or threonine residue to form an ester (e.g., using acetyl chloride) or using acetic anhydride; oxidized, e.g., converting a cysteine ​​residue to cysteic acid; reduced (e.g., using a reducing agent such as dithiothreitol, β-mercaptoethanol, or TCEP); contacted with a protecting group, e.g., a phosphorylated residue can be protected (e.g., using β-elimination of the phosphate group, Michael addition with an optional sulfhydryl group, e.g., as described in Knight, et al. Nat. Biotechnology. 21, 1047-1054 (2003), which is incorporated herein by reference in its entirety), etc. The polymeric analyte or monomer can be modified with a protecting group or moiety, such as methyl, formyl, ethyl, acetyl, tert-butyl, anisyl, benzyl, trifluoroacetyl, N-hydroxysuccinimide, tert-butyloxycarbonyl (Boc), benzoyl, 4-methylbenzyl, thioformyl, thiocresol, benzyloxymethyl, 4-nitrophenyl, benzyloxycarbonyl, 2-nitrobenzoyl, 2-nitrobenzenesulfinyl, 4-toluenesulfonyl, pentafluorophenyl, diphenylmethyl, 2-chlorobenzyloxycarbonyl, 2,4,5-trichlorophenyl, 2-bromobenzyloxycarbonyl, 9-fluorenylmethyloxycarbonyl (FMOC), trityl, or 2,2,5,7,8-pentamethylchroman-6-sulfonyl. Polymer analytes or monomers can be treated with protecting agents such as carboxyethyl methanethiosulfonate (CEMTS), thiazolidine, mercaptophenylacetic acid, cyanobenzothiazole (e.g., for lipidation of N-terminal cysteine), acetamidomethyl, 2-methylsulfonylethyloxycarbonyl, etc. In some cases, lysine residues can be blocked using isothiocyanates (e.g., PITC) (e.g., the primary amine of the lysine residue can react), and optionally a single round of Edman degradation can be performed to generate a new N-terminal exposed end.

[0237] In some cases, the monomers of the polymer analyte can be modified to facilitate cleavage of the monomers from the polymer analyte. For example, the amino acid monomers of the peptide polymer analyte can be modified (e.g., acetylation of the amino acid) so that they are recognized by the enzyme and can facilitate cleavage of the acetylated amino acid by the acyl peptide hydrolase. Additional or alternative modifications to the monomers, such as those described herein, can also facilitate recognition by or interaction with the engineered cleavage enzyme.

[0238] In some cases, the monomer comprising naturally occurring modifications can be treated to remove or change naturally occurring modifications so that polymeric analytes or monomers are more suitable for processing operations disclosed herein. For example, acetylation, formylation, methylation and pyrrolidonecarboxylic acid post-translational modifications can be removed before sequencing. Acetylation modifications can be removed with acyl peptide hydrolases or acid treatment (for example, using 1N HCl). Methylation can be removed using aminopeptidase. Formylation modifications can be removed, for example, using acid treatment (for example, 0.6M HCl treatment). Pyrrolidonecarboxylic acid (PCA) can be removed with pyroglutamate aminopeptidase. Exemplary C-terminal modifications can include amidation and methylation, both of which can be removed using carboxypeptidase.

[0239] Binder: The binding agent can be contacted with the monomer (e.g., after cleavage and monomer-capture moiety coupling). The binding agent can be any useful molecule that can be coupled to a monomer or monomer-capture moiety complex. For example, the binding agent can be or comprise a protein or peptide (e.g., an antibody, antibody fragment, single-chain variant fragment (scFv), nanobody, anticalin, tRNA synthetase or tRNA-acyl synthetase, fibronectin domain), a simulated peptide, a peptide mimetic (e.g., a peptidomimetic, a β-peptide, a D-peptide peptide mimetic), a polysaccharide, a nucleic acid molecule (e.g., an aptamer), a somamer, a polymer, an inorganic compound, an organic compound, a small molecule, or a derivative thereof (e.g., an engineered variant), or a combination. In the case where the polymer analyte comprises a peptide, the binding agent can be capable of binding to a modified amino acid (e.g., an amino acid coupled to a sequencing reagent) or a portion thereof. The binding agent can comprise a recognition site that specifically recognizes an amino acid, a modified amino acid (e.g., a sequencing reagent-amino acid complex), or a derivatized (and optionally modified) amino acid. For example, the binding agent can be configured to recognize the portion of the modified amino acid or have binding specificity to the portion of the modified amino acid, such as a specific amino acid residue, a residue-sequencing reagent complex or a derived amino acid (e.g., a thiocarbamoyl-derived residue, a thiazolone-derived residue, a thiohydantoin-derived residue, etc.) or a portion of the modified amino acid. In some cases, the binding agent can be derived from or engineered from a naturally occurring enzyme or protein, such as an aminopeptidase, an exopeptidase, a metalloproteinase, an antibody, anticalin, an N-recognition protein, a Clp protease, an endoproteinase (e.g., trypsin) or a tRNA synthetase. In some examples, the binding agent can be a cleavage enzyme (e.g., trypsin, endoproteinase) that has been modified to remove peptidase activity. The binding agent can also recognize the terminal amino acid attached to the substrate; for example, after all monomers except the final monomer of the polymer analyte have been coupled to the capture portion or multiple capture portions and cut, the final monomer can remain coupled to the substrate. Therefore, the binding agent can recognize and bind to surface-coupled monomers.

[0240] The binding agent can contact and specifically bind to the cleaved monomer, sequencing reagent-monomer complex, monomer-sequencing reagent-capture moiety, or monomer-capture moiety complex (collectively referred to herein as "monomer analyte"). For example, the monomer analyte can fall into any size or size range that is smaller than the entire polymer analyte. The size of the monomer analyte complex can be about 0.1 nanometer (nm), about 0.5 nm, about 1 nm, about 5 nm, about 10 nm, about 20 nm, about 30 nm, about 40 nm, about 50 nm, about 60 nm, about 70 nm, about 80 nm, about 90 nm, about 100 nm, about 200 nm, about 300 nm, about 400 nm, about 500 nm, about 600 nm, about 700 nm, about 800 nm, about 900 nm, about 1 micrometer (μm), about 10 μm, about 100 μm, about 1 millimeter mm, or larger. The monomer analyte can have any molecular weight or molecular weight range. Monomeric analytes can be about 1 Dalton (Da), 10 Da, 100 Da, 500 Da, 1 kiloDalton (kDa), 10 kDa, 100 kDa, 1,000 kDa, 10,000 kDa, 100,000 kDa, or greater. The molecular weight or length of monomeric analytes can vary, for example, depending on the amino acid residues.

[0241] The binding agent may comprise or be coupled directly or indirectly to a polymerizable molecule. The polymerizable molecule may be the same type of molecule as the binding agent (e.g., both are peptides, both are nucleic acid molecules, etc.), or they may be different. In some cases, the binding agent comprises a peptide (e.g., an antibody or antibody fragment) and the polymerizable molecule comprises a nucleic acid molecule. The polymerizable molecule can be conjugated to the binding agent via chemical conjugation methods, for example, using linkers such as SMCC, (Ne-maleimidocaproyloxy) succinimide ester (EMCS), succinimidyl-4-(p-maleimidophenyl) butyrate (SMPB), succinimidyl-(N-maleimidopropionamido-ethylene glycol) ester (SMPEG), succinimidyl (NHS) esters, succinimidyl-4-formylbenzamide (S-4FB), succinimidyl-6-hydrazino-nicotinamide (S-HyNic), 4-phenyl-3H-1,2,4-triazolidine-3,5(4H)dione (PTAD) or other diazonium salts, 1-ethyl-3-3-dimethylaminopropylcarbodiimide hydrochloride (EDC), and the like. The synthesis of peptide-nucleic acid molecule conjugates can also be performed using solid phase synthesis, fragment conjugation (e.g., using a heterobifunctional cross-linker such as a cross-linker comprising an aliphatic chain and a maleimide group on one end and NHS on the other end), click chemistry (e.g., strain-promoted azide-alkyne cycloaddition, reverse electron demand Diels-Alder reaction), or a combination of methods or chemistries. In some cases, enzymatic methods can be used to conjugate the polymerizable molecule to the binding agent. For example, DNA-protein conjugates can be generated using truncated nucleases (e.g., Cas proteins such as Cas9), relaxases (e.g., VirD2), or other enzymes, ribozymes, or deoxyribozymes. In some cases, the polymerizable molecule can be conjugated to the binding agent using a SpyTag and SpyCatcher interaction, a biotin-avidin interaction, a SNAP-tag, or other interaction. Optional purification can be performed, for example, using ion exchange chromatography, HPLC, affinity chromatography, or other purification techniques.

[0242] The binding agent can be coupled to the polymerizable molecule via non-covalent interactions. For example, the binding agent can include an avidin or streptavidin tag to which a biotin-conjugated polymerizable molecule can bind. Alternatively, the binding agent can include a biotin tag to which an avidin or streptavidin-conjugated polymerizable molecule can bind.

[0243] The polymerizable molecule may contain identification information for the binding agent. For example, the polymerizable molecule may include a nucleic acid barcode molecule comprising a barcode sequence. The barcode sequence may encode the identity of the binding agent or binding partner. For example, monomers (e.g., amino acids) of a polymer analyte (e.g., a peptide comprising multiple amino acids) may be cleaved and coupled to a capture portion (e.g., on a substrate) and may be contacted with a binding agent (e.g., an antibody, an antibody fragment, a nanobody). The binding agent may specifically recognize an amino acid residue or derivative thereof (e.g., a form derived from PTH, PTC, ATZ) relative to other amino acid residues or derivatives thereof. The nucleic acid barcode molecule may contain information for identifying the binding agent, which may also identify a specific amino acid residue (or derivative) due to the binding agent's specificity for its target.

[0244] The polymerizable molecules of the binding agent can include additional multiplexing information. For example, the polymerizable molecules (e.g., nucleic acid molecules) can include sequences encoding cycles or other time information or spatial information. In one such example, an array of peptides and capture moieties can be provided on a substrate. The array can include multiple individually addressable units, wherein each individually addressable unit of the array (or a subset thereof) includes peptides to be analyzed and capture moieties. The binding agent and the polymerizable molecules contained therein or coupled thereto can include spatial information (e.g., spatial barcode sequences), which uniquely identify the individual addressable units and therefore identify the position of the array. The polymerizable molecules can additionally include time information (e.g., indicating a round or iterative cycle barcode in which the binding agent or polymerizable molecules are provided). Subsequently, sequencing of the polymerizable molecules can be used to reveal spatial information (e.g., the origin position of the peptide or amino acid in the array). In some cases, the polymerizable molecules can include a unique molecular identifier (UMI), which can be used to determine the amount of a given binding agent or monomer (e.g., amino acid) of a given peptide, substrate, array or sample.

[0245] Alternatively, the binding agent that recognizes the monomer-capture moiety complex may not contain or be coupled to a polymerizable molecule. In such cases, after the binding agent is bound to the monomer-capture moiety complex, another molecule (e.g., a second binding agent) containing a detectable label such as a fluorophore, a radioisotope, a mass tag, or an identification polymerizable molecule (e.g., a nucleic acid barcode molecule) can be brought into contact with and bound to the binding agent bound to the monomer-capture moiety complex. In some examples, the additional molecule includes an identification polymerizable molecule, and the identification polymerizable molecule can be coupled or transferred to another polymerizable molecule. In one non-limiting example, the binding agent includes a first antibody or antibody fragment that recognizes a monomer-capture moiety complex (e.g., a terminal amino acid-sequencing reagent-capture moiety complex) or a portion thereof (e.g., a terminal amino acid or a terminal amino acid-sequencing reagent complex); after the first antibody or antibody fragment is bound to the monomer-capture moiety complex or a portion thereof, a second antibody or antibody fragment that contains or is coupled to a polymerizable molecule (e.g., a nucleic acid barcode molecule) is coupled to the first antibody. The polymerizable molecule of the second antibody or antibody fragment may contain information about the second antibody or antibody fragment, the first antibody or antibody fragment, or other information. The transfer of the polymerizable molecule of the second antibody or antibody fragment to another polymerizable molecule or coupling thereto may be mediated by any suitable technique, for example, hybridization of nucleic acid molecules, optionally mediated by a splint molecule, click chemistry, or association of high-affinity molecules (e.g., streptavidin and biotin).

[0246] In some cases, the method may include contacting the monomer-capture moiety complex with a binding agent library. The binding agent library may include a variety of binding agents that are specific for different analytes. For example, the binding agent library may include multiple binding agents that recognize different amino acids or derivatives thereof (e.g., derived amino acids such as PTH, PTC, or ATZ forms), amino acid clusters (e.g., dipeptides, tripeptides, etc.), or amino acid combinations (e.g., amino acids with similar side chain groups). In one such example, a given binding agent can recognize and bind to more than one amino acid, optionally with different affinities or binding kinetics. A given binding agent can recognize and bind to a single amino acid, two different amino acids, three different amino acids, four different amino acids, etc. For example, a given binding agent can bind to amino acids with similar residues, such as amino acids with positively charged side chains (e.g., arginine, histidine, lysine), amino acids with negatively charged side chains (aspartic acid, glutamic acid), amino acids with polar uncharged side chains (e.g., serine, threonine, asparagine, glutamine), amino acids with hydrophobic side chains (e.g., alanine, valine, isoleucine, leucine, methionine, phenylalanine, tyrosine, tryptophan), or a combination thereof. In general, a library of binding agents can specifically recognize or bind to any number of different amino acids; for example, a library of binding agents can be configured to specifically bind to at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 11, at least 12, at least 13, at least 14, at least 15, at least 16, at least 17, at least 18, at least 19, or at least 20 different proteinogenic amino acids or derivatives thereof.

[0247] The binding agent library can contain any useful number of binding agents, each of which can have different binding specificities. For example, a first binding agent can recognize one amino acid, a second binding agent can recognize two amino acids, and a third binding agent can recognize three amino acids. In another example, a first binding agent can recognize one amino acid, a second binding agent can recognize different amino acids, and a third binding agent can recognize multiple amino acids. It will be understood that any number of binding agents can be used, and each binding agent can be specific for one or more amino acids. In summary, the binding agent library can bind to all 20 proteinogenic amino acids or derivatives thereof, or a subset of amino acids (e.g., 10 or more, 15 or more).

[0248] Inactivation of the binder can be performed before or during contact with the cleaved monomer. Inactivation can be achieved using blocking agents or solutions such as lactoproteins (e.g., lactoglobulin, lactalbumin, lactoferrin, casein, whey, immunoglobulins, insulin, growth factors, osteopontin), albumin (e.g., bovine serum albumin), Tween 20, commercially available blocking solutions, or combinations thereof. Alternatively or additionally, inactivation of the binder can be performed using polymers (e.g., polyethylene glycol), organic compounds (e.g., oils, lipids), sugars, nanoparticles, inorganic compounds, ions, and the like.

[0249] Coupling of polymerizable molecules: polymerizable molecules can be coupled to each other using any useful method. Such coupling can include covalent interactions or non-covalent interactions (e.g., ionic interactions, hydrophobic interactions, van der Waals forces, etc.). In some cases, the first polymerizable molecule and the second polymerizable molecule include nucleic acid molecules and can be coupled via hybridization, connection, or both. For example, the first polymerizable molecule can include a first sequence that is complementary to the second sequence of the second polymerizable molecule, and the coupling can occur via hybridization of the first sequence with the second sequence. Alternatively, the first sequence and the second sequence may not be complementary to each other, but may be complementary to the third sequence and the fourth sequence of a splint or a bridging oligonucleotide, respectively. Therefore, the coupling of the first polymerizable molecule to the second polymerizable molecule can be mediated by hybridization of the first sequence and the second sequence with the third sequence and the fourth sequence of a splint or a bridging oligonucleotide, respectively.

[0250] In some cases, nucleic acid reaction can be carried out as part of or in addition to the coupling of the first polymerizable molecule to the second polymerizable molecule. For example, the first sequence of the first polymerizable molecule can hybridize with the second sequence of the second polymerizable molecule, and a nucleic acid extension reaction (for example, using a polymerase) can be carried out. Such extension reactions can allow the coding information of one of the polymerizable molecules (for example, the first polymerizable molecule) to be transferred to another polymerizable molecule (for example, the second polymerizable molecule). In another example, the first sequence of the first polymerizable molecule can be connected to the second sequence of the second polymerizable molecule to provide a first polymerizable molecule covalently coupled to the second polymerizable molecule.

[0251] Polymerizable molecules can be chemically coupled covalently or non-covalently. In some cases, the first polymerizable molecule can be chemically connected to the second polymerizable molecule. For example, the first polymerizable molecule can include a first reactive moiety, and the second polymerizable molecule can include a second reactive moiety that can react with the first reactive moiety. The first reactive moiety can be contacted with the second reactive moiety, and subjected to conditions sufficient to connect the first reactive moiety to the second reactive moiety, for example, via click chemistry. In other cases, the first polymerizable molecule can be coupled to the second polymerizable molecule via non-covalent or indirect interaction (for example, biotin-streptavidin).

[0252] In some cases, the polymerizable molecules of a binding agent can be coupled to other polymerizable molecules. For example, a substrate can include a polymer analyte and a capture moiety coupled thereto, together with a plurality of other polymerizable molecules. After a monomer is coupled to the capture moiety and the monomer is cut from the polymer analyte, the monomer can be contacted with the same or different binding agent any number of times. The polymerizable molecules of a single binding agent can iteratively contact and be coupled to any number of other polymerizable molecules for repeated inquiries; for example, a polymerizable molecule of a binding agent can be coupled to a first other polymerizable molecule, as described herein, and then subsequently cut or removed (for example, via dehybridization), and contacted and coupled with a second other polymerizable molecule. Such methods can be beneficial for transferring several copies of the polymerizable molecules of a binding agent to a substrate.

[0253] Uncoupling monomers, binding agents: In some cases, after the first polymerizable molecule is coupled to the second polymerizable molecule, the monomer can be uncoupled from the monomer-capture moiety complex or substrate. Uncoupling can be performed chemically, mechanically or enzymatically. For example, in some cases, the monomer is coupled to the capture moiety via a connection nucleic acid molecule (e.g., a sequencing reagent comprising a monomer reactive group and a connection nucleic acid molecule). The connection nucleic acid molecule can include a cleavage site, such as a restriction site, and can be uncoupled by enzymatic cleavage at the cleavage site using, for example, a restriction endonuclease. Alternatively or additionally, the capture moiety or any polymerizable molecule can include a cleavage site that allows the monomer to be uncoupled from a portion of the capture moiety. In a non-limiting example, other examples of enzymatic cleavage include glycosylases (e.g., uracil glycosylases), restriction endonucleases, micrococcal nucleases, transposases, Cas proteins (e.g., Cas9), Argonaut endonucleases, etc. Beneficially, removing the monomer from the monomer-capture moiety complex can allow the capture moiety to be used for subsequent reactions or iterations, or prevent additional binding agents from binding to the capture moiety, which can help reduce erroneous or repeated coupling of the polymerizable molecules of the additional binding agents to the capture moiety. Alternatively or additionally, uncoupling can occur using stimulation (e.g., light stimulation (such as UV, γ, X-ray irradiation), thermal stimulation, chemical stimulation, etc.). In some cases, the sequencing reagent can contain a cleavable group, and applying an appropriate stimulus can result in cleavage of the sequencing reagent, as described elsewhere herein.

[0254] Alternatively or additionally, the monomer can be altered so that it cannot be detected by the binding agent, for example, to prevent the binding agent from binding to the cleaved monomer in subsequent iterations or cycles of cleavage, coupling to a capture moiety, and contact with another binding agent. For example, the monomer can be contacted with a blocking agent or derivatized so that the binding agent no longer recognizes the derivatized form. Such blocking strategies can be useful in eliminating the need to detect or transfer information from the polymerizable molecule to which the binding agent is coupled, thereby removing the cleaved monomer. Other strategies for inhibiting binding of binding agents to cleaved monomers are described elsewhere herein.

[0255] Similarly, in some cases, the binding agent can be removed from the monomer-capture moiety complex by any useful or convenient operation, such as after the polymerizable molecule is coupled. Removal of the binding agent can be performed using chemical or enzymatic methods, such as using chemical denaturants, detergents, acidic or alkaline conditions, heat or proteases. Alternatively or additionally, if the polymerizable molecule is coupled to the binding agent, the polymerizable molecule can be removed from the binding agent, such as via cleavage or restriction sites and the use of a cleavage enzyme (e.g., UDG, restriction enzyme), chemical cleavage, photolysis or other methods. In some cases, the polymerizable molecule is coupled to the binding agent via non-covalent interactions, such as desthiobiotin-avidin; therefore, uncoupling of the polymerizable molecule from the binding agent can be achieved by competitively replacing desthiobiotin using a competitor, such as biotin of higher affinity.

[0256] Identification of polymerizable molecules: Polymerizable molecules can be sequenced to determine the identity of individual monomers (e.g., amino acids). For example, monomers can be cleaved from the capture moiety (e.g., Figure 1A F) or after any number of iterations of workflow 100, the polymerizable molecules containing the binder information and, therefore, the identity of the monomer can be removed from the substrate and prepared for sequencing (e.g., DNA sequencing, NGS). Removal of the polymerizable molecules can be achieved using any useful method, such as chemical or enzymatic cleavage. In some cases, for example, prior to removing the polymerizable molecules containing the monomer information, any excess or uncoupled polymerizable molecules or capture moieties can be removed. For example, again referring to Figure 1A Subgraph B and Figure 1A Panel E, the coupling event can generate a double-stranded or partially double-stranded molecule; thus, a single-stranded polymerizable molecule or capture portion that does not contain monomers or additional polymerizable molecules (e.g., from a binding agent) can be digested using an enzyme such as a Type II restriction endonuclease, S1 endonuclease.

[0257] Alternatively or additionally, the polymerizable molecules can be amplified (e.g., using nucleic acid amplification methods such as polymerase chain reaction (PCR), isothermal amplification, ligation-mediated amplification, transcription-based amplification, etc.) to generate amplicons for sequencing. Amplification can be performed, for example, using a capture moiety or polymerizable molecule as a primer binding site. Any number of useful preparative operations can be performed, such as purification or enrichment, cleanup, nucleic acid reactions (e.g., ligation, extension, amplification, marker cleavage, restriction enzyme cleavage), fragmentation, barcoding, addition of adapters, enzymatic treatment, etc. In some cases, the polymerizable molecules or substrates comprising the polymerizable molecules can be filtered based on any useful feature or property. Filtering based on features or properties can achieve higher accuracy or lower noise by removing low-quality molecules or enriching high-quality polymerizable molecules prior to sequencing. For example, polymerizable molecules or substrates containing polymerizable molecules (e.g., beads or particles) can be filtered by size or length, number, presence of specific sequences (e.g., primer sequences, sequences of interest), GC content, polarity, polarization, birefringence, fluorescence (or other optical properties), anisotropy, charge, secondary structure (e.g., hairpins), or other useful indicators, features, or properties, or combinations thereof. Such filtering or enrichment can be performed using any suitable method, such as affinity or hybridization methods (e.g., bead-based affinity sequence or hybridization assays that can enrich for specific sequences), chromatography, size-based filtration, electrophoresis, electrofocusing, optoelectronics, digital fluidics, magnetic-activated sorting, fluorescence-activated sorting, flow cytometry, or other suitable techniques.

[0258] Sequencing can be performed using commercially available nanopore systems (e.g., Oxford Nanopore Technologies, Genia Technologies, NobleGen, or Quantum Biosystem) or other sequencing and next-generation sequencing systems (e.g., Illumina, BGI, Qiagen, ThermoFisher, PacBio, and Roche), including formats such as parallel bead arrays, sequencing by synthesis, sequencing by ligation (e.g., SOLiD), capillary electrophoresis, electronic microchips, "biochips," microarrays, parallel microchips, single molecule arrays, and Sanger sequencing, as described elsewhere herein.

[0259] Sequencing can output the identity of the polymerizable molecule or the sequence of polymerizable molecules coupled together. For example, referring again to Figure 1AAfter one or more iterations of the workflow 100, the polymerizable molecule 107 can comprise the nucleic acid sequence of the polymerizable molecule 117 of the binder (or its complement), or a stack of polymerizable molecules obtained from multiple rounds of binding of the binder to its target monomer or monomer-capture moiety complex (or its complement). Thus, sequencing the polymerizable molecule 107 can generate sequencing reads that identify the nucleic acid sequence of the polymerizable molecule 117 of the binder and the information encoded therein, such as the cycle number and identity of the binder or monomer (e.g., one of the 20 proteinogenic amino acids). In the case where the polymerizable molecule 117 of the binder includes a nucleic acid molecule that encodes additional information (e.g., including a barcode sequence, a UMI, cycle information, spatial information, etc.), multiple types of information can be revealed from the nucleic acid sequencing of the polymerizable molecule 107.

[0260] Figure 3 Schematic diagram of barcoded DNA sequences (e.g., Figure 1A Multiple iterations of the workflow 100 generate a stack of polymerizable molecules (a stack of molecules) that are input into a sequencing instrument and output a representation of a peptide or protein sequence. The barcoded DNA sequence comprises a stack of barcode sequences obtained by individual binders bound to their target monomer or monomer-capture moiety complex; each barcode sequence encodes the identity of a monomer (e.g., an amino acid).

[0261] Sequencing reads can be assembled using a de novo method to identify peptides or proteins. For example, a common barcode sequence can be used to mark the fragmented peptides produced from a common parent protein, as described elsewhere herein. Therefore, it is possible to assemble inferred peptide reads based on a common barcode sequence, amino acid identity, and, if applicable, the number of cycles. Error reads can be identified by probabilistic modeling of read accuracy, thereby generating reconstructed fragment peptide sequences (contigs) with possible gaps in lost or unidentified rounds / amino acids. The alternative option for de novo read reconstruction can adopt an end-to-end, unsupervised machine learning-based reconstruction of peptide reads. This option can adopt a machine learning algorithm, which refers to a model based on deep learning, which uses the NGS sequencing reads associated with the parent protein / peptide barcode as its input, and outputs the possible reconstruction of peptide reads (contigs). The training of the model can be performed using known protein / peptide standards with protein sequencing operations. De novo reconstruction can output reconstructed fragment peptide sequences (contigs), with probability assigned to each amino acid and assembled peptide sequence. In some cases, fixed length nucleotide string (k-mer) or De Brujin method can be used for peptide sequence reconstruction.For example, the read segment produced from each polymerizable molecule can be decomposed into a shorter fixed length nucleotide string sequence.The fixed length nucleotide string sequence from the read segment pool can be assembled into a longer contig sequence.De Brujin diagram can be generated, for example, to represent splice variants, post-translational modifications or other protein variants.Isoforms (isoforms) can be assembled, and expression levels can be determined using Bayesian methods.The assembled isoforms of proteins can be evaluated and error corrected, for example, by comparing with the standard protein incorporated into the sample, and evaluating the missing segments, incorrect or redundant assembly, uniform coverage, etc. of the sequence.

[0262] Alternatively, the identity of the polymerizable molecule can be obtained without using a sequencing method. For example, a probe can be used to couple a specific region of the polymerizable molecule. The probe can include a nucleic acid probe with a probe sequence that can be used to specifically detect a monomer type. In one such example, the polymer analyte includes a peptide, and a single amino acid (monomer unit) can be coupled to a capture portion and cleaved from the peptide. The monomer-capture portion complex can be contacted with a binding agent (e.g., an antibody, nanobody, scFv) containing a nucleic acid barcode molecule (polymerizable molecule) that identifies the binding agent. The binding agent may be specific for an amino acid (e.g., one of the 20 proteinogenic amino acids), and thus the nucleic acid barcode molecule encodes a specific amino acid. Therefore, a nucleic acid probe having a complementary sequence to the nucleic acid barcode molecule of the binding agent can be used to identify the presence of the binding agent (e.g., via in situ hybridization). In some cases, the probe can include a detectable label or portion, such as a fluorophore, a radioisotope, a mass tag, etc. For example, a hybridization-based assay such as SeqFISH or Nanostring can be performed to detect or assay a specific region of the polymerizable molecule to determine its identity. In other examples, amplification-based methods can be used to determine the presence and identity of polymerizable molecules. For example, PCR or nested PCR methods can be used to selectively detect specific sequences of polymerizable molecules.

[0263] Alternatively or additionally, the binding agent may comprise a detectable label or portion. For example, the binding agent may include a fluorophore, a radioisotope, a mass label, a chromogenic enzyme (e.g., horseradish peroxidase), etc., which may be detected using appropriate imaging techniques. Different binding agents (e.g., binding agents that recognize different monomers or amino acids) may be labeled with different markers (e.g., different fluorophores), which may be used to identify the presence of a monomer or amino acid. In some examples, single molecule imaging (e.g., total internal reflection, confocal, wide field of view, or super-resolution microscopy (e.g., PALM, STORM, STED)) may be used to detect binding agents labeled with fluorophores.

[0264] In some cases, a substrate comprising a polymerizable molecule (e.g., Figure 1AThe method of the present invention provides a method for sequencing a plurality of beads on an array. The method comprises the steps of: (a) performing one or more iterations of a workflow of a plurality of peptides on an array; (b) performing one or more iterations of a workflow of a plurality of peptides on an array; (c) performing one or more iterations of a workflow of a plurality of peptides on an array; (d) performing one or more iterations of a workflow of a plurality of peptides on an array; (e.g., ...

[0265] Fingerprint analysis: The methods described herein can be used for complete de novo protein or peptide sequencing (e.g., identifying each amino acid in a peptide) or for protein fingerprint analysis (e.g., identifying only a subset of the amino acid types in a peptide and inferring the identity of the peptide using a reference database). For fingerprint analysis, a subset of amino acids can be identified, for example, using the methods described herein, without the need for specific binders for all 20 protein amino acids. For example, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, or 19 different binders with single or multiple amino acid specificity can be sufficient to determine the identity of a protein or peptide. For a known proteome database, a reference-based reconstruction can be performed by simulating the NGS reads generated from the possible peptide sets in the workflow. For each possible peptide, the simulation can produce an NGS read that mimics the output of the protein sequence system. Next, the real (experimental) NGS reads from the run can be matched to the simulated reads of the candidate peptides from the database based on likelihood. This results in reconstructed, fragmented peptide sequences (contigs) with probabilities assigned to the assembled peptide sequence.

[0266] High-throughput sequencing / parallelization: The methods described herein can be performed in a parallelized high-throughput format. Such parallelization can be achieved by coupling a substrate comprising multiple polymer analytes thereto and iteratively performing operations across the substrate (e.g., coupling a monomer to a capture portion, cutting, contacting with a binding agent or binding agent library, coupling a polymerizable molecule, optionally cutting the monomer from the capture portion). In some cases, a binding agent library can be used to identify different monomer types (e.g., different amino acids or derived amino acids of a peptide analyte), such that different polymer analytes (e.g., different peptides) can be processed on a single substrate.

[0267] Binding agent libraries can be used to identify different monomer types to facilitate high-throughput readout. As described herein, binding agent libraries can include binding agents that can identify a single monomer (e.g., a single cleaved amino acid) or multiple monomers (e.g., multiple cleaved amino acids). In some cases, binding agents with different levels of specificity can be used sequentially or in order, which can help make the binding agent with less specificity more specific based only on the sequence provided. For example, a first binding agent can be able to specifically bind to a first monomer analyte and a second binding agent can be able to bind to both the first monomer analyte and the second monomer analyte. A first binding agent can be provided and contacted with the first monomer analyte and the second monomer analyte. Since the first binding agent has specificity for the first monomer analyte, the first binding agent will uniquely bind to the first monomer analyte. A second binding agent can subsequently be provided; however, since the first monomer analyte is bound to the first binding agent, the first monomer analyte may not be able to approach (e.g., spatially blocked) the second binding agent. Therefore, the second binding agent can only bind to the second monomer analyte. Thus, identification of the first and second binding agents (eg, by detecting a label / tag or by sequencing a polymerizable molecule coupled to the binding agent and optionally transferred to a substrate) can allow identification of the first and second monomeric analytes.

[0268] In some cases, it may be useful to barcode the polymer analyte prior to processing. A barcode sequence can be attached to the polymer analyte at a single location (e.g., at the end), multiple locations, near the polymer analyte (e.g., on a substrate), etc., as described elsewhere herein. For example, a peptide can be labeled with a nucleic acid barcode molecule at the N-terminus, C-terminus, or internal amino acid. The nucleic acid barcode sequence can contain information or be unique to a partition or compartment, sample, peptide, etc., such that each unique barcode sequence can be traced back (e.g., after nucleic acid sequencing or other detection methods) to the origin partition or compartment, sample, peptide, etc.

[0269] Alternatively or additionally, capture moiety or polymerizable molecule can include barcode sequence.Barcode sequence can be specific to specific partition, sample or spatial position.For example, substrate can include multiple units formed individually or by introducing chaotropic agent (for example, guanidine, formamide, urea).The polymerizable unit or capture moiety of each individual addressable unit can include a unique barcode (for example, spatial barcode) that is specific to the individual addressable unit. Polymer analyte can be coupled to substrate so that each individual addressable unit includes an average of no more than one polymer analyte.For example, this type of distribution of polymer analyte can use limiting dilution method (for example, diluting polymer analyte to reduce the number of polymer analytes that may be attached to a given individual addressable unit) or by introducing chaotropic agent (for example, guanidine, formamide, urea) to obtain.Polymer analyte can be distributed across individual addressable units according to Poisson distribution. Thus, for a given substrate, approximately 6%, 10%, 18%, 20%, 30%, 36%, 40%, or 50% of the individually addressable units may contain one or fewer polymeric analytes.

[0270] Modification of polymeric analytes and polymerizable molecules: The present disclosure also provides methods for modifying polymeric analytes or monomers of polymeric analytes (e.g., amino acids of a peptide) and polymerizable molecules described herein. Such modifications can be useful, for example, to render the monomer more resistant to certain reaction conditions (e.g., Edman degradation), to increase or decrease the binding affinity of a binding agent for the modified monomer, to aid in docking or interfacing the modified monomer with an enzyme (e.g., a protease, a cleavage enzyme, or an enzyme analog, such as a ribozyme or deoxyribozyme, a binding agent), or for other purposes.

[0271] Polymer analytes such as peptides can be modified to make the peptide or constituent amino acids more resistant to the reaction conditions for cleaving amino acids from the peptide. For example, the peptide can be alkylated, for example, using 4-vinylpyridine, iodoacetamide, which can be used to prevent oxidation of cysteine ​​residues. The peptide can be acetylated, for example, O-acetylated to form esters such as acetyl chloride, which can be used to prevent dehydration, racemization or destruction of the derivatization (for example, PTH form) of serine or threonine. The peptide can be subjected to β-elimination of phosphoric acid, followed by Michael addition of sulfhydryl groups (for example, as described in Knight et al. 2003.2003.Nature Biotechnology 21, 1047-1054, which is incorporated herein by reference) to detect phosphorylation events. The peptide can be contacted with phenyl isothiocyanate, acetic anhydride or other amine reactive groups to protect lysine residues. Additional examples of peptide processing for Edman degradation can be found in Tarr, Methods of Protein Microcharacterization. pp. 155-194 (incorporated herein by reference).

[0272] The polymer analyte or monomer can be modified to affect the interaction of the binding agent with the polymer analyte or monomer, for example, by derivatizing the cleaved monomer, adding a chemical group to the cleaved monomer, or performing other chemical treatments (e.g., adding or removing a group) from the cleaved monomer. Blocking of the binding agent can be achieved by attaching a blocking agent (e.g., a chemical group or adduct) to the monomer-capture moiety complex; for example, conjugation of a synthetic polymer (e.g., PEG), a nucleic acid molecule, a fluorophore, a quencher, a nanotube, a nanoparticle, a small molecule, a polypeptide or protein, a fatty acid chain, or other large sterically hindered molecule. The blocking agent can be attached to the monomer using chemical methods (e.g., reacting with an amino acid, for example, via a photoreaction) or enzymatic means (e.g., using a methyltransferase, a tRNA synthetase, an acetyltransferase, etc.).

[0273] Additional examples of modifications of monomers, cleaved monomers, and binding agents can be found in International Patent Publication No. WO2023196642, which is herein incorporated by reference in its entirety.

[0274] Order of Operations: It will be understood that the operations set forth in the methods described herein can be performed in any useful or convenient order, and in some cases, some operations can be optional. For example, in some cases, the coupling of a monomer to a capture moiety can occur before, during, or after the monomer is cut from the polymer analyte. Similarly, the coupling of a sequencing reagent (e.g., with or without a connecting nucleic acid molecule) to a capture moiety can occur before, during, or after the sequencing reagent is coupled to the monomer. In another example, a substrate can be provided with a cut monomer coupled thereto so that cutting of the monomer from the polymer analyte is avoided. In yet another example, where a sequencing reagent is used to couple to a monomer (e.g., an amino acid) and a capture moiety or substrate, the sequencing reagent can comprise a monomer-coupling group and subsequently react with a substrate-binding group (e.g., an oligonucleotide); alternatively, the sequencing reagent can be provided with a substrate-binding group as part of the sequencing reagent (e.g., pre-conjugated to the substrate-binding group).

[0275] Additional manipulations can be performed at any useful or convenient step, e.g., before providing the polymer analyte (e.g., peptide) or after one or more processing operations (e.g., after sequencing reagents, coupling of polymerizable molecules, contact with binding agents, etc.). For example, purification or enrichment or purification of a population of polymerizable molecules (e.g., in Figure 1B Course 110 or Figure 1F Such enrichment or purification can be performed using any useful technique, such as bead-based enrichment, immunoprecipitation, chromatography, electrophoresis, DNA purification, and the like. In one such example, nucleic acid molecules can be purified using beads containing or coupled to a complementary sequence to the nucleic acid molecule, and optionally, the nucleic acid molecule can be eluted after capture. Similarly, proteins can be purified using beads containing antibodies that recognize a protein or a portion of a protein.

[0276] Substrate conjugation

[0277] The present disclosure provides a method for coupling a molecule (e.g., a biomolecule such as a nucleic acid molecule, a peptide, a lipid, a carbohydrate, etc.) to a substrate. The substrate can be functionalized to allow the molecule to be covalently or non-covalently coupled to the substrate. The substrate can include any useful functional moiety, for example, a reactive moiety. In a non-limiting example, the reactive moiety can include a click chemistry moiety, such as an azide, an alkyne, a nitrone, an alkene (e.g., a strained alkene), a tetrazine, a methyl tetrazine, a triazole, a tetrazole, a phosphite, a phosphine, etc. The click chemistry moiety can be reactive in a copper-catalyzed Huisgen cycloaddition or a 1,3- dipolar cycloaddition between an azide and a terminal alkyne, a Diels-Alder reaction (e.g., a cycloaddition between a diene and a dienophile), or a nucleophilic substitution reaction in which one of the active species is an epoxy or aziridine. The molecule to be coupled to the substrate can comprise a click chemistry moiety that is complementary to the click chemistry moiety of the substrate; for example, the substrate can comprise an alkyne moiety and the molecule to be coupled can comprise an azide moiety that can react with the alkyne moiety of the substrate to produce a covalent linkage. In one such example, the substrate can comprise a dibenzocyclooctyne (DBCO) moiety, and an azide-containing molecule (e.g., azide-DNA, azide-polymer, azide-peptide) can react with the dibenzocyclooctyne moiety and be conjugated.

[0278] The reactive moiety can include a photoreactive moiety that can be activated when exposed to a light stimulus (e.g., light such as UV or visible light). Examples of photoreactive moieties include aryl(phenyl)azides (e.g., phenylazide, o-hydroxyphenylazide, m-hydroxyphenylazide, tetrafluorophenylazide, o-nitrophenylazide, m-nitrophenylazide), diaziridones, azido-methyl-coumarin, benzophenones, anthraquinones, diazo compounds, diaziridones, psoralens, and analogs or derivatives thereof.

[0279] The reactive moiety may comprise a carboxyl-reactive crosslinker group such as diazomethane, diazoacetyl, carbonyldiimidazole, carbodiimide (e.g., 1-ethyl-3-(3-dimethylaminopropyl)carbodiimide hydrochloride (EDC), dicyclohexylcarbodiimide (DCC)), or an amine-reactive group (e.g., N-hydroxysulfosuccinimide (NHS), sulfo-NHS, or NHS-ester). The reactive group may comprise a crosslinker that may comprise an NHS group, an EDC group, a maleimide, a thiol, a cystamine, an aldehyde, a succinimide group, an epoxy compound, or an acrylate. Examples of cross-linking agents include, for example, NHS (N-hydroxysuccinimide); sulfo-NHS (N-hydroxysulfosuccinimide); EDC (1-ethyl-3-[3-dimethylaminopropyl]); carbodiimide hydrochloride; SMCC (succinimidyl 4-(N-maleimidomethyl)cyclohexane-1-carboxylate); DSS (disuccinimidyl suberate); DSG (disuccinimidyl glutarate); DFDNB (1,5-difluoro-2,4-dinitrobenzene); BS3 (bis(sulfosuccinimidyl) suberate); TSAT (tris-(succinimidyl)aminotriacetate); BS(PEG)5 ( PEGylated bis(sulfosuccinimidyl) suberate); BS(PEG)9 (PEGylated bis(sulfosuccinimidyl) suberate); DSP (disulfobis(succinimidyl propionate)); DTSSP (3,3'-disulfobis(sulfosuccinimidyl propionate)); DST (disuccinimidyl tartrate); BSOCOES (bis(2-(succinimidyloxycarbonyloxy)ethyl)sulfone); EGS (ethylene glycol bis(succinimidyl succinate)); DMA (dimethyl adipimide); DMP (dimethyl pimelimidate); DMS (dimethyl suberimidate); DTBP (Wang and Richard's Reagent); BM(PEG)2(1,8-bismaleimido-diethylene glycol); BM(PEG)3(1,11-bismaleimido-triethylene glycol); BMB(1,4-bismaleimidobutane); DTME(disulfide bismaleimidoethane); BMH(bismaleimidohexane); BMOE(bismaleimidoethane); TMEA(tris(2-maleimidoethyl)amine); SPDP(3-(2-pyridyldisulfide) ) propionate); SMCC (trans-4-(maleimidomethyl) cyclohexane-1-carboxylic acid succinimidyl ester); SIA (succinimidyl iodoacetate); SBAP (succinimidyl 3-(bromoacetamido) propionate); STAB (succinimidyl (4-iodoacetyl) aminobenzoate); sulfo-SIAB (sulfosuccinimidyl (4-iodoacetyl) aminobenzoate); AMAS (N-α-maleimidoacetyl-oxysuccinimide ester);BMPS (N-β-maleimidopropyl-oxysuccinimide ester); GMBS (N-γ-maleimidobutyryl-oxysuccinimide ester); Sulfo-GMBS (N-γ-maleimidobutyryl-oxysulfosuccinimide ester); MBS (m-maleimidobenzoyl-N-hydroxysuccinimide ester); Sulfo-MBS (m-maleimidobenzoyl-N-hydroxysulfosuccinimide ester); SMCC (4-(N-maleimidomethyl)cyclohexane-1-carboxylic acid succinimide ester); Sulfo-SMCC (4-(N-maleimidomethyl)cyclohexane-1-carboxylic acid succinimide ester); EMCS (N-ε-maleimidocaproyl-oxysuccinimide ester) imide ester); sulfo-EMCS (N-ε-maleimidocaproyl-oxysulfosuccinimide ester); SMPB (4-(p-maleimidophenyl)butyric acid succinimide ester); sulfo-SMPB (4-(N-maleimidophenyl)butyric acid succinimide ester); SMPH (6-((β-maleimidopropionylamino)hexanoic acid succinimide ester)); LC-SMCC (succinimidyl 4-(N-maleimidomethyl)cyclohexane-1-carboxy-(6-aminohexanoate)); sulfo-KMUS (N-κ-maleimidoundecanoyl-oxysulfosuccinimide ester); SPDP (3-(2-pyridyldithio)propionic acid succinimide ester); LC-SPD P (6-(3(2-pyridyldithio)propionamido)hexanoic acid succinimidyl ester); LC-SPDP (6-(3(2-pyridyldithio)propionamido)hexanoic acid succinimidyl ester); Sulfo-LC-SPDP (6-(3'-(2-pyridyldithio)propionamido)hexanoic acid succinimidyl ester); SMPT (4-succinimidyloxycarbonyl-α-methyl-α(2-pyridyldithio)toluene); PEG4-SPDP (PEGylated long-chain SPDP crosslinker); PEG12-SPDP (PEGylated long-chain SPDP crosslinker); SM(PEG)2 (PEGylated SMCC crosslinker); SM(PEG)4 (PEGylated SMCC crosslinker); SM (PEG)6 (PEGylated long-chain SMCC cross-linker); SM(PEG)8 (PEGylated long-chain SMCC cross-linker); SM(PEG)12 (PEGylated long-chain SMCC cross-linker); SM(PEG)24 (PEGylated long-chain SMCC cross-linker); BMPH (N-β-maleimidopropionic acid hydrazide); EMCH (N-ε-maleimidocaproic acid hydrazide); MPBH (4-(4-N-maleimidophenyl)butyric acid hydrazide); KMUH (N-κ-maleimidoundecanoic acid hydrazide); PDPH (3-(2-pyridyldisulfide)propionohydrazide); ATFB-SE (4-azido-2,3,5,6-tetrafluorobenzoic acid, succinimidyl ester);ANB-NOS (N-5-azido-2-nitrobenzoyloxysuccinimide); SDA (NHS-diazopropene) (succinimidyl 4,4'-azidopentanoate); LC-SDA (NHS-LC-diazopropene) (succinimidyl 6-(4,4'-azidopentanoylamino)hexanoate); SDAD (NHS-SS-diazopropene) (succinimidyl 2-((4,4'-azidopentanoylamino)ethyl)-1,3'-dithiopropionate); Sulfo-SDA (Sulfo-NHS-diazopropene) (succinimidyl 4,4'-azidopentanoate); Sulfo-LC-SDA (Sulfo-NHS-LC-diazopropene) cyclopropene) (sulfosuccinimidyl 6-(4,4'-azidopentanoylamino)hexanoate); sulfo-SDAD (sulfo-NHS-SS-diazacyclopropene) (sulfosuccinimidyl 2-((4,4'-azidopentanoylamino)ethyl)-1,3'-dithiopropionate); SPB (succinimidyl-[4-(psoralen-8-yloxy)]-butyrate); sulfo-SANPAH (sulfosuccinimidyl 6-(4'-azido-2'-nitrophenylamino)hexanoate); DCC (dicyclohexylcarbodiimide); EDC (1-ethyl-3-(3-dimethylaminopropyl)carbodiimide hydrochloride); glutaraldehyde; formaldehyde; and combinations or derivatives thereof.

[0280] A linker can also be used to attach a molecule to a substrate. The linker can have any useful number of functional groups or reactive groups and can be monofunctional (having one functional group), difunctional, trifunctional, tetrafunctional or contain a greater number of functional groups. In some cases, a heterobifunctional linker can be used to attach a molecule (e.g., a nucleic acid molecule, a peptide or a polymer) to a substrate. A heterobifunctional linker can contain any useful functional group, as described herein.Non-limiting examples of heterobifunctional linkers include: p-azidobenzohydrazide (ABH), N-5-azido-2-nitrobenzoyloxysuccinimide (ANB-NOS), N-[4-(p-azidosalicylamido)butyl]-3'-(2'-pyridyldithio)propionamide (APDP), p-azidophenylglyoxal monohydrate (APG), bis[B-(4-azidosalicylamido)ethyl]disulfide (BASED), bis[2-(succinimidyloxycarbonyloxy)ethyl]sulfone (BSOCOES), BMPS, 1,4-bis[3'-(2'-pyridyldithio)propionamido]butane (DPDPB), disulfide bis(succinimidyloxycarbonyloxy)ethyl]sulfone (BSOCOES), bis[B-(4-azidosalicylamido)ethyl]sulfone ...bis[B-(4-azidosal bis(succinimidyl) propionate (DSP), disuccinimidyl suberate (DSS), disuccinimidyl tartrate (DST), 3,3'-disulfide bis(sulfosuccinimidyl) propionate (DTSSP), EDC, ethylene glycol bis(succinimidyl) succinate (EGS), N-(E-maleimidocaproyl) hydrazide (EMCH), N-(E-maleimidocaproyloxy)-succinimide ester (EMCS), N-maleimidobutyryloxysuccinimide ester (GMBS), hydroxylamine-HCl, MAL-PEG-SCM, m-maleimidobenzoyl-N-hydroxysuccinimide ester (MBS), N-hydroxysuccinimide-4- Azidosalicylic acid (NHS-ASA), PDPH, N-succinimidyl bromoacetate (SBA), SIA, sulfo-SIA, succinimidyl 4-(N-maleimidomethyl)cyclohexane-1-carboxylate (SMCC), succinimidyl 4-(p-maleimidophenyl)butyrate (SMPB), succinimidyl-6-[β-maleimidopropionylamino]hexanoate (SMPH), N-succinimidyl 3-[2-pyridyldithio]-propionate (SPDP), sulfo-LC-SPDP, N-(p-maleimidophenyl isocyanate (PMPI), N-succinimidyl (4-iodoacetyl) aminobenzoate (SIAB) ), sulfo-MBS, sulfo-SANPAH, sulfo-SMCC, sulfo-DST, sulfo-EMCS, sulfo-GMBS, N-hydroxysulfosuccinimidyl-4-azidobenzoate (sulfo-HSAB), (4-azidophenyl)-1,3-dithiopropionic acid sulfosuccinimidyl ester (sulfo-SADP), 2-(m-azido-o-nitrobenzamido)-ethyl-1,3'-dithiopropionic acid sulfosuccinimidyl ester (sulfo-SAND), sulfosuccinimidyl-2-(p-azidosalicylamido)ethyl-1,3-dithiopropionic acid ester (sulfoSASD), sulfo-SIAB, sulfo-SMCC, sulfo-SMPB, etc.

[0281] Additional examples of conjugation reactions that can be used to attach a molecule to a substrate include Ullmann reaction, Heck reaction, Negishi reaction, Stille reaction, Suzuki reaction, Buchwald-Hartwig coupling, Castro-Stevens coupling, Glaser coupling, Kumada coupling, Larock indole synthesis, Miyaura borylation, Sonagashira cross-coupling, Grubbs reaction.

[0282] More than one type of molecule can be coupled to a substrate. For example, a substrate can be coupled to a nucleic acid molecule and a peptide. Alternatively, a substrate can be coupled to only one type of molecule (e.g., only nucleic acid molecules, only peptides, only lipids, only carbohydrates, etc.). The substrate can be coupled to any useful combination of molecules, joints, reactive moieties, or functional groups, which can be coupled with any useful density, as described elsewhere herein. For example, a multifunctional joint can be used to attach both nucleic acid barcode molecules and peptides to a substrate. Alternatively, a substrate can include a joint and a reactive site; a joint can be used to attach one type of molecule (e.g., peptide or nucleic acid molecule), and a reactive site can be used to attach another type of molecule (e.g., nucleic acid molecule or peptide).

[0283] Various methods (e.g., self-assembled monolayers, patterning methods, linking moieties, etc.) can be used to control the proximity of a molecule coupled to a substrate to its nearest neighbor (e.g., another molecule). In some cases, it can be advantageous to have two molecules (e.g., two polymerizable molecules, such as a peptide and a nucleic acid molecule, or two nucleic acid molecules) that are in close proximity. For example, with respect to the sequencing methods described herein, a capture moiety can be used to couple a monomer of a polymer analyte, and after cleavage of the monomer, additional polymerizable molecules may be required to be adjacent to the capture moiety to allow transfer of information encoded by the polymerizable molecule of the binder. The proximity of molecules (e.g., capture moieties and polymerizable molecules) can be mediated using tethering molecules such as nucleic acid molecule "staples" or multifunctional linkers.

[0284] Nucleic acid molecules can be coupled with substrate by direct coupling. In such cases, substrate or nucleic acid molecules can include a functional part that can interact. For example, substrate and nucleic acid molecules can include complementary click chemistry pairs, such as alkynes and azide. In one such example, substrate can include alkyne moieties (e.g., DBCO), which can react with azide-functionalized nucleic acid molecules. Nucleic acid molecules can react with alkyne moieties in click chemistry to covalently link substrate to nucleic acid molecules. In another example, substrate can include avidin or streptavidin moieties, and biotinylated nucleic acid molecules can interact with this moiety and non-covalently bind. Alternatively or additionally, substrate can include nucleic acid molecules, and other nucleic acid molecules (e.g., nucleic acid analytes, nucleic acid linkers) are conjugated to the nucleic acid molecules using hybridization, connection, click chemistry, crosslinking (e.g., photocrosslinking such as CNVK).

[0285] Alternatively or additionally, a linker can be used to couple the nucleic acid molecule to the substrate, for example, as described elsewhere herein. The linker can comprise at least two functional groups that can couple to both the substrate and the nucleic acid molecule (e.g., a heterobifunctional linker). In an example, the substrate can comprise an amine group, and an alkyne-functionalized DNA primer (e.g., a DBCO-DNA primer) can be attached using a linker such as azidoacetic acid NHS ester. In another example, an amine-functionalized substrate can be coupled to an azide-functionalized DNA primer using a DBCO-NHS ester or DBCO-PEG-NHS ester linker.

[0286] Similarly, peptide can be coupled to substrate by direct coupling or by using joint.Peptide can be coupled to substrate at the end of peptide (for example, C-terminal or N-terminal), at the internal residue or amino acid of peptide, or along multiple positions of peptide.In the example of direct coupling, peptide functionalization can be used with the part that can interact with the part of substrate (for example, click chemistry pair, avidin-biotin).For example, substrate and peptide can include complementary click chemistry pair, for example, alkyne and azide, or the binding partner such as avidin and biotin.In an example of click chemistry pair, substrate can include alkyne moiety (for example, DBCO), which can react with azide functionalized peptide.Peptide can react with alkyne moiety in click chemistry reaction to covalently link substrate to peptide.In another example, substrate can include avidin or streptavidin moiety, and biotinylated peptide can interact with this part and non-covalently combine.

[0287] Alternatively or additionally, a linker can be used to couple the peptide to the substrate, for example, as described elsewhere herein. The linker can comprise at least two functional groups (e.g., heterobifunctional linkers) that can couple to both the substrate and the nucleic acid molecule. In an example, the substrate can comprise an amine group, and the alkyne-functionalized peptide can be attached using a linker such as azidoacetic acid NHS ester. In another example, an amine-functionalized substrate can be coupled to an azide-functionalized peptide using a DBCO-NHS ester or DBCO-PEG-NHS ester linker. In another example, an amine-functionalized substrate can be coupled to an azide-functionalized peptide using EDC and sulfo-NHS.

[0288] Peptide can be functionalized with functional moieties so that peptide can be attached or coupled to substrate.Functional moieties can include silanes such as aminosilanes (such as APTES), amino-PEG-silanes, click chemistry moieties or other connecting moieties and can be attached to peptides at peptide ends (N- or C-), at internal amino acids or at multiple positions (such as multiple internal amino acids, one or two ends, etc.).Chemical methods for peptide functionalization can include C-terminal specific conjugation (for example, via C-terminal decarboxylation alkylation) using photoredox catalysis, for example, as described by Bloom et al., Nature Chemistry 10,205-211.2018. and Zhang et al., ACS Chem.Biol.2021,16,11,2595–2603, each of which is incorporated herein by reference in its entirety, or coupled to amine-functionalized surface amides.N-terminal attachment can include coupling N-terminal amino amides to carboxyl-functionalized surfaces or using 2-pyridinecarboxaldehyde variants. Alternatively or additionally, functionalization of the termini of the peptide can be achieved enzymatically or using enzyme analogs such as ribozymes or deoxyribozymes. In examples of enzymatic functionalization and attachment, carboxypeptidases or amidases are used for C-terminal functionalization (e.g., as described in Xu et al., ACS ChemBiol. 2011 Oct 21; 6(10): 1015–1020; Zhu et al., Chinese Chemical Letters. 2018, Vol. 29 No. 7, pp. 1116-1118; and Zhu et al., ACS Catal. 2022, 12, 13, 8019–8026, each of which is incorporated herein by reference in its entirety), which can allow click chemistry moieties to be added to the peptide. Then, click chemistry functionalized peptide can be directly attached to substrate via another clickable group (such as BCN-azide or DBCO-azide coupling), or in other cases, can react with another joint or polymerizable molecule (such as, with the bait nucleic acid molecule of clickable group), then this joint or polymerizable molecule can be directly or indirectly connected to substrate (such as, using capture nucleic acid molecule and hybridizing with bait nucleic acid molecule). Other examples that can be used for functionalized or attached enzyme include sortase A, subtiligase, Butelase I or trypsiligase. In some examples, ubiquitin ligase can be used for ubiquitin protein with linker moiety attached to substrate. Then, these linker moieties can be used for chemical attachment of protein to ubiquitin coupled substrate. In some examples, glycosylases can be used to conjugate functionalized sugar groups (e.g., click chemistry functionalized sugars, polymer-conjugated sugars, biotinylated sugars) to amino acid residues, which can allow attachment to substrates (e.g., via click chemistry, polymer cross-linking or nucleic acid hybridization, avidin-biotin interactions), etc.Internal amino acid residues or post-translationally modified residues can be coupled to the substrate using, for example, thiol labeling, amide coupling to glutamic acid or aspartic acid residues using EDC / NHS chemistry or DMT-MM, esterification of glutamic acid or aspartic acid residues, alkylation or disulfide bridge labeling of cysteine, or amide coupling to lysine residues.

[0289] The peptide can be processed before, during or after the peptide is coupled to the substrate. In some cases, the peptide is conjugated to a tag that enables attachment to the substrate, for example, using a His tag, a SNAP tag, a CLIP tag, a spy trap, a spy tag, a nucleic acid tag (for example, a bait oligonucleotide that can be attached to a capture oligonucleotide of a substrate). In some examples, it may be advantageous to block or protect primary amines or carboxyl groups and optionally to block or protect the N-terminal primary amine or the C-terminal carboxyl group to promote attachment of the N-terminal or C-terminal to the substrate. In an example, the single point (for example, C-terminal) selective attachment of a peptide can be achieved by reacting the peptide with a linker comprising an amine reactive group (for example, an isothiocyanate such as PITC) and a reactive group (for example, a click chemistry group). The linker can be, for example, a PITC-conjugated click chemistry moiety, such as PITC-azide, PITC-alkyne, optionally with a spacer moiety in between, such as PITC-alkyl-azide, PITC-PEG-azide, PITC-alkyl-alkyne, PITC-PEG-azide. The linker reacts with and "blocks" primary amines (e.g., modified lysines), including the N-terminus. Subsequent cleavage of the N-terminal amino acid can be performed (e.g., using an Edman reagent, such as an acid), and one of the remaining modified lysines can be attached to a substrate (e.g., using a click chemistry moiety coupled to an amine-reactive group). Optionally, the peptide can be treated with a protease, e.g., LysC, which cleaves the peptide such that the remaining peptide has a C-terminal lysine and such that the remaining peptide contains only a primary amine at the C-terminal lysine residue and the N-terminus; such cleavage can be performed prior to reacting the amine-reactive group, e.g., as shown by Xie et al. Langmuir 2022, 38, 30, 9119–9128, which is incorporated herein by reference in its entirety.

[0290] In another embodiment, the carboxyl group of the present invention can be used for the preparation of the peptide of the present invention.Similarly, carboxyl group can react in a manner that can carry out C-terminal or internal residue attachment.In the example of C-terminal conjugation, when processed with activating reagent (for example, acetic anhydride) to generate " blocking " carboxyl group on peptide-thiohydantoin (at C-terminal) and aspartic acid and glutamic acid residues, carboxyl group can be marked with C-terminal sequencing reagent (such as isothiocyanate).Thiohydantoin reaction can then be made to be coupled to substrate.Alternatively, single reactive carboxyl group is only exposed at C-terminal amino acid place via single round C-terminal sequencing degradation or via protease cutting C-terminal amino acid.Then single reactive C-terminal carboxyl group can be used as the reactive part of single attachment site.

[0291] In another approach, peptides or proteins can be attached via the N-terminus using the specific reactivity of the N-terminal amine group. Amine-based reactions, such as amide coupling, can be performed at low pH, where only the N-terminal amine group is reactive. In addition, 2-pyridinecarboxaldehyde and variants can be used to react with N-terminal amine groups.

[0292] In some cases, the peptide can be conjugated to the substrate using polymerization reactions, e.g., free radical polymerization, such as Michael-type addition using PEGylated peptides, methacrylamide-modified peptides, maleimide-terminated oligoNIPAAM-conjugated peptides; photocrosslinking of phenylazo-conjugated peptides or other polymerization reactions of peptides conjugated to monomers, e.g., as described by Krishna et al. Biopolymers. 2010; 94(1): 32-48, which is incorporated herein by reference in its entirety.

[0293] The molecule of various types can be attached to substrate.Substrate can include any combination of its coupled molecule, including but not limited to peptide, protein (for example, enzyme, antibody, nano antibody, antibody fragment), nucleic acid molecule, lipid, carbohydrate or sugar, metabolite, small molecule, polymer, metal, viral particle, biotin, avidin, streptavidin, neutravidin etc.The molecule of various types can be attached to substrate simultaneously or attach in a sequential manner.For example, substrate can be processed to put together nucleic acid molecule and subsequently processed to put together peptide, or alternatively, substrate can be processed to put together peptide before nucleic acid molecule.Any number of putting together or attachment can be used.For example, when a plurality of types of molecules are attached to substrate, any number of putting together chemistry can be used for the molecule of each type.

[0294] In some embodiments, the substrate or its part can be subjected to the conditions sufficient to passivate substrate or its part.The passivation of substrate can be used for various purposes, such as preventing the non-specific binding of binding agent, changing the surface density of molecule (for example, increasing the density of nucleic acid molecules or peptide), blocking reactive sites (for example, molecule blocks available click chemistry part after being conjugated on substrate) etc..Passivation can be realized using chemical methods, for example, deposition blockers such as protein (for example, albumin), Tween-20, polymer, metal or metal oxide, or using biochemical method to realize, for example, using metallomicron. The substrate comprising reactive part can also be deactivated after molecule conjugation (for example, nucleic acid molecules, peptide etc.) by making any unreacted site and suitable molecular reaction. For example, the substrate comprising click chemistry part (for example, DBCO beads) can use click chemistry (for example, azide-nucleic acid molecules, azide-peptide) to be coupled to molecule of interest (for example, polymerizable molecule such as nucleic acid molecules, peptide) with useful density. Unreacted sites can be passivated by providing and reacting a complementary click chemistry molecule (e.g., an azide-polymer (e.g., PEG-azide)), which can reduce downstream nonspecific interactions.

[0295] Substrate deactivation can occur at any useful time or step. For example, deactivation for blocking unreacted DBCO sites can be performed before, during, or after conjugation of an analyte or other molecule of interest (e.g., peptides and nucleic acid molecules). Deactivation can be controlled by the stoichiometry or density of the deactivating agent relative to the molecule of interest, or by physical methods (e.g., photopatterning, self-assembled monolayers, etc.).

[0296] Sample processing

[0297] The present disclosure also provides systems, compositions, devices, and methods for processing samples. One or more methods for processing a sample can include preparing a biological sample for analysis, which in some cases includes partitioning cells for single-cell analysis. Methods for processing a biological sample can include extracting or isolating one or more peptides or proteins from the biological sample for further processing and analysis, as described elsewhere herein.

[0298] Preparation of cell suspensions for single cell analysis: Methods described herein can relate to preparing single cell suspensions from biological samples. Single cell suspensions can be prepared from biological samples by dissociating cells and optionally cultivating them in liquid culture medium. In some cases, the biological sample comprises a liquid sample. For example, the biological sample can comprise a bacterial liquid culture, a mammalian liquid culture, blood, plasma, or serum sample. The processing of such liquid samples can comprise centrifugation (for example, to separate cells), resuspending the cells in a suitable culture medium (for example, Dulbecco's phosphate buffered saline (DPBS)), and optionally cultivating the cells separated.

[0299] In some embodiments, the present invention provides the method for the treatment of the cell of the present invention.Biological sample can comprise the cell of cultivation, for example the cell of suspension culture, or the cell that adheres to solid surface, such as petri dish or tissue culture dish.The adherent cell sample of cultivation can be processed to produce cell suspension, for example, via protease (such as trypsin), to separate cells from the surface.Biological sample can comprise tissue or biopsy sample.Can mechanically or enzymatically process tissue or biopsy sample to generate cell suspension.Such processing can comprise ultrasonic treatment (mechanical treatment) or enzyme treatment, such as using other enzymes of pronase, collagenase, hyaluronidase, metalloproteinase, trypsin or digestion extracellular matrix components.Then the cell of dissociation can be stored in a suitable buffer, such as DPBS.

[0300] Cell Sorting: Biological samples or cell suspensions can be sorted to isolate cells of interest. Sorting can be performed to select or separate cells based on their mass or characteristics (e.g., expression of a protein target, size, deformability, fluorescence or other optical properties, or other physical properties of the cells). Sorting can be accomplished using any number of methods, such as immunosorting (e.g., fluorescence activated cell sorting (FACS) or magnetic activated cell sorting (MACS)), electrophoretic methods, chromatography, microfluidic methods (e.g., using inertial focusing, cell traps, electrophoresis, isoelectric focusing), acoustic sorting, optical sorting (e.g., optoelectronic tweezers), mechanical cell picking (e.g., using a manual or robotic pipette), or passive methods (e.g., gravity sedimentation).

[0301] Partitions: Cells of a biological sample or cell suspension can be divided into separate partitions such that at least a subset of the separate partitions contain single cells. The separate partitions can contain barcode molecules (e.g., fluorophores or sets of fluorophores, nucleic acid barcode molecules, etc.). The barcode molecules can be unique to the partitions such that each separate partition contains a different barcode sequence than other partitions. The barcode molecules can be loaded into the separate partitions at any useful ratio of barcode molecules to sample types (e.g., cells, proteins, nucleic acid molecules). The barcode molecules can be loaded into the partitions such that approximately 0.0001, 0.001, 0.1, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 50, 100, 500, 1000, 5000, 10,000, or 200,000 barcodes are loaded per sample type. In some cases, the barcodes are loaded into the partitions so that more than about 0.0001, 0.001, 0.1, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 50, 100, 500, 1000, 5000, 10000, or 200,000 barcodes are loaded per sample type. In some cases, the barcodes are loaded into the partitions so that less than about 0.0001, 0.001, 0.1, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 50, 100, 500, 1000, 5000, 10000, or 200,000 barcodes are loaded per sample type.

[0302] Partitions can take on any useful geometric shape, such as droplets, microwells, solid substrates, gels (e.g., cells encapsulated in gel beads), beads, flasks, tubes, spots, capsules, channels, chambers, or other compartments or containers. Partitions can be part of a partition array, such as a droplet in a microfluidic device, a microwell in a microplate, a spot on a multi-spot array, etc.

[0303] Dissolution, permeabilization and analyte extraction: Single cells can be processed (e.g., in partitions) to obtain one or more analytes contained therein. Methods for processing single cells can include dissolving the cells to release the contents into separate compartments or partitions. Dissolution can be performed using detergents (e.g., Triton-X 100, sodium dodecyl sulfate, sodium deoxycholate, CHAPS), RIPA buffer, temperature changes (e.g., elevated or lower temperatures, freezing, freeze-thaw), enzymes, mechanical dissolution (e.g., ultrasonic treatment, application of mechanical force), electrolysis, or a combination thereof. Dissolution can be performed in the presence of protease inhibitors to prevent degradation or digestion of proteins from the cells. The contents can optionally be further processed, for example, by purification or extraction, denaturation of proteins or peptides, enzymatic or chemical digestion, etc. In some cases, the contents can be enzymatically digested to remove nucleic acid molecules, for example, using nucleases (such as DNA enzymes or RNA enzymes). Alternatively or additionally, the cells can be fixed (e.g., using a fixative) and / or permeabilized. Examples of fixatives include aldehydes (e.g., glutaraldehyde, formaldehyde, paraformaldehyde), alcohols (e.g., methanol, ethanol), acetone, acids (e.g., acetic acid, Davidson's AFA), oxidants (e.g., osmium tetroxide, potassium dichromate, chromic acid, permanganate), Zenker's fixative, picrates, Hepes-glutamate buffer-mediated organic solvent protection effect (HOPE), or Karnovsky's fixative. Cell permeabilization can be achieved mechanically (e.g., using sonication, electroporation, shearing) or chemically (e.g., using organic solvents such as methanol or acetone, or detergents such as saponin, Tween-20, Triton X-100).

[0304] Protein treatment: The biological sample (or single cell suspension or partitioned cells) can be further treated to enable proteomic analysis. For example, protein disaggregation in the sample can be performed, for example using chemical or mechanical methods. Chemical disaggregation methods can include, but are not limited to, sodium dodecyl sulfate (SDS), Triton-X100, 3-((3-cholamidopropyl)dimethylammonium)-1-propanesulfonate (CHAPS), ethylene carbonate, or formamide. Mechanical disaggregation methods can include, but are not limited to, sonication or high temperature treatment. The biological sample (or single cell suspension or partitioned cells) can be subjected to conditions sufficient to denature one or more proteins. Denaturation can be achieved using heat, chemicals (e.g., SDS, urea, guanidine), reducing agents (e.g., dithiothreitol (DTT), β-mercaptoethanol, TCEP), urea, enzymes (e.g., ClpX, ClpS, depolymerase), ribozymes, or deoxyribozymes. Similarly, the peptide or protein can be subjected to conditions that dissolve the peptide or protein in solution, such as using detergents, organic solvents, spermidine, or tagging the peptide or protein with a polyionic tag (e.g., DNA, PEG or other polymer). Alternatively or additionally, the peptide or protein can be enriched or purified; in examples, the peptide or protein of interest can be precipitated using trichloroacetic acid, chloroform, trizole or other chemical reagents. Other biological or chemical reagents can be included during protein processing, such as lysozyme, papain, cruzain, trypsin, protease inhibitors, nucleases or nuclease-containing proteins (e.g., DNA enzymes, RNA enzymes, DNA glycosylases, restriction endonucleases, transposases, micrococcal nucleases, Cas proteins).

[0305] Can be before analysis by peptide or protein fragmentation.Protein fragmentation can be used to reduce the size of protein, and allow the effective processing of peptide, as described in other places herein.Can use protease (such as trypsin, chymotrypsin, pepsin, Lys-C, Glu-C, proteinase K, furin, thrombin, endopeptidase, papain, subtilisin, elastase, enterokinase, genenanse, endoproteinase, metalloproteinase), or with chemical treatment (such as cyanogen bromide, hydrazine, hydroxylamine, formic acid, BNPS-skatole, iodine benzoic acid, 2-nitro-5-thiocyanatobenzoic acid etc.) carry out fragmentation.Alternately or additionally, mechanical method can be used to carry out fragmentation, such as ultrasonic process, eddy current, mechanical agitation, use temperature variation (such as, freeze / thaw, heating) or other fragmentation methods.

[0306] Enrichment of proteins or peptides in biological samples can be carried out, for example, for separating proteins and peptides from cell debris or other types of analytes (for example, nucleic acids, lipids, carbohydrates, metabolites). Such enrichment can include, for example, using affinity columns (for example ion exchange), size exclusion columns, affinity precipitation (for example immunoprecipitation), chemical precipitation (for example, using trichloroacetic acid, chloroform, trizole), chromatography (for example HPLC) or electrophoresis. Before enrichment, in the case of cell partitioning, microbeads, affinity microcolumns, affinity beads etc. can be used for enrichment. In some cases, protein or peptide can be subjected to fractional separation, which can be used to separate proteins by size, hydrophobicity, charge, affinity, size, quality, density etc.

[0307] Peptides can be barcoded in batches or in partitions. Peptides can be barcoded with any useful type of barcode molecule (e.g., spectral or fluorescent barcodes, mass tags, nucleic acid barcode molecules, etc.). Barcode molecules can allow identification of origin peptides, partitions, samples, cells, or cell compartments. For example, a cell sample can be partitioned so that the partition contains at most one cell; a partition can contain a unique barcode molecule (e.g., a nucleic acid barcode molecule) that identifies the partition and therefore identifies the cell. Subsequently, peptides within the partition are labeled with barcode molecules (e.g., by permeabilization or dissolution of cells) to identify the peptides as being produced from or originating from the same cell or partition. In other examples, a substrate can include a nucleic acid molecule that includes a unique barcode sequence that is different from the barcode sequence of other substrates. Therefore, a barcode sequence can be used to identify a substrate. In some cases, a cell sample can be used to partition a barcoded substrate so that at least a subset of the partition includes a single cell and a single barcoded substrate. Therefore, peptides generated from a single cell and transferred to a barcoded substrate can all be identified as being derived from the single cell. Barcode molecules can contain additional useful functional sequences, such as UMIs, primer sites, restriction sites, cleavage sites, transposition sites, sequencing sites, read sites, etc.

[0308] Attachment of the barcode molecule to the peptide can be achieved using any suitable chemistry. For example, C-terminal conjugation of a nucleic acid barcode molecule can be achieved by amide coupling of an amine-conjugated DNA barcode molecule to the peptide or by thiol alkylation (e.g., reacting a thiolated peptide with an alkylated (e.g., iodoacetamide) DNA barcode molecule). N-terminal conjugation can be achieved, for example, using a 2-pyridinecarboxaldehyde label of the DNA barcode and reacting with the N-terminus of the peptide. Internal residues, such as glutamic acid, can also be labeled with amine-conjugated DNA barcode molecules or carboxylated DNA barcodes (e.g., to react with the primary amine in lysine).

[0309] For a given peptide, a separate peptide can be barcoded at multiple locations. Peptides can be marked with the same or different barcode sequences at multiple sites. For example, a peptide can be partitioned into partitions comprising multiple identical barcode molecules, each comprising a barcode sequence unique to the partition. Peptides can be marked with a unique partition barcode sequence at a single or multiple sites, optionally each including a unique molecule identifier (UMI) so that subsequent downstream analysis (e.g., sequencing) can be attributed to the same peptide using the barcode sequence. In some cases, the end (e.g., N-terminus or C-terminus) or internal amino acids of the peptide can be barcoded. In some cases, the peptide can be fragmented before analysis or sequencing; therefore, multiple identical barcode molecules attached upstream to the same peptide can allow sequence analysis to be attributed back to a single peptide. Barcoding of the peptide can occur before, during, or after fragmentation. Peptides can be labeled with barcodes (e.g., nucleic acid barcode molecules) using any suitable chemistry (e.g., as described above) or using bifunctional or trifunctional linkers comprising multiple linking moieties (e.g., as described elsewhere herein, such as click chemistry moieties, NHS-esters, EDC, etc.). For example, C-terminal attachment can include amide coupling to the C-terminal carboxyl group or photoredox tagging of the C-terminal carboxyl group (e.g., adding an electrophile tag). N-terminal attachment can include amide coupling to the N-terminal amine group, where specific attachment can occur at low pH, or using 2-pyridinecarboxaldehyde variants for specific attachment to the N-terminus. Internal attachment can include, for example, amide coupling to glutamic acid or aspartic acid using EDC / NHS chemistry or DMT-MM; alkylation or disulfide bridge labeling of cysteine; or amide coupling to lysine residues.

[0310] In some examples, peptides can be labeled with different barcode molecules, which can be indexed by proximity to each other, for example, using primers that can anneal to adjacent barcode molecules. In one such method, after a protein has been labeled with multiple barcodes having different barcode sequences, a polymerase extension based on proximity can be used to copy and associate the sequence of the adjacent barcode. For example, each barcode molecule can include a primer binding site, to which a double primer adapter sequence including two sequences is annealed. The double primer adapter sequence can be bound to the primer binding sites of two adjacent barcodes. An extension reaction, for example, using a polymerase, can extend and copy the barcode sequence of the adjacent barcode. Subsequently, the double primer adapter sequence now having a copy of two adjacent barcodes can be removed and sequenced. From the sequencing reads, an adjacency matrix of the barcode sequence can be generated (for example, corresponding to the barcode sequence on a single double primer adapter adjacent in space). Thus, each of the barcode sequences can be associated with nearby neighboring barcode sequences, and thus, peptide portions can be aligned or grouped as contiguous. Such methods can be useful in cases where peptides are fragmented, so that individual fragments of the peptide can be aligned to the nearest neighbor fragments using the barcode sequences.

[0311] In another example, bridge amplification can be used to barcode peptides at multiple positions of a given peptide. In such methods, peptides or proteins can be labeled with nucleic acid primers at multiple sites. A nucleic acid barcode molecule can be provided that can anneal to a nucleic acid primer (not shown) or be connected to a nucleic acid primer. Subsequent rounds of bridge amplification can be performed to copy the nucleic acid barcode molecule to other primers located at other sites of a given peptide. In some examples, peptides can be tagged with multiple copies of a nucleic acid primer, and a barcode sequence can be sparsely provided so that only one nucleic acid primer per peptide is extended by polymerase extension. Subsequent rounds of bridge amplification can obtain peptides with the same barcode sequence at each nucleic acid primer. Subsequent fragmentation of the peptide can be performed so that the peptide fragments contain a single barcode on average. Therefore, in some cases, the output of such amplification methods can be a peptide with a separate barcode, which is generated from fragmenting a multi-tagged protein, wherein peptides from the same protein have the same barcode.

[0312] The cell sample can be partitioned into separate partitions or compartments (e.g., droplets, microwells) such that at least a subset of the partitions contain single cells. The partitions can then be treated with a lytic agent to lyse the cells and release proteins from the cells into the partitions. The proteins can then be labeled with a partition-specific barcode (e.g., using barcoded beads) such that all peptides or proteins produced from a single compartment contain the same barcode. In some examples, the barcode comprises a nucleic acid barcode molecule, and the barcode sequence can be used in downstream processing (e.g., via sequencing) to determine the partition or cell from which the peptide originated. The nucleic acid barcode molecule can comprise any additional useful sequence, such as a UMI, a primer sequence, etc.

[0313] Batch processing: Biological samples can be processed in batches. For example, biological samples can be processed to obtain a suspension of cells that can be directly dissolved in the suspension without partitioning the cells into separate compartments. Cells can be dissolved in batches using any useful method (e.g., as described above) and optionally further processed, such as homogenization, protease inhibition, denaturation, protein processing (e.g., chemical treatment, fragmentation), or a combination thereof. Biological samples can be pretreated prior to cell lysis or protein extraction. Such pretreatments can include removing debris, purification, filtration, concentration, or sorting.

[0314] Spatial barcoding: A biological sample can include a tissue sample comprising multiple cells. Tissue samples can be processed using methods that retain spatial information (e.g., to identify peptides from separate cells), for example, using spatial barcodes. For example, a 2D or 3D tissue sample can be provided, and separate cells or positions within the tissue sample can be contacted with multiple spatial barcodes (e.g., nucleic acid barcode molecules) comprising different barcode sequences. Different barcode sequences can be attributed to specific positions in a 2D or 3D tissue sample, which can correspond to the positions of the cells. For example, deterministic methods (such as two-photon patterning) or random methods (such as PCR) can be used to provide spatial barcodes to assign unique spatial barcodes to different sections of a 2D or 3D tissue sample. Thus, peptides labeled with spatial barcodes can be returned to a single position within a tissue sample, or to a single cell.

[0315] Computer system

[0316] The present disclosure provides computer systems programmed to implement the methods of the present disclosure. Figure 4A computer system 401 is shown that is programmed or otherwise configured to detect a detectable complex. In some embodiments, the computer system is programmed or otherwise configured to detect a detectable complex using a sequencing agent of Formula I. In some embodiments, the detectable complex comprises an amino acid. In some embodiments, the detectable complex comprises a peptide. In some embodiments, the computer system is programmed or otherwise configured to provide sequencing data for a polypeptide. Computer system 401 can adjust various aspects of the detection of the present disclosure, such as providing sequencing data for a polypeptide. Computer system 401 can be a user's electronic device or a computer system remotely located relative to the electronic device. The electronic device can be a mobile electronic device.

[0317] Computer system 401 includes a central processing unit (CPU, also referred to herein as a "processor" and "computer processor") 405, which can be a single-core or multi-core processor, or multiple processors for parallel processing. Computer system 401 also includes memory or storage location 410 (e.g., random access memory, read-only memory, flash memory), electronic storage unit 415 (e.g., a hard disk), a communication interface 420 (e.g., a network adapter) for communicating with one or more other systems, and peripheral devices 425 (such as cache, other memory, data storage, and / or an electronic display adapter). Memory 410, storage unit 415, interface 420, and peripheral devices 425 communicate with CPU 405 via a communication bus (solid lines), such as a motherboard. Storage unit 415 can be a data storage unit (or data repository) for storing data. Computer system 401 can be operatively coupled to a computer network ("network") 430 via communication interface 420. Network 430 can be the Internet, an internetwork, and / or an extranet, or an intranet and / or extranet in communication with the Internet. In some cases, network 430 is a telecommunications and / or data network. Network 430 may include one or more computer servers that can enable distributed computing, such as cloud computing. In some cases, with the aid of computer system 401, network 430 may implement a peer-to-peer network that enables devices coupled to computer system 401 to act as clients or servers.

[0318] The CPU 405 can execute a series of machine-readable instructions, which can be embodied in a program or software. The instructions can be stored in a storage location, such as memory 410. The instructions can be directed to the CPU 405, which can then program or otherwise configure the CPU 405 to implement the methods of the present disclosure. Examples of operations performed by the CPU 405 can include fetching, decoding, executing, and writing back.

[0319] CPU 405 may be part of a circuit, such as an integrated circuit. One or more other components of system 401 may be included in the circuit. In some cases, the circuit is an application-specific integrated circuit (ASIC).

[0320] Storage unit 415 can store files such as drivers, libraries, and saved programs. Storage unit 415 can store user data, such as user preferences and user programs. In some cases, computer system 401 can include one or more additional data storage units external to computer system 401, such as located on a remote server that communicates with computer system 401 via an intranet or the Internet.

[0321] Computer system 401 can communicate with one or more remote computer systems via network 430. For example, computer system 401 can communicate with a user's remote computer system. Examples of remote computer systems include personal computers (e.g., portable PCs), tablets or tablet-like PCs (e.g., iPad, Galaxy Tab),...

Claims

1. A sequencing reagent of formula I: or a stereoisomer, tautomer or salt thereof, wherein: A comprises a first reactive group configured to form a covalent bond with the N-terminal amino acid of the polypeptide, wherein the first reactive group is a dithioester or a thiocarbamoyl group; B comprises a second reactive group; and L 1 Includes a linker coupled to A and B.

2. The sequencing reagent of claim 1, wherein the second reactive group is covalently attached or configured to be covalently attached to a polymer.

3. The sequencing reagent of claim 2, wherein the polymer comprises polyethylene glycol (PEG).

4. The sequencing reagent according to claim 2 or 3, wherein the polymer comprises deoxyribonucleic acid (DNA) or ribonucleic acid (RNA).

5. The sequencing reagent of any one of claims 2 to 4, wherein the polymer is covalently attached or configured to be covalently attached to a surface.

6. The sequencing reagent of any one of claims 1 to 5, wherein the second reactive group is covalently linked or configured to be covalently linked to a surface-bound linker.

7. The sequencing reagent of claim 6, wherein the second reactive group is covalently attached to the surface-bound adapter, and wherein the surface-bound adapter is attached to a surface.

8. The sequencing reagent of claim 6 or 7, wherein the surface-bound linker comprises an ethyl group.

9. The sequencing reagent of claim 6 or 7, wherein the surface-bound linker comprises a propyl group.

10. The sequencing reagent of claim 6 or 7, wherein the surface-bound adapter comprises a nucleic acid molecule.

11. The sequencing reagent of any one of claims 1 to 10, wherein the second reactive group comprises: in Indicates the orientation of the second reactive group relative to the first reactive group.

12. The sequencing reagent of any one of claims 1-10, wherein the second reactive group comprises a click chemistry moiety.

13. The sequencing reagent of claim 12, wherein the click chemistry moiety is azide, alkyne, dibenzocyclooctyne (DBCO), tetrazine, or trans-cyclooctene (TCO).

14. The sequencing reagent of any one of claims 1 to 13, wherein the first reactive group comprises: in Indicates the orientation of the first reactive group relative to the second reactive group.

15. The sequencing reagent of any one of claims 1 to 13, wherein the first reactive group comprises: in Indicates the orientation of the first reactive group relative to the second reactive group.

16. The sequencing reagent of claim 15, wherein the first reactive group comprises: in Indicates the orientation of the first reactive group relative to the second reactive group.

17. The sequencing reagent of any one of claims 1 to 13, wherein the first reactive group comprises: in Indicates the orientation of the first reactive group relative to the second reactive group.

18. The sequencing reagent of any one of claims 1-13, wherein the first reactive group comprises a thioacetyl group.

19. The sequencing reagent of any one of claims 1-13, wherein the first reactive group comprises a thiobenzoyl group.

20. The sequencing reagent of any one of claims 1-13, wherein the first reactive group comprises a derivative of N-thiobenzoylsuccinimide.

21. The sequencing reagent of any one of claims 1-13, wherein the first reactive group comprises a derivative of cyanomethyldithiobenzoate.

22. The sequencing reagent of any one of claims 1 to 13, wherein the first reactive group comprises: in Indicates the orientation of the first reactive group relative to the second reactive group.

23. The sequencing reagent of any one of claims 1 to 13, wherein the first reactive group comprises in Indicates the orientation of the first reactive group relative to the second reactive group.

24. The sequencing reagent of any one of claims 1 to 23, wherein L1 comprises a cleavable linker.

25. The sequencing reagent of claim 24, wherein the cleavable linker comprises a disulfide bond, a hydrazone, a DNA molecule comprising a cleavage site, an enzyme-cleavable peptide, or a click chemistry moiety.

26. The sequencing reagent of claim 25, wherein the cleavable linker comprises a hydrazone.

27. The sequencing reagent of claim 24, wherein the cleavable linker comprises o-aminobenzylhydrazone.

28. The sequencing reagent of any one of claims 1-24, wherein L1 comprises a non-cleavable linker.

29. The sequencing reagent of any one of claims 1 to 28, wherein the first reactive group is linked to the N-terminal amino acid of the polypeptide via a covalent bond.

30. The sequencing reagent of any one of claims 1-29, wherein the second reactive group is directly or indirectly attached or is configured to be directly or indirectly attached to a substrate.

31. The sequencing reagent of claim 30, wherein the first reactive group is linked to the N-terminal amino acid of the polypeptide via a covalent bond, wherein the polypeptide is linked to a substrate; and wherein the second reactive group is directly or indirectly linked to the substrate.

32. A method of using the sequencing reagent of any one of claims 1 to 31, comprising: a. Providing a capture moiety and a polymer analyte; b. contacting the polymer analyte with the sequencing reagent, wherein the sequencing reagent binds to the monomer of the polymer analyte to form a sequencing reagent-monomer complex; c. tethering the sequencing reagent-monomer complex to the capture moiety; d. cleaving the sequencing reagent-monomer complex from the polymer analyte to provide a detectable complex; and e. detecting the detectable complex.

33. The method of claim 32, wherein the polymer analyte comprises a polypeptide and the method comprises contacting the polypeptide with an alkylating agent prior to contacting the polypeptide with the sequencing reagent.

34. The method of claim 33, wherein the alkylating agent comprises 4-vinylpyridine.

35. The method of claim 33, wherein the alkylating agent comprises iodoacetamide.

36. The method of claim 32, wherein the polymer analyte comprises a polypeptide.

37. The method of any one of claims 32-36, wherein the monomer comprises a terminal amino acid residue.

38. The method of any one of claims 32-37, wherein the capture moiety is bound to a substrate.

39. The method of any one of claims 32-37, wherein the capture moiety is coupled to the polymer analyte.

40. The method of any one of claims 32-39, wherein the capture moiety comprises a DNA molecule.

41. The method of any one of claims 32-40, wherein detecting the detectable complex comprises contacting the sequencing reagent-monomer complex with a binding agent.

42. The method of claim 41, wherein the binding agent comprises an antibody, a nanobody, a single-chain variable fragment (scFv), or an aptamer.

43. The method of claim 41 or 42, wherein the binding agent comprises a polymerizable molecule.

44. The method of claim 43, further comprising coupling the polymerizable molecule to the capture moiety or another polymerizable molecule, and wherein the detecting comprises sequencing the polymerizable molecule.

45. The method of any one of claims 41-43, wherein the capture moiety is coupled to the polymer analyte, and further comprising partitioning the sequencing reagent-monomer complex into partitions using the binding agent, and in the partitions, coupling the barcode molecule to the capture moiety.

46. ​​The method of claim 41 or 42, wherein the binding agent comprises a fluorophore, and wherein the detecting comprises imaging the fluorophore.

47. The method of any one of claims 32-46, further comprising repeating (b)-(d).

48. The method of any one of claims 32-47, wherein the detecting comprises translocating the detectable complex through a nanopore and identifying the monomer.

49. The method of any one of claims 32-48, further comprising repeating the method to sequence the polymer analyte.

Citation Information

Patent Citations

  • Single-molecule protein and peptide sequencing

    US11499979B2

  • Single-Molecule Protein and Peptide Sequencing

    US20200217853A1

  • Methods of generating nanoarrays and microarrays

    WO2019195633A1

  • Methods and systems for processing polymeric analytes

    WO2023196642A1