Continuous encoding method and related kit

The method of nucleic acid encoding and barcoding addresses proteomics challenges by enabling efficient, parallelized, and sensitive polymer analysis, particularly for protein sequencing, through the use of binders and enzymes to transfer information from coding to recording tags.

JP2026136372APending Publication Date: 2026-08-25ENCODIA INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2026093982
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2020-08-19
Filing Date
2026-06-04
Publication Date
2026-08-25

AI Technical Summary

Technical Problem

Current proteomics techniques face challenges in efficiently, accurately, and high-throughput identification and quantification of proteins, with limitations in sensitivity, dynamic range, cross-reactivity, and background signaling, particularly in multiplexing readouts of affinityant sets for homogeneous macromolecules.

Method used

A method involving nucleic acid encoding and barcoding of molecular recognition events, utilizing binders, nucleic acid conjugation reagents, polymerases, and cleavage enzymes to transfer information from coding tags to recording tags, enabling high-throughput polymer analysis.

Benefits of technology

Enables efficient, parallelized, and sensitive polymer analysis, including protein sequencing, with reduced DNA-DNA interactions and simplified reaction steps, facilitating accurate characterization and quantification of macromolecules.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026136372000003
    Figure 2026136372000003
  • Figure 2026136372000004
    Figure 2026136372000004
  • Figure 2026136372000005
    Figure 2026136372000005
Patent Text Reader

Abstract

Providing a continuous encoding method and related kits. [Solution] This disclosure relates to methods and kits for analyzing polymers. In some embodiments, this disclosure relates to polymer analysis methods using barcoding and nucleic acid encoding of molecular recognition events. Also provided herein are methods and related kits for transferring information using multiple enzymes, which include performing ligation, extension, and cleavage reactions with nucleic acid molecules associated with a polymer for analysis. In some embodiments, the polymer for analysis includes peptides, polypeptides, or proteins.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] Related applications This application claims priority to U.S. Provisional Patent Application No. 63 / 067,744, filed on 19 August 2020, the disclosure of which is incorporated herein by reference in its entirety for all purposes.

[0002] Sequence list in ASCII text This patent or application file contains a sequence listing file submitted in computer-readable ASCII text format (filename: 4614-2002540_SeqList_ST25.txt, filed date: July 27, 2021, size: 2,469 bytes). The contents of the sequence listing file are incorporated herein by reference in their entirety.

[0003] This disclosure relates to methods and kits for analyzing polymers. In some embodiments, this disclosure relates to polymer analysis methods using barcoding and nucleic acid encoding of molecular recognition events. Also provided herein are methods and related kits for transferring information using multiple enzymes, which include performing ligation, extension, and cleavage reactions with nucleic acid molecules associated with the polymer for analysis. In some embodiments, the polymer for analysis includes peptides, polypeptides, or proteins. [Background technology]

[0004] The highly parallel characterization and recognition of macromolecules such as proteins remains a challenge. In proteomics, one goal is to identify and quantify a large number of proteins in a sample, a task that is difficult to achieve in a high-throughput manner. Assays such as immunoassays and mass spectrometry-based methods are used, but they are limited to both sample and analyte levels, have limited sensitivity and dynamic range, and have potential problems with cross-reactivity and background signaling. For example, multiplexing the readout of an affinityant set to an aggregate of homogeneous macromolecules using affinityants with detectable labels remains difficult. Along with improved techniques related to macromolecular analysis, there remains a need for protein sequencing and / or analysis, as well as the application of products, methods, and kits to achieve it. Proteomics techniques are needed to perform macromolecular analysis that is efficient, highly parallelized, accurate, sensitive, and high-throughput. This disclosure satisfies these and other related needs.

[0005] These and other aspects of the present invention will become apparent by reference to the following embodiments for carrying out the invention. For this purpose, various references describing in more detail specific background information, procedures, compounds, and / or compositions are listed herein, each of which is incorporated herein in whole by reference. [Overview of the project] [Means for solving the problem]

[0006] This summary of the present invention is not intended to be used to limit the scope of the claimed subject matter. Other features, details, usefulness, and advantages of the claimed subject matter will become apparent from the embodiments for carrying out the invention, including the embodiments disclosed in the accompanying drawings and the accompanying claims.

[0007] Provided herein is a method for analyzing a polymer, comprising the steps of: providing a polymer and an associated recording tag bonded to a support; contacting the polymer with a binder capable of binding to the polymer and comprising a coding tag having identification information relating to the binder, thereby enabling binding between the polymer and the binder; bonding the 5' end of the recording tag to the 3' end of the coding tag using a nucleic acid conjugation reagent; extending the recording tag using the coding tag as a template with a polymerase to produce a double-stranded extended recording tag; cleaving the double-stranded extended recording tag with a double-stranded nucleic acid cleavage reagent to generate a 3' overhang within the extended recording tag; thereby transferring information from the coding tag to the recording tag to produce the extended recording tag.

[0008] Provided herein is a binder including a coding tag, comprising a binder, a nucleic acid conjugation reagent, a polymerase, and a double-strand nucleic acid cleavage reagent, wherein the coding tag includes identification information relating to the binder, and the binder is configured to bind to a polymer associated with a recording tag, and the identification information from the coding tag is configured for transfer from the coding tag to the recording tag associated with the polymer. In some embodiments, the kit comprises multiple binders. In some embodiments, the nucleic acid conjugation reagent, polymerase, and double-strand nucleic acid cleavage reagent are provided as a mixture. [Brief explanation of the drawing]

[0009] Non-limiting embodiments of the present invention will be described by reference to the accompanying drawings. The drawings are schematic and not intended to be to exact scale. Not all components are indicated in all drawings for illustrative purposes, and not all components of each embodiment of the present invention are shown if not necessary to illustrate them to those skilled in the art.

[0010] [Figure 1A]An exemplary polymer analysis assay involving information transfer using the method provided herein is shown. In Figure 1A, the polymer to be analyzed (e.g., a peptide) is ligated to a recording tag hairpin immobilized on a support using a ligation reaction, and the structure is cleaved for subsequent steps. The first (far left) panel of Figure 1A shows a capture nucleic acid hairpin with a recessed 5' phosphorylated end. The second panel of Figure 1A shows the peptide to be analyzed attached to a bait nucleic acid that hybridizes (via a reactive coupling moiety) to the capture nucleic acid hairpin immobilized on a support. Ligating the bait nucleic acid to the capture nucleic acid. Ligating the recording tag to the hairpin. Block-labeled barcodes (BCs) indicate any barcode that can be attached to the polymer and incorporated into the recording tag, e.g., a sample-specific barcode and / or UMI. Block-labeled restriction enzyme recognition sites (RSs) represent the incorporated sequence for recognition and cleavage by an IIS-type restriction enzyme. The third panel of Figure 1A shows polymerase elongation for creating a double-stranded DNA (dsDNA) construct to which the peptide is attached. The fourth panel of Figure 1A shows the dsDNA construct after digestion using an IIS-type RE, which generates a 3' overhang (2-base pair sequence) with a recessed 5' phosphorylated end. In this way, a record tag containing one or more barcodes is prepared and available for information transfer from a coding tag. [Figure 1B]Figure 1A shows the encoding cycle with the generated structure. The left panel of Figure 1B shows the binder bound to the peptide, bringing the coding tag attached to the binder close to the recording tag. In some embodiments, the binder can attach to or ligate to the coding tag at locations not depicted (e.g., the coding tag or other loop regions). The shown binder is attached to the coding tag by a linker. The coding tag contains a binder-specific barcode (BBC), a 2bp spacer, and an IIS-type restriction enzyme site (RS). The center panel of Figure 1B shows the products of the first two enzymatic reactions. Upon ligation of the 5' end of the recording tag to the 3' end of the coding tag, the polymerase extends the 3' (unligated) end of the recording tag to create a dsDNA molecule containing a 2base pair spacer adjacent to each IIS-type RE site. Following the double helix, the IIS-type RE binds and cleaves adjacent to its recognition site. The right panel of Figure 1B shows the final product after all three enzymatic steps, in which the dsDNA contains a 2nt 3' overhang (OH) that functions as a binder-specific barcode and spacer sequence. In some embodiments, after the information transfer cycle, a portion of the polypeptide for analysis can be removed from the polypeptide. The cycle of steps shown in Figure 1B can be repeated one or more times with additional binders and coding tags to further extend the recording tag. [Figure 2]Exemplary encoding results generated by the encoding methods shown in Figures 1A and 1B. Two test polypeptides (F-peptide, SEQ ID NO: 6 and L-peptide, SEQ ID NO: 7) were ligated to immobilized bead-attached nucleic acid recording tags. The test polypeptides were brought into contact with two binders (F-binder and L-binder) attached to the corresponding coding tags. In the ligation step, three different concentrations of T4 DNA ligase were tested. After ligation, a Krenow fragment was added for the extension step and a BtsI-V2 enzyme was added for the cleavage step. Fractions of the encoded recording tags were evaluated by NGS sequencing, showing specific encoding results for both binders. Each error bar is constructed using one standard error from the mean. [Modes for carrying out the invention]

[0011] Provided herein are methods and kits for analyzing polymers. In some embodiments, the analysis utilizes barcoding and nucleic acid encoding of molecular recognition events. The provided method includes (a) providing a polymer and associated recording tags bonded to a support; (b) contacting the polymer with a binder that can bond to the polymer and contains a coding tag having identification information relating to the binder, thereby enabling bonding between the polymer and the binder; (c) bonding the 5' end of the recording tag to the 3' end of the coding tag using a nucleic acid conjugation reagent; (d) extending the recording tag using the coding tag as a template with polymerase to produce a double-stranded extended recording tag; and (e) cleaving the double-stranded extended recording tag with a double-stranded nucleic acid cleavage reagent to produce a 3' overhang within the extended recording tag. Performing these steps transfers information from the coding tag to the recording tag to produce the extended recording tag. Also provided are kits containing components and / or reagents for performing the methods provided for polymer sequencing and / or analysis. In some embodiments, the kit also includes instructions for using the kit to perform one of the methods provided herein.

[0012] The highly parallel characterization and recognition of macromolecules such as proteins remains a challenge. In proteomics, one goal is to identify and quantify a large number of proteins in a sample, a task that is difficult to achieve in a high-throughput manner. Assays such as immunoassays and mass spectrometry-based methods are used, but they are limited to both the sample and analyte levels, and have limitations in sensitivity, dynamic range, cross-reactivity, and background signaling. For example, multiplexing the readout of an affinityant set to an aggregate of homogeneous macromolecules using affinityants with detectable labels remains difficult. Along with improved techniques related to macromolecular analysis, there remains a need for protein sequencing and / or analysis, and its application to products, methods, and kits to achieve this. Proteomics techniques are needed to perform macromolecular analysis that is efficient, highly parallelized, accurate, sensitive, and high-throughput. This disclosure satisfies these and other related needs.

[0013] In some embodiments, the disclosure provides methods for analyzing polymers, including information transfer, which involves, in part, the characterization, quantification, and / or direct application to protein and peptide characterization and sequencing. In some examples, the information transferred includes identification information relating to a binder configured to bind to the polymer. In some embodiments, a plurality of polymers obtained from a sample are analyzed. In some embodiments, the sample is obtained from a subject. In some embodiments, the polymer sequencing or analysis method includes detecting a plurality of polymers to be analyzed using a plurality of binders associated with coding tags.

[0014] The information transfer methods provided herein utilize multiple enzymes to perform ligation, extension, and cleavage reactions with nucleic acid molecules. In some embodiments, the methods provided include oligonucleotides containing hairpin structures and restriction enzyme sites (or parts thereof). In some embodiments, the methods include the use of a reaction system in which mixed enzymes are provided for the reaction. For example, the activity of polymerase, nucleic acid junctioning reagent, and double-strand nucleic acid cleavage reagent is provided under suitable conditions to transfer information from a coding tag to a recording tag to produce an extended recording tag. In the methods provided, the recording tag used includes at least a partially double-stranded DNA structure. Some advantages of using the described methods include high success in information transfer (encoding), a simple design for stepwise reactions, the option to perform in a single step / as a single pot reaction, reduced need for spacers or shortened spacer length, and / or minimization of DNA-DNA interactions within the system.

[0015] To provide a complete understanding of this disclosure, the following description contains many specific details. These details are provided for illustrative purposes only, and the claimed subject matter can be carried out in accordance with the claims without some or all of these specific details. It should be understood that other embodiments can be used and structural modifications can be made without departing from the scope of the claimed subject matter. It should be understood that the various features and functions described in one or more of the individual embodiments are not limited in terms of their applicability to the particular embodiment in which they are described. Rather, they can be applied, either alone or in some combination, to one or more of the other embodiments of this disclosure, whether such embodiments are described and whether such features are presented as part of the embodiment in which they are described. For clarity, technical materials known in the art relating to the claimed subject matter are not described in such detail that the claimed subject matter becomes unnecessarily obscured.

[0016] All publications, including patent documents, scientific papers, and databases, referenced in this application are incorporated by reference in whole for any purpose to the same extent that each individual publication is incorporated by reference individually. No citation of any publication or document is intended to acknowledge that any of them is relevant prior art, nor does it constitute any acknowledgment of the content or date of any such publication or document.

[0017] All headings are provided for the reader's convenience and should not be used to limit the meaning of the text that follows them, unless otherwise specified.

[0018] definition Unless otherwise defined, all technical and scientific terms used herein have the same meanings as those commonly understood by those skilled in the art in which this disclosure pertains. If any definition set forth in this chapter contradicts or is inconsistent with any definition set forth in the patents, applications, published applications and other publications incorporated herein by reference, the definition set forth in this chapter shall prevail over the definition set forth in this chapter.

[0019] As used herein, the singular forms "a," "an," and "the" encompass multiple referents unless the context clearly indicates otherwise. For example, a reference to "a peptide" encompasses one or more peptides or mixtures of peptides. Also as used herein, unless specifically stated or evident from the context, the term "or" is understood to be inclusive and encompasses both "or" and "and."

[0020] As used herein, the term “about” refers to the normal range of error for each value, which is readily known to those skilled in the art. References to “about” values ​​or parameters herein encompass (and describe) embodiments relating to that value or parameter itself. For example, a statement referring to “about X” encompasses a statement relating to “X.”

[0021] As used herein, the term "antibody" is used in its broadest sense and includes polyclonal and monoclonal antibodies, including intact antibodies, as well as antigen-binding fragments (Fab) fragments, F(ab')2 fragments, Fab' fragments, Fv fragments, recombinant IgG (rIgG) fragments, single-chain antibody fragments including single-chain variable fragments (scFv), and functional (antigen-binding) antibody fragments including single-domain antibody (e.g., sdAb, sdFv, nanobody) fragments. This term includes immunoglobulins in modified forms, such as intrabodies, peptibodies, chimeric antibodies, fully human antibodies, humanized antibodies, and heteroconjugate antibodies, multispecific, e.g., bispecific, antibodies, diabodies, triabodies, and tetra-bodies, tandem di-scFv, tandem tri-scFv, etc., which are modified by genetic engineering and / or other methods. Unless otherwise specified, it should be understood that the term "antibody" includes its functional antibody fragments. This term also includes intact or full-length antibodies, including antibodies of any class or subclass, including IgG and its subclasses, IgM, IgE, IgA, and IgD.

[0022] "Individual" or "subject" includes mammals. Examples of mammals include, but are not limited to, livestock (e.g., cows, sheep, cats, dogs, and horses), primates (e.g., non-human primates such as humans and monkeys), rabbits, and rodents (e.g., mice and rats). "Individual" or "subject" can include birds such as chickens, vertebrates such as fish, and mammals such as mice, rats, rabbits, cats, dogs, pigs, cows, male cows, sheep, goats, horses, monkeys, and other non-human primates. In certain embodiments, the individual or subject is human.

[0023] As used herein, the term “sample” refers to any material that may contain the analyte of choice in an analyte assay. As used herein, “sample” may be a solution, suspension, liquid, powder, paste, aqueous, non-aqueous, or any combination thereof. A sample may be a biological sample, such as a biological fluid or biological tissue. Examples of biological fluids include urine, blood, plasma, serum, saliva, semen, feces, sputum, cerebrospinal fluid, tears, mucus, and amniotic fluid. Biological tissue is usually an aggregate of certain types of cells, along with intercellular material that forms one of the structural materials of human, animal, plant, bacterial, fungal, or viral structures, including connective tissue, epithelial tissue, muscle, and nerve tissue. Examples of biological tissue also include organs, tumors, lymph nodes, arteries, and individual cells.

[0024] In some embodiments, the sample is a biological sample. Biological samples as used herein include samples in the form of solutions, suspensions, liquids, powders, pastes, aqueous samples, or non-aqueous samples. As used herein, “biological sample” includes any sample obtained from a living organism or a viral (or prion) source or other source of macromolecules and biomolecules, and includes any cell type or tissue from which nucleic acids, proteins, and / or other macromolecules can be obtained. A biological sample may be a sample obtained directly from a biological source or a processed sample. For example, amplified isolated nucleic acids constitute a biological sample. Examples of biological samples include, but are not limited to, body fluids such as blood, plasma, serum, cerebrospinal fluid, synovial fluid, urine, and sweat, tissue and organ samples of animal and plant origin, and processed samples obtained therefrom. In some embodiments, the sample may be derived from tissue or body fluid, for example, tissue selected from the group consisting of connective tissue, epithelial tissue, muscle tissue, or nerve tissue, brain, lungs, liver, spleen, bone marrow, thymus, heart, lymph, blood, bone, cartilage, pancreas, kidneys, gallbladder, stomach, intestines, testes, ovaries, uterus, rectum, nervous system, glands, and internal blood vessels, or body fluid selected from the group consisting of blood, urine, saliva, bone marrow, sperm, ascites, and fractions thereof, such as serum or plasma.

[0025] The terms “level” or “levels” are used to refer to the presence and / or quantity of a target, for example, a substance or organism that is part of the pathogenesis of a disease or disorder and can be determined qualitatively or quantitatively. A “qualitative” change at a target level refers to the appearance or disappearance of a target that is undetectable or present in a sample obtained from a normal control. A “quantitative” change at the level of one or more targets refers to a measurable increase or decrease in the target level compared to a healthy control.

[0026] As used herein, the term “polymer” encompasses large molecules composed of smaller subunits. Examples of polymers include, but are not limited to, peptides, polypeptides, proteins, nucleic acids, carbohydrates, lipids, macrocyclic molecules, or combinations or complexes thereof. Polymers also include chimeric polymers (e.g., peptides bound to nucleic acids) composed of a combination of two or more types of polymers covalently bonded together. Polymers may also include “polymer aggregates” composed of non-covalent complexes of two or more polymers. Polymer aggregates may consist of polymers of the same type (e.g., protein-protein) or two or more different types of polymers (e.g., protein-DNA).

[0027] As used herein, the term “polypeptide” encompasses peptides and proteins and refers to molecules containing chains of two or more amino acids joined by peptide bonds. In some embodiments, polypeptides contain 2 to 50 amino acids, for example, having more than 20 to 30 amino acids. In some embodiments, peptides do not contain secondary, tertiary, or higher-order structures. In some embodiments, polypeptides are proteins. In some embodiments, proteins contain 30 or more amino acids, for example, having more than 50 amino acids. In some embodiments, proteins include secondary, tertiary, or higher-order structures in addition to the primary structure. The amino acids in polypeptides are most typically L-amino acids, but may also be D-amino acids, modified amino acids, amino acid analogs, amino acid mimes, or any combination thereof. Polypeptides may be naturally occurring, synthetically produced, or recombinantly expressed. Polypeptides may be synthetically produced, isolated, recombinantly expressed, or produced by a combination of the methodologies described above. Polypeptides may also contain additional groups that modify the amino acid chain, for example, functional groups added via post-translational modification. The polymer may be linear or branched, may contain modified amino acids, or may be interrupted by non-amino acids. This term also encompasses naturally modified or intervened amino acid polymers, and any other operations or modifications such as disulfide bond formation, glycosylation, lipidation, acetylation, phosphorylation, or conjugation with labeling components.

[0028] As used herein, the term “amino acid” refers to an organic compound comprising an amine group, a carboxylic acid group, and a side chain specific to each amino acid, which functions as a monomeric subunit of a peptide. Examples of amino acids include 20 standard, natural, or canonical amino acids, as well as non-standard amino acids. Standard natural amino acids include alanine (A or Ala), cysteine ​​(C or Cys), aspartic acid (D or Asp), glutamic acid (E or Glu), phenylalanine (F or Phe), glycine (G or Gly), histidine (H or His), isoleucine (I or Ile), lysine (K or Lys), leucine (L or Leu), methionine (M or Met), asparagine (N or Asn), proline (P or Pro), glutamine (Q or Gln), arginine (R or Arg), serine (S or Ser), threonine (T or Thr), valine (V or Val), tryptophan (W or Trp), and tyrosine (Y or Tyr). Amino acids can be L-amino acids or D-amino acids. Non-standard amino acids may be naturally occurring or chemically synthesized modified amino acids, amino acid analogs, amino acid mimes, non-standard proteinogenic amino acids, or non-proteinogenic amino acids. Examples of non-standard amino acids include, but are not limited to, selenocysteine, pyrrolicin, and N-formylmethionine, β-amino acids, homoamino acids, proline and pyruvate derivatives, 3-substituted alanine derivatives, glycine derivatives, ring-substituted phenylalanine and tyrosine derivatives, linear core amino acids, and N-methyl amino acids.

[0029] As used herein, the term “post-translational modification” refers to modifications that occur on a peptide after its translation, for example, ribosome translation, is complete. Post-translational modifications may be covalent chemical modifications or enzymatic modifications. Examples of post-translational modifications include acylation, acetylation, alkylation (including methylation), biotinylation, butyrylation, carbamylation, carbonylation, deamidation, deiminoation, diphthamide formation, disulfide crosslinking, eliminylation, flavin attachment, formylation, gammacarboxylation, glutamylation, glycylation, glycosylation, glycieation, heme C attachment, hydroxylation, hypsin formation, Post-translational modifications include, but are not limited to, iodization, isoprenylation, lipidation, lipoylation, malonylation, methylation, myristolylation, oxidation, palmitoylation, pegylation, phosphopantetheination, phosphorylation, prenylation, propionylation, retinilidenci base formation, S-glutathione addition, S-nitrosylation, S-sulfenylation, selenization, succinylation, sulfinization, ubiquitination, and C-terminal amidation. Post-translational modifications include modifications of the amino and / or carboxyl terminals of peptides. Modifications of terminal amino groups include, but are not limited to, desamino, N-lower alkyl, N-di-lower alkyl, and N-acyl modifications. Modifications of terminal carboxyl groups include, but are not limited to, amide, lower alkylamide, dialkylamide, and lower alkyl ester modifications (for example, lower alkyls are C1-C4 alkyls). Post-translational modifications also include modifications of amino acids between the amino and carboxyl terminals (e.g., those mentioned above, but not limited to these). The term post-translational modification may also include peptide modifications that include one or more detectable labels.

[0030] As used herein, the term “binding agent” refers to a binding target, such as a nucleic acid molecule, peptide, polypeptide, protein, carbohydrate, or small molecule that binds to, associates with, integrates with, recognizes, or combines with a polypeptide or a component or mechanism of a polypeptide. A binding agent may form a covalent or non-covalent association with a polypeptide or a component or mechanism of a polypeptide. A binding agent may also be a chimeric binding agent composed of two or more types of molecules, such as a nucleic acid molecule-peptide chimeric binding agent or a carbohydrate-peptide chimeric binding agent. A binding agent may be a naturally occurring, synthetically produced, or recombinantly expressed molecule. A binding agent may bind to a single monomer or subunit of a polypeptide (e.g., a single amino acid of a polypeptide) or to multiple linked subunits of a polypeptide (e.g., a dipeptide, tripeptide, or higher-order peptide of a longer peptide, polypeptide, or protein molecule). A binding agent may bind to a linear molecule or a molecule having a three-dimensional structure (also called a conformation). For example, antibody binders may bind to linear peptides, polypeptides, or proteins, or to conformational peptides, polypeptides, or proteins. Binders may bind to the N-terminal peptide, C-terminal peptide, or intervening peptide of peptide, polypeptide, or protein molecules. Binders may bind to the N-terminal amino acid, C-terminal amino acid, or intervening amino acid of peptide molecules. Binders may preferably bind to chemically modified or labeled amino acids (e.g., amino acids labeled with chemical reagents) rather than to unmodified or unlabeled amino acids. For example, binders may preferably bind to labeled or modified amino acids rather than to unlabeled or unmodified amino acids. Binders may bind to post-translational modifications of peptide molecules. Binders may exhibit selective binding to components or mechanisms of polypeptides (e.g., a binder may selectively bind to one of 20 possible native amino acid residues, while binding to the other 19 native amino acid residues with very low affinity or not at all).If the binder is capable of binding to or configured to bind to multiple components or mechanisms of a polypeptide, it may exhibit low selective binding (for example, the binder may bind to two or more different amino acid residues with similar affinity). The binder may include a coding tag that can be attached to the binder by a linker.

[0031] As used herein, the term “linker” refers to one or more nucleotides, nucleotide analogs, amino acids, peptides, polypeptides, polymers, or non-nucleotide chemical moieties used to join two molecules. Linkers are used, for example, to join a binder to a coding tag, a recording tag to a polypeptide, a polypeptide to a support, and a recording tag to a solid support. In certain embodiments, the linker joins the two molecules by an enzymatic reaction or a chemical reaction (e.g., click chemistry).

[0032] As used herein, the term “ligand” refers to any molecule or portion attached to a compound described herein. A “ligand” may also refer to one or more ligands attached to a compound. In some embodiments, the ligand is a pendant group or a binding site (e.g., a site to which a binder is attached).

[0033] As used herein, the term “proteome” may include the entire set of proteins, polypeptides, or peptides (including their conjugates or complexes) expressed by the genome, cells, tissues, or organism of any organism at a given time. In one embodiment, the proteome is the set of proteins expressed in a given type of cell or organism at a given time under specified conditions. Proteomics is the study of the proteome. For example, the “cellular proteome” may include the collection of proteins found in a particular cell type under a particular set of environmental conditions, such as exposure to hormonal stimulation. The complete proteome of an organism may include the complete set of proteins from all of the various cellular proteomes. The proteome may also include the collection of proteins in a biological system at a particular cellular level. For example, all the proteins within a virus may be called the viral proteome. As used herein, the term “proteome” includes, but is not limited to, a subset of the proteome defined by quinomes; secretomes; receptomes (e.g., GPCRomes); immunoproteomes; nutriprotomes; a subset of the proteome defined by post-translational modifications (e.g., phosphorylation, ubiquitination, methylation, acetylation, glycosylation, oxidation, lipidation, and / or nitrosylation), such as phosphoproteomes (e.g., phosphotyrosine-proteomes, tyrosine-quinomes, and tyrosine-phosphatoms), glycoproteomes, etc.; a subset of the proteome associated with a tissue or organ, developmental stage, or physiological or pathological condition; a subset of the proteome associated with cellular processes such as the cell cycle, differentiation (or dedifferentiation), cell death, senescence, cell migration, transformation, or metastasis; or any combination thereof. As used herein, the term “proteomics” refers to the quantitative analysis of the proteome in cells, tissues, and body fluids, as well as the corresponding spatial distribution of the proteome within cells and tissues. Furthermore, proteomics research includes the dynamic state of the proteome, which continuously changes over time as a function of biological and defined biological or chemical stimuli.

[0034] The terminal amino acid at one end of a peptide or polypeptide chain having a free amino group is referred to herein as the “N-terminal amino acid” (NTAA). The terminal amino acid at the other end of a chain having a free carboxyl group is referred to herein as the “C-terminal amino acid” (CTAA). If the length of a peptide is n amino acids, the amino acids constituting the peptide may be numbered sequentially. As used herein, NTAA is n 目 This is considered the nth amino acid (also referred to herein as "nNTAA"). Using this nomenclature, the next amino acids are n-1 amino acids, then n-2 amino acids, and so on, descending the length from the N-terminus to the C-terminus of the peptide. In certain embodiments, NTAA, CTAA, or both may be modified or labeled with a partial or chemical moiety.

[0035] As used herein, the term “barcode” refers to a nucleic acid molecule of about 2 to about 30 bases (e.g., 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, or 30 bases) that provides a unique identifier tag or origin information for a polypeptide, a binder, a set of binders from a binding cycle, a sample polypeptide, a set of samples, a polypeptide in a compartment (e.g., a droplet, a bead, or a separated location), a polypeptide in a set of compartments, a polypeptide fraction, a set of polypeptide fractions, a spatial region or set of spatial regions, a library of polypeptides, or a library of binders. A barcode may be an artificial sequence or a naturally occurring sequence. In a particular embodiment, each barcode within a collection of barcodes is different. In other embodiments, a portion of the barcodes in a group of barcodes may differ, for example, by at least about 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 97%, or 99% of the barcodes in a group of barcodes. The group of barcodes may be generated randomly or not. In certain embodiments, the group of barcodes may be error-corrected or error-tolerant barcodes. Barcodes can be used to computationally deconvolve multiplexed sequencing data to identify sequence reads originating from individual polypeptides, samples, libraries, etc. Barcodes can also be used to deconvolve polypeptide assemblies that have been distributed into smaller compartments to enhance mapping. For example, instead of remapping peptides to a proteome, peptides can be remapping to their original protein molecules or protein complexes.

[0036] As used herein, the term “coding tag” refers to a nucleic acid molecule of approximately 2 to approximately 100 bases, containing any integer between 2 and 100 (including 2 and 100), and having any suitable length, including identification information relating to its associated binder. “Coding tags” can also be fabricated from “sequencing polymers” (e.g., Niu et al., 2013, Nat. Chem. 5:282-292, Roy et al., 2015, Nat. Commun. 6:7237, Lutz). See Macromolecules 48:4759-4767, 2015. (Each of these is incorporated by reference in whole.) The coding tag may include an encoder sequence, which may optionally have one spacer adjacent to one side, or optionally have spacers adjacent to both sides. The coding tag may also consist of an optional UMI and / or any binding cycle-specific barcode. The coding tag may be single-stranded or double-stranded. A double-stranded coding tag may include a blunt end, a protruding end, or both. The coding tag may refer to a coding tag directly attached to the binder, a complementary sequence hybridized to a coding tag directly attached to the binder (e.g., in the case of a double-stranded coding tag), or coding tag information present in the extension record tag. In certain embodiments, the coding tag may further include a binding cycle-specific spacer or barcode, a unique molecular identifier, a universal priming site, or any combination thereof.

[0037] As used herein, the term “spacer” (Sp) refers to a nucleic acid molecule located on the end of a recording tag or coding tag, having a length of approximately 1 to 20 bases (e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 bases). In certain embodiments, the spacer sequence is adjacent to the encoder sequence of the coding tag at one or both ends. Following the binding of the binder to the polypeptide, annealing between the complementary spacer sequences on their respective coding and recording tags enables the transfer of binding information via a primer extension reaction or ligation to the recording tag, coding tag, or jitag construct, respectively. “Sp’” refers to a spacer sequence complementary to Sp. Preferably, spacer sequences in the binder library have the same number of bases. Common (shared or identical) spacers can be used in the binder library. Spacer sequences may have “cycle-specific” sequences to track the binder used in a particular binding cycle. Spacer sequences (Sp) may be constant across all binding cycles, specific to a particular class of polypeptide, or specific to the number of binding cycles. Polypeptide class-specific spacers allow the coding tag information of homogeneous binders present in the extension record tag from a completed binding / extension cycle to anneal via the class-specific spacer to the coding tag of another binder that recognizes the same class of polypeptide in a subsequent binding cycle. Only through the sequential binding of precise homogeneous pairs will interacting spacer elements and effective primer extension be achieved. Spacer sequences may contain enough bases to anneal to complementary spacer sequences in the record tag to initiate a primer extension (also known as polymerase extension) reaction, or to provide a “splint” for a ligation reaction, or to mediate a “sticky end” ligation reaction. Spacer sequences may contain fewer bases than encoder sequences in the coding tag.

[0038] As used herein, the term “recording tag” refers to a portion, such as a chemical coupling portion, a nucleic acid molecule, or a sequenceable polymer molecule (see, e.g., Niu et al., 2013, Nat. Chem. 5:282-292, Roy et al., 2015, Nat. Commun. 6:7237, Lutz, 2015, Macromolecules 48:4759-4767; each of these is incorporated by reference as a whole), to which identification information of a coding tag may be transferred, or from there identification information about the polymer associated with the recording tag (e.g., UMI information) may be transferred to the coding tag. Identification information may include any information that characterizes the molecule, such as information relating to the sample, fraction, partition, spatial position, interacting neighboring molecules, cycle number, etc. Furthermore, the presence of UMI information can also be classified as identification information. In certain embodiments, after the binder has bound to the polypeptide, information from the coding tag linked to the binder can be transferred to the recording tag associated with the polypeptide while the binder is still bound to the polypeptide. In other embodiments, after the binder has bound to the polypeptide, information from the recording tag associated with the polypeptide can be transferred to a coding tag linked to the binder while the binder is bound to the polypeptide. The recording tag may be linked directly to the polypeptide, linked to the polypeptide via a polyfunctional linker, or associated with the polypeptide by its proximity (or colocalization) on a support. The recording tag may be linked via its 5' or 3' end, or at an internal site, insofar as its linkage is compatible with the method used to transfer coding tag information to or from the recording tag. The recording tag may further include other functional components, such as a universal priming site, a unique molecular identifier, a barcode (e.g., a sample barcode, fraction barcode, spatial barcode, compartment tag, etc.), a spacer sequence complementary to the spacer sequence of the coding tag, or any combination thereof.The spacer array of the recording tag is preferably located at the 3' end of the recording tag in embodiments in which coding tag information is transferred to the recording tag using polymerase elongation.

[0039] As used herein, the term "primer elongation," also referred to as "polymerase elongation," refers to a reaction catalyzed by a nucleic acid polymerase (e.g., DNA polymerase) in which a nucleic acid molecule (e.g., oligonucleotide primer, spacer sequence) that anneals to a complementary strand is elongated by the polymerase using the complementary strand as a template.

[0040] As used herein, the terms “Unique Molecular Identifier” or “UMI” refer to nucleic acid molecules of approximately 3 to approximately 40 bases (3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, or 40 bases) that provide a unique identifier tag to each polymer, polypeptide, or binder to which the UMI is linked. Polypeptide UMIs can be used to computationally deconvolve sequencing data from multiple extension record tags to identify the extension record tags derived from individual polypeptides. Polypeptide UMIs can be used to accurately count the original polypeptide molecules by folding NGS reads into unique UMIs. Binder UMIs can be used to identify each individual molecular binder that binds to a particular polypeptide. For example, UMI can be used to identify the number of individual binding events of a single-amino acid-specific binder that occur for a particular peptide molecule. When both UMI and barcodes are referenced in relation to a binder or polypeptide, it is understood that the barcode refers to identification information other than the UMI for the individual binder or polypeptide (e.g., sample barcode, compartment barcode, binding cycle barcode).

[0041] As used herein, the terms “universal priming site,” “universal primer,” or “universal priming sequence” refer to nucleic acid molecules that can be used in library amplification and / or sequencing reactions. Universal priming sites may include, but are not limited to, priming sites (primer sequences) for PCR amplification, flow cell adapter sequences that anneal to complementary oligonucleotides on the flow cell surface to enable bridge amplification in some next-generation sequencing platforms, sequencing priming sites, or combinations thereof. Universal priming sites can be used for other types of amplification, including those commonly used in combination with next-generation digital sequencing. For example, extension record tag molecules can be cyclically formed, and the universal priming site can be used in rolling circle amplification to form DNA nanoballs that can be used as sequencing templates (Drmanac et al., 2009, Science 327:78-81). Alternatively, the recording tag molecule can be cyclized and directly sequenced by polymerase elongation from a universal priming site (Korlach et al., 2008, Proc. Natl. Acad. Sci. 105: 1176-1181). The term "forward" as used in relation to "universal priming site" or "universal primer" may also be referred to as "5" or "sense". The term "reverse" as used in relation to "universal priming site" or "universal primer" may also be referred to as "3" or "antisense".

[0042] As used herein, the term “extension record tag” refers to a record tag to which information of at least one binding agent coding tag (or its complementary sequence) has been transferred following the binding of the binding agent to the polypeptide. The coding tag information may be transferred to the record tag directly (e.g., by ligation) or indirectly (e.g., by primer extension). The coding tag information may be transferred to the record tag enzymatically or chemically. The extension record tag may contain binder information for 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 125, 150, 175, and 200 or more coding tags. The nucleotide sequences of the extension record tag may reflect the temporal and sequential order of binder binding identified by those coding tags, the sequential order of some of the binder binding identified by the coding tags, or may not reflect any order of binder binding identified by the coding tags. In certain embodiments, the coding tag information present in the extension record tag represents the polypeptide sequence being analyzed with at least 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identity. In certain embodiments where the extension record tag does not represent the polypeptide sequence being analyzed with 100% identity, the error may be due to off-target binding by the binder, or "failure" of the binding cycle (e.g., failure of the primer extension reaction due to the binder being unable to bind to the polypeptide during the binding cycle), or both.

[0043] As used herein, the terms “solid support,” “solid surface,” “solid substrate,” “sequencing substrate,” or “substrate” refer to any solid material, including porous and non-porous materials, to which polypeptides can be directly or indirectly associated by any means known in the Art, including covalent and non-covalent interactions, or any combination thereof. A solid support may be two-dimensional (e.g., planar) or three-dimensional (e.g., gel matrix or beads). A solid support may be any support surface, including but not limited to beads, microbeads, arrays, glass surfaces, silicon surfaces, plastic surfaces, filters, membranes, PTFE membranes, PTFE films, nitrocellulose membranes, nitrocellulose-based polymer surfaces, nylon, silicon wafer chips, flow-through chips, flow cells, biochips including signal-converting electronics, channels, microtiter wells, ELISA plates, spin interference disks, nitrocellulose membranes, nitrocellulose-based polymer surfaces, polymer matrices, nanoparticles, or microspheres. Materials for the solid support include, but are not limited to, acrylamide, agarose, cellulose, dextran, nitrocellulose, glass, gold, quartz, polystyrene, polyethylene vinyl acetate, polypropylene, polyester, polymethacrylate, polyacrylate, polyethylene, polyethylene oxide, polysilicate, polycarbonate, polyvinyl alcohol (PVA), Teflon®, fluorocarbon, nylon, silicone rubber, polyanhydride, polyglycolic acid, polyvinyl chloride, polylactic acid, polyorthoester, functionalized silane, polypropyl fumerate, collagen, glycosaminoglycan, polyamino acid, dextran, or any combination thereof. The solid support further includes molded polymers such as thin films, membranes, bottles, dishes, fibers, woven fibers, and tubes, particles, beads, microspheres, fine particles, or any combination thereof.For example, if the solid surface is beads, the beads may include, but are not limited to, ceramic beads, polystyrene beads, polymer beads, polyacrylate beads, methylstyrene beads, agarose beads, cellulose beads, dextran beads, acrylamide beads, solid core beads, porous beads, paramagnetic beads, glass beads, controlled porous beads, silica-based beads, or any combination thereof. The beads may be spherical or irregular in shape. The beads or support may be porous. The size of the beads may range from nanometers, e.g., 100 nm to several millimeters, e.g., 1 mm. In certain embodiments, the size of the beads is in the range of about 0.2 microns to about 200 microns, or about 0.5 microns to about 5 microns. In some embodiments, the beads may have a diameter of about 1, 1.5, 2, 2.5, 2.8, 3, 3.5, 4, 4.5, 5, 5.5, 6, 6.5, 7, 7.5, 8, 8.5, 9, 9.5, 10, 10.5, 15, or 20 μm. In certain embodiments, the “beads” solid support may refer to individual beads or a group of beads. In some embodiments, the solid surface is nanoparticles. In certain embodiments, the size of the nanoparticles is in the range of diameters from about 1 nm to about 500 nm, for example, between about 1 nm and about 20 nm, between about 1 nm and about 50 nm, between about 1 nm and about 100 nm, between about 10 nm and about 50 nm, between about 10 nm and about 100 nm, between about 10 nm and about 200 nm, between about 50 nm and about 100 nm, between about 50 nm and about 150 nm, between about 50 nm and about 200 nm, between about 100 nm and about 200 nm, or between about 200 nm and about 500 nm. In some embodiments, the nanoparticles may have diameters of about 10 nm, about 50 nm, about 100 nm, about 150 nm, about 200 nm, about 300 nm, or about 500 nm. In some embodiments, the nanoparticles have a diameter of less than approximately 200 nm.

[0044] As used herein, the terms “nucleic acid molecule” or “polynucleotide” refer to single- or double-stranded polynucleotides and polynucleotide analogs containing deoxyribonucleotides or ribonucleotides linked by 3'-5' phosphodiester bonds. Examples of nucleic acid molecules include, but are not limited to, DNA, RNA, and cDNA. Polynucleotide analogs may have a backbone other than the standard phosphodiester bond found in natural polynucleotides, and optionally, a modified sugar moiety other than ribose or deoxyribose. Polynucleotide analogs contain bases capable of hydrogen bonding to standard polynucleotide bases via Watson-Crick base pairing, and the analog backbone presents the bases in a manner that enables such hydrogen bonding between the oligonucleotide analog molecule and the bases of standard polynucleotides in a sequence-specific manner. Examples of polynucleotide analogs include, but are not limited to, heteronucleotides (XNAs), cross-linked nucleic acids (BNAs), glycol nucleic acids (GNAs), peptide nucleic acids (PNAs), γPNAs, morpholinopolynucleotides, locked nucleic acids (LNAs), threose nucleic acids (TNAs), 2'-O-methyl polynucleotides, 2'-O-alkylribosyl-substituted polynucleotides, phosphorothioate polynucleotides, and boronophosphate polynucleotides. Polynucleotide analogs may include, for example, 7-deazapurine analogs, 8-halopurine analogs, 5-halopyrimidine analogs, or universal base analogs that can pair with any base, including hypoxanthine, nitroazole, isocarbostyryl analogs, azolecarboxamides, and aromatic triazole analogs, or base analogs having additional functionalities such as a biotin moiety for affinity binding. In some embodiments, the nucleic acid molecule or oligonucleotide is a modified oligonucleotide. In some embodiments, the nucleic acid molecule or oligonucleotide is DNA with pseudocomplementary bases, DNA with protected bases, RNA molecule, BNA molecule, XNA molecule, LNA molecule, PNA molecule, γPNA molecule, or morpholino DNA, or a combination thereof.In some embodiments, nucleic acid molecules or oligonucleotides are modified by backbone modification, sugar modification, or nucleic acid base modification. In some embodiments, nucleic acid molecules or oligonucleotides have nucleic acid base protecting groups such as Alloc, electrophilic protecting groups such as tilane, acetyl protecting groups, nitrobenzyl protecting groups, sulfonate protecting groups, or conventional base instability protecting groups.

[0045] As used herein, “nucleic acid sequencing” means determining the order of nucleotides in a nucleic acid molecule or a sample of nucleic acid molecules.

[0046] As used herein, “next-generation sequencing” refers to high-throughput sequencing methods that enable the parallel sequencing of millions to billions of molecules. Examples of next-generation sequencing methods include synthetic sequencing, ligation sequencing, hybridization sequencing, polony sequencing, ion-semiconductor sequencing, and pyrosequencing. By attaching primers to a solid substrate and complementary sequences to nucleic acid molecules, nucleic acid molecules can be hybridized to the solid substrate via the primers, and then amplified using polymerase to generate multiple copies in separate regions on the solid substrate (these groupings may also be called polymerase colonies or polonys). As a result, nucleotides at a particular position can be sequenced multiple times (e.g., hundreds or thousands of times) during the sequencing process. This depth of coverage is referred to as “deep sequencing.” Examples of high-throughput nucleic acid sequencing technologies include platforms offered by Illumina, BGI, Qiagen, Thermo-Fisher, and Roche, and include forms such as parallel bead arrays, synthesis sequencing, ligation sequencing, capillary electrophoresis, electronic microchips, "biochips," microarrays, parallel microchips, and single-molecule arrays (see, e.g., Service, Science (2006) 311:1544–1546).

[0047] As used herein, “single-molecule sequencing” or “third-generation sequencing” refers to a next-generation sequencing method in which a read from a single-molecule sequencing instrument is produced by sequencing a single molecule of DNA. Unlike next-generation sequencing methods that rely on amplification to clone a large number of DNA molecules in parallel for sequencing in a stepwise approach, single-molecule sequencing interrogates a single molecule of DNA and does not require amplification or synchronization. Single-molecule sequencing includes methods that require pausing the sequencing reaction after the incorporation of each base (“wash-and-scan” cycle) and methods that do not require pausing between read steps. Examples of single-molecule sequencing methods include single-molecule real-time sequencing (Pacific Biosciences), nanopore-based sequencing (Oxford Nanopore), duplex-interrupted nanopore sequencing, and direct imaging of DNA using advanced microscopy.

[0048] As used herein, “analyzing” a polypeptide means identifying, detecting, quantifying, characterizing, distinguishing all or part of the components of the polypeptide, or a combination thereof. For example, the analysis of a peptide, polypeptide, or protein includes determining all or part of the amino acid sequence (continuous or discontinuous) of the peptide. The analysis of a polypeptide also includes the partial identification of the components of the polypeptide. For example, the partial identification of amino acids in a polypeptide protein sequence can be done by identifying amino acids in the protein as belonging to a subset of possible amino acids. The analysis usually begins with the analysis of nNTAAs and proceeds to the next amino acids of the peptide (i.e., n-1, n-2, n-3, etc.). This is achieved by eliminating the nNTAAs, thereby converting the n-1 amino acids of the peptide to the N-terminal amino acid (referred to herein as “n-1 NTAA”). The analysis of a peptide may also include the determination of the presence and frequency of post-translational modifications of the peptide. This may or may not include information regarding the order of post-translational modifications of the peptide. The analysis of a peptide may also include the determination of the presence and frequency of epitopes in the peptide. This may or may not include information regarding the order or position of epitopes within the peptide. Peptide analysis may involve a combination of different types of analysis, such as obtaining epitope information, amino acid sequence information, post-translational modification information, or any combination thereof.

[0049] The aspects and embodiments of the present invention described herein are understood to encompass aspects and embodiments "consisting of" and / or aspects and embodiments "consisting essentially of".

[0050] Through this disclosure, various aspects of the present invention are presented in range form. It should be understood that the range form is merely for convenience and conciseness and should not be interpreted as an immutable limitation to the scope of the invention. Therefore, range descriptions should be considered to specifically disclose all possible subranges and the individual numbers within those ranges. For example, a range description such as 1-6 should be considered to specifically disclose subranges such as 1-3, 1-4, 1-5, 2-4, 2-6, 3-6, and the individual numbers within those ranges (e.g., 1, 2, 3, 4, 5, and 6). This applies regardless of the width of the range.

[0051] Other objects, advantages, and features of the present invention will become apparent from the following specification, linked with the accompanying drawings.

[0052] I. Continuous Encoding Provided herein are methods and kits for the analysis of polymers, such as peptides, polypeptides, and proteins, which include the step of transferring information to a recording tag. The analysis uses nucleic acid encoding of molecular recognition events. In some embodiments, the information to be transferred includes identification information relating to a binder configured to bind to the polymer. The method provided for information transfer includes (a) providing a polymer and an associated recording tag conjugated to a support; (b) contacting the polymer with a binder that can bind to the polymer and includes a coding tag having identification information relating to the binder, thereby enabling binding between the polymer and the binder; (c) conjugating the 5' end of the recording tag to the 3' end of the coding tag with a nucleic acid conjugation reagent; (d) extending the recording tag using the coding tag as a template with a polymerase to produce a double-stranded extended recording tag; and (e) cleaving the double-stranded extended recording tag with a double-stranded nucleic acid cleavage reagent to produce a 3' overhang within the extended recording tag. By performing these steps, information is transferred from the coding tag to the recording tag to produce the extended recording tag. In some embodiments, the polymerase, nucleic acid conjugation reagent, and double-strand nucleic acid cleavage reagent are provided in a mixture or simultaneously. In some embodiments, steps (c), (d), and (e) are performed as a one-pot reaction. In some other embodiments, steps (c), (d), and (e) are performed sequentially and separately. In some cases, each reagent can be provided separately, or two of the three enzyme reagents can be provided simultaneously. In some cases, the recording tag and coding tag of the provided method contain nucleic acids.

[0053] One or more cycles of steps (b) to (e) can be repeated to generate an extended record tag containing information transferred from multiple coding tags. For example, steps (b), (c), (d), and (e) can be repeated one or more times consecutively in a periodic manner. In some embodiments, the generated extended record tag includes a nucleic acid hairpin. The methods provided herein may include providing multiple binders and multiple polymers, enabling the binders and polymers to interact. In some embodiments, the multiple binders are provided as a mixture in each cycle. In some embodiments, the method includes contacting a single polymer with a single binder, contacting multiple polymers with a single binder, or contacting multiple polymers with multiple binders.

[0054] In some embodiments, this disclosure provides methods for analyzing polymers, including information transfer, which in part involve direct application to the characterization, quantification, and / or sequencing of proteins and peptides. In some specific embodiments, the polymer for analysis is not a nucleic acid. In some specific embodiments, the binder is not a nucleic acid. Provided herein are methods for transferring information from a coding tag associated with or conjugated to a binder to a recording tag associated with the polymer (e.g., polypeptide) to be analyzed. The information transfer is carried out using a three-part reaction including ligation, extension, and cleavage with a double-strand nucleic acid cleavage reagent (e.g., restriction enzyme). The information transferred from the coding tag includes identification information relating to the identity of the binder, thereby providing information about the polymer or a portion thereof conjugated by the binder. For example, if a protein / polypeptide / peptide polymer is conjugated by the binder, the identification information may include information about the identity of one or more amino acids conjugated by the binder. In some embodiments, the information about the identity of the polymer (or a portion thereof) conjugated by the binder is from a coding tag associated with the binder and is transferred to a recording tag. A polymer analysis assay may involve one or more cycles of transferring binder identification information from a coding tag to a record tag associated with the polymer being analyzed. The final extension record tag associated with the polymer for analysis may contain information from one or more coding tags. If multiple cycles are performed, the resulting extension record tag will contain information constructed from a series of binding events and multiple information transfer events from the coding tags. In general, improvements for information transfer using the described methods, including the activity of nucleic acid conjugation reagents, polymerases, and double-strand nucleic acid cleavage reagents, may offer specific advantages to polymer analysis assays.

[0055] In some embodiments, the sequential encoding systems used in methods provided for analyzing polymers offer certain advantages in the overall design of the assay. In particular, several advantages are provided by the sequential steps performed by the nucleic acid conjugation reagent, polymerase, and double-strand nucleic acid cleavage reagent. In some embodiments, the method includes or uses a reaction system in which a mixture of enzymes is supplied to the reaction, resulting in it being performed in one step and / or the reaction being performed as a one-pot reaction. The design of the system allows the ligation, extension, and cleavage in steps (c), (d), and (e) to occur in a stepwise or sequential manner, respectively. In some other embodiments, the stepwise nature of the steps can be introduced by the use of blocking groups or by introducing additional requirements such that the ligation, extension, and / or cleavage are performed in a manner conditioned on the completion of the previous step. The activity of the polymerase, nucleic acid conjugation reagent, and double-strand nucleic acid cleavage reagent is provided under favorable conditions for transferring information from the coding tag to the recording tag to generate an extension recording tag. Some advantages of using the described method include high success in information transfer (encoding), a simple design for stepwise reactions, the option to run in a single step / as a single pot reaction, reduced need for spacers or shortened spacer length, minimizing DNA-DNA-protein interactions within the system and / or minimizing DNA-protein interactions within the system. For example, compared to single-stranded nucleic acid molecules that are plastic and can provide bases exposed for interaction, the double-stranded recording tags provided herein may show less interaction for DNA to interact with other components in the system (e.g., nucleic acids, peptides, binders, etc.).

[0056] In some embodiments, methods provided for transferring information using sequential encoding enable specific advantages related to the format of the nucleic acid components used. For example, in extension-based information transfer methods that rely on spacer elements to form polymerase priming sites, the size of the spacers may need to be a specific length, such as at least 6 base pairs. Shorter spacers may have advantages such as avoiding nonspecific interactions in the assay. In some cases, sequence enrichment is straighter if repeating elements are minimized or eliminated. In some cases, the methods provided avoid specific problems with specificity, bias, stability, and efficiency compared to extension-based methods. In ligation-based methods for information transfer, the efficiency of joining suitable ends of nucleic acids and the complexity with other necessary steps for transferring information can be problematic. In some embodiments, methods provided using polymerase, nucleic acid joining reagents, and double-strand nucleic acid cleavage reagents offer advantages over simply extension-based or ligation-based systems. By using a cleavage step that occurs after the polymerase step, the A-tailing problem, where the polymerase leaves an "A" overhang, can be avoided. The use of double-stranded DNA minimizes potential DNA-DNA interactions that may occur in other systems utilizing single-stranded DNA components. In some cases, shorter spacers that can be used in the provided method can reduce the length to, for example, 2 bp spacers, and in some cases, the use of spacers can be eliminated entirely. In some embodiments, the sequential encoding method self-terminates after one cycle of information transfer due to incompatible overhangs.

[0057] In some embodiments, both the recording tag and the coding tag include spacer sequences. For example, the spacer is a nucleic acid molecule with 10 or fewer bases, 9 or fewer bases, 8 or fewer bases, 7 or fewer bases, 6 or fewer bases, 5 or fewer bases, 4 or fewer bases, 3 or fewer bases, or 2 or fewer bases. The spacer may be a cycle-specific spacer or a cycle-alternating spacer. A particular design of the system for transferring information may utilize spacers such that information is transferred from the second coding tag to the extended recording tag only if the spacer added by the previous coding tag matches at least a portion of the spacer on the second coding tag. In this way, subsequent encoding events are dependent on and conditional on the previous encoding event. In some other embodiments, neither the recording tag nor the coding includes spacer sequences.

[0058] In some embodiments, several steps within a cycle can be optionally repeated. For example, after one cycle of providing a binder and transferring information from a coding tag, the binder can be removed and then provided again to allow the binder to recombine, thereby allowing another opportunity for information transfer to occur, for example, if the first information transfer did not occur. In some embodiments, the method provided using alternating spacers is suitable for this repetition to allow a second opportunity for information transfer. If the first binding and information transfer event is successful, the spacer transferred with the successful coding tag information does not allow a repeated information transfer event to occur (even if the binder recognizes the polymer for analysis). In contrast, if the first information transfer event is unsuccessful, recombination results in a coding tag with a spacer matching the spacer on an existing recording tag, and the second binding allows an information transfer event to occur after the recombination step, thereby providing a mechanism to compensate for the first missed event.

[0059] The provided method utilizes polymerase, a nucleic acid ligation reagent, and a double-strand nucleic acid cleavage reagent. Any suitable enzyme for double-strand nucleic acid extension, ligation, and cleavage can be used (see, for example, Japanese Patent Publications CN104212791B, CN101560538A, and CN100510069C).

[0060] In the provided method, the transfer of information from a coding tag to a recording tag begins with a ligation step. The 5' end of the recording tag is ligated (e.g., ligated) to the 3' end of the coding tag by a nucleic acid ligating reagent. In some embodiments, after the recording tag is ligated to the coding tag by the nucleic acid ligating reagent, the binder does not need to remain bound or remain bound to the polymer.

[0061] In some examples, the nucleic acid conjugation reagent is either a chemical ligation reagent or an enzymatic ligation reagent. The ligation step may be performed by chemical ligation or enzymatic ligation (e.g., sticky-end ligation, single-stranded (ss) ligation such as ssDNA ligation, or any combination thereof). In some designs, ligation may be blunt-end ligation and / or sticky-end ligation. The provided sequential encoding method is designed to generate a double-stranded structure containing information transferred from a coding tag, and any conjugation method to achieve this objective can be applied. Examples of ligases include, but are not limited to, CV DNA ligase, Circ ligase, Circ ligase II, T4 DNA ligase, T7 DNA ligase, T3 DNA ligase, Taq DNA ligase, E. coli DNA ligase, and 9°N DNA ligase (see, for example, U.S. Patent Publication 2014 / 0378315A1 or U.S. Patent No. 10,494,671B2). In some preferred embodiments, the nucleic acid conjugation reagent refers to an enzyme (e.g., a ligase) that conjugates two DNA segments, such as T4 DNA ligase.

[0062] Following the conjugation or ligation step, an extension step may be performed, which may include a polymerase-mediated reaction (e.g., primer extension of single-stranded or double-stranded nucleic acid). Extension of the 3' end of the record tag (or an already extended record tag) occurs using the ligated coding tag as a template. The extension step causes the restriction enzyme / endonuclease site introduced from the coding tag and ligated to become double-stranded on the extended record tag. In some embodiments, the DNA polymerase used for primer extension has strand displacement activity and restricted or absent 3'-5 exonuclease activity. Some of the many examples of such polymerases include Krenow exo- (Krenow fragment of DNA Pol1), T4 DNA polymerase exo-, T7 DNA polymerase exo (Sequenase 2.0), Pfu exo-, Vent exo-, and Deep The exoskeletons include Vent exo-, Bst DNA polymerase large fragment exo-, Bca Pol, 9°N Pol, and Phi29 Pol exo-. In a preferred embodiment, the DNA polymerase is active at room temperature and up to 45°C. In another embodiment, a “warm-start” version of the thermophilic polymerase is used so that the polymerase is activated and used at approximately 40°C to 50°C. An exemplary warm-start polymerase is Bst2.0 warm-start DNA polymerase (New England Biolabs). Preferred conditions for the extension reaction, including any additives and buffers for the reaction, may also be provided.

[0063] Following the extension step, a double-stranded extension record tag containing a recognition sequence recognizable by the double-stranded nucleic acid cleavage reagent is generated. Cleavage of the double-stranded extension record tag generates a 3' overhang on the record tag. In some embodiments, the 3' overhang of the extension record tag generated by the double-stranded nucleic acid cleavage reagent is available for hybridization with a second coding tag when step (b) is repeated. In some embodiments, cleavage of the double-stranded extension record tag removes or detaches the binder from the extension record tag. In some specific cases, cleavage of the double-stranded record tag releases the binder from the extension record tag. In some cases, cleavage makes the binder available for release from the polymer being analyzed. The double-stranded nucleic acid cleavage reagent may be a restriction enzyme or restriction endonuclease. In some embodiments, the restriction enzyme is an IIS-type restriction enzyme. The advantages of using IIS-type restriction enzymes include fewer single-stranded nucleotides remaining after the cleavage reaction compared to classical palindromic restriction enzymes, as shown in Figure 1B and Example 2 (2nt spacer only). In some preferred embodiments, the sequence after cleavage by the cleavage reagent leaves a 3' overhang. For example, the cleavage reagent may be the enzyme BtsI, Mva1269I, BsaI, BsmBI, etc. IIS-type restriction enzymes can recognize base sequences containing 5'…GCAGTGNN…3' / 3'…CGTCACNN…5'. In some specific cases, the restriction enzyme is Nb.BtsI or BtsI-v2 or a derivative thereof.

[0064] In some other embodiments, the double-strand nucleic acid cleavage reagent is an IIP-type restriction enzyme that has palindromic specificity and cleaves within its recognition sequence. After ligating the coding tag to the recording tag, the coding tag sequence adjacent to the ligation site should be different from the restriction enzyme cleavage site to avoid potential reproducibility of the restriction enzyme cleavage site at the ligation site (see, e.g., Figure 1B, first step). In some preferred embodiments, the IIP-type restriction enzyme generates a 3' overhang on the recording tag after cleavage, as shown in Figure 1B. In some other embodiments, an IIP-type restriction enzyme that generates a blunt end may be employed for use in the disclosed method. In these embodiments, the 3' end of the coding tag should be ligated to the blunt end of the recording tag generated after restriction enzyme cleavage, and the coding tag sequence adjacent to the ligation site should be different from the restriction enzyme cleavage site.

[0065] In some other embodiments, a combination of enzymes or endonucleases can be used to achieve cleavage of both strands of the extension recording tag. For example, if a nickeling enzyme or endonuclease is used to cleave the first strand of the extension recoding tag, a secondary mechanism can be used to cleave the second strand. In some specific embodiments, the nickeling enzyme or nickeling endonuclease is not used. The appropriate restriction enzyme can be selected based on considerations of preferred cleavage sites, distance between digestion sites and recognition sites, precision of digestion sites, and ability to generate nucleic acid ends required for other steps of the method.

[0066] In some embodiments, the nucleic acid conjugation reagent and the polymerase are provided simultaneously. In some embodiments, the polymerase and the double-strand nucleic acid cleavage reagent are provided simultaneously.

[0067] In some of the embodiments provided, steps (a), (b), (c), (d), and (e) are performed sequentially. In certain embodiments, information about the binding event of a binder to a polymer (e.g., a peptide) is periodically transferred from the coding tag to a recording tag associated with the immobilized polymer. In some embodiments, steps that are repeated one or more times include: (b) contacting the polymer with a binder that can bind to the polymer and includes a coding tag having identification information relating to the binder, thereby enabling binding between the polymer and the binder; (c) conjugating the 5' end of the recording tag to the 3' end of the coding tag with a nucleic acid conjugation reagent; (d) extending the recording tag using the coding tag as a template with polymerase to produce a double-stranded extended recording tag; and (e) cleaving the double-stranded extended recording tag with a double-stranded nucleic acid cleavage reagent to produce a 3' overhang within the extended recording tag.

[0068] In some embodiments, this method further includes removing the binder. For example, the binder can be removed after the information from the coding tag has been transferred to the recording tag. Once the encoding cycle has been performed, the binder can be removed before repeating step (b). The binder can be removed or released by providing appropriate conditions (e.g., reagents and temperature) during washing.

[0069] In some embodiments, the method further includes removing a portion of the polymer before repeating step (b). For example, if a polypeptide is being analyzed, the method may include removing one or more amino acids (e.g., from the terminal) of the polypeptide before repeating step (b). In some cases, before repeating step (b), the N-terminal amino acid (NTAA) polypeptide is removed from the polypeptide to expose the new NTAA of the polypeptide. In some embodiments, the removed portion of the polymer is treated with or modified with a chemical or enzymatic agent. In some cases, the polypeptide is treated with a reagent for modifying the terminal amino acids of the polypeptide. Modification of the polypeptide, such as NTAA, may be performed before step (b). In some cases, modification of the polypeptide, such as NTAA, may be performed after step (e).

[0070] In some cases, this method further includes one or more washing steps before, during, or after one or more of the steps. For example, after the binder is provided in step (b), the washing step may be performed before any of the enzymatic reagents (e.g., nucleic acid conjugation reagent, polymerase, double-strand nucleic acid cleavage reagent) are introduced. Such a washing step can remove nonspecifically bound binder. The stringency of the washing step may be adjusted according to the affinity of the binder. In some cases, it is preferable that a washing step is not required between steps (c), (d), and (e). In some other embodiments, the washing step is performed after ligation in step (c). After ligation and washing, subsequent steps including polymerase and double-strand nucleic acid cleavage reagent can be provided in a single step. In some other embodiments, the washing step is performed after extension in step (d). In some cases, the washing step is performed before steps (c), (d), and / or step (e). After ligation and extension have occurred in one step, a washing step may be performed before the double-strand nucleic acid cleavage reagent is provided. In some cases, this method further includes removing the binder after the information has been transferred to the recoding tag, such as before steps (e) and / or (b) are repeated in subsequent cycles.

[0071] In some embodiments, the final information transfer cycle is performed to provide a capping sequence to the extension record tag, and optionally, the capping sequence includes a universal priming site for amplification, sequencing, or both. The capping sequence may be provided to the extension record tag by a binder configured to bind to a universal feature of the polymer. Thus, the binder for delivering the capping sequence is configured to bind to a target region contained by all or many of the polymers in the sample. In some examples, the universal feature is a chemical modification of the polypeptide. The capping sequence may include sequences useful for analysis of the extension record tag, such as a universal reverse priming site.

[0072] A. Recording tags Polymers (e.g., proteins or polypeptides) for analysis may be labeled with recording tags containing nucleic acid molecules or oligonucleotides. In some embodiments, multiple polymers in a sample are provided with recording tags. The recording tags may be directly or indirectly associated with or attached to the polymers using any suitable means. In some embodiments, a polymer may be associated with one or more recording tags. In some embodiments, the recording tags may be directly or indirectly associated with or attached to the polymers before contact with a binder.

[0073] In some embodiments, at least one record tag is directly or indirectly associated with or colocalized with a polymer (e.g., a polypeptide). Providing the polymer and associated record tag in step (a) may include processing the record tag and any associated nucleic acid to conjugate, cleave, or otherwise prepare the record tag for assay. In some embodiments, step (a) includes providing a barcode and / or UMI to the record tag using ligation and / or extension. In some embodiments, step (a) includes cleaving the record tag using a restriction enzyme to generate a 3' overhang. For example, the 3' overhang of the record tag is generated by extension using polymerase and / or cleavage with a double-strand nucleic acid cleavage reagent. An exemplary workflow for preparing and providing the record tag is shown in Figure 1A.

[0074] In certain embodiments, a single record tag is attached to a polypeptide, for example, by attachment to an N-terminal or C-terminal amino acid. In other embodiments, multiple record tags are attached to a polypeptide, for example, a lysine residue or a peptide backbone. In some embodiments, a polypeptide labeled with multiple record tags is fragmented or digested into smaller peptides, each peptide being labeled with an average of one record tag.

[0075] The recording tag may contain DNA, RNA, or polynucleotide analogs including PNA, gPNA, GNA, HNA, BNA, XNA, TNA, or combinations thereof. The recording tag may be single-stranded, or partially or completely double-stranded. In some specific embodiments, certain advantages are associated with recording tags containing double-stranded regions. For example, the advantage may be reduced DNA-DNA interactions with other nucleic acid components of the system. In some cases, the recording tag may contain a nucleic acid hairpin. The recording tag may have blunt or protruding ends. In some specific embodiments, the recording tag associated with the polymer is processed or treated (e.g., via digestion) so that it has a 3' overhang. In certain embodiments, all or substantial amounts of the polymer in a sample (e.g., at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100%) are labeled with the recording tag. In other embodiments, a subset of polymers in a sample is labeled with a recording tag. In certain embodiments, a subset of polymers from a sample undergoes targeted (analyte-specific) labeling with a recording tag. For example, targeted recording tag labeling of a protein can be achieved using a target protein-specific binder (e.g., an antibody, aptamer, etc.). In some embodiments, the recording tag (or part thereof) is attached to the polymer before the sample is provided on the support. In some embodiments, the recording tag is attached to the polymer after the sample has been provided on the support.

[0076] In some embodiments, the recording tag may include other nucleic acid components. In some embodiments, the recording tag may include a unique molecular identifier, a compartment tag, a partition barcode, a sample barcode, a fraction barcode, a spacer sequence, a universal priming site, or any combination thereof. In some embodiments, the recording tag may include a blocking group, such as the 3' end of the recording tag. In some cases, the 3' end of the recording tag is blocked to prevent extension of the recording tag by polymerase.

[0077] In some embodiments, the recording tag may include a sample that identifies the barcode. The sample barcode is useful for multiplex analysis of a set of samples in a single reaction vessel, or it may be immobilized on a single solid substrate or a collection of solid substrates (e.g., a planar slide, a collection of beads in a single tube or container). For example, polymers from many different samples can be labeled with recording tags bearing sample-specific barcodes, and then all samples can be pooled together before immobilization on a support, periodic binding of a binder, and recording tag analysis. Alternatively, samples can be kept separate until after the creation of a DNA coding library, and the sample barcodes can be attached during PCR amplification of the DNA coding library, then mixed together and sequenced. This approach may be useful when assaying analytes (e.g., proteins) of different abundance classes.

[0078] In certain embodiments, the recording tag includes an optional UMI, which provides a unique identifier tag for each polymer (e.g., polypeptide) associated with a unique molecular identifier (UMI). The UMI may be about 3 to 40 bases, about 3 to 30 bases, about 3 to 20 bases, or about 3 to 10 bases, or about 3 to 8 bases. In some embodiments, the UMI is about 3 bases, 4 bases, 5 bases, 6 bases, 7 bases, 8 bases, 9 bases, 10 bases, 11 bases, 12 bases, 13 bases, 14 bases, 15 bases, 16 bases, 17 bases, 18 bases, 19 bases, 20 bases, 25 bases, 30 bases, 35 bases, or 40 bases in length. The UMI can be used to deconvolve sequencing data from multiple extension recording tags to identify sequence reads from individual polymers. In some embodiments, within a library of polymers, each polymer is associated with a single record tag, and each record tag contains a unique UMI. In other embodiments, multiple copies of a record tag are associated with a single polymer, and each copy of the record tag contains the same UMI. In some embodiments, the UMI has a different nucleotide sequence from a spacer or coding tag to facilitate the distinction of these components during sequence analysis. In some embodiments, the UMI may function as a positional identifier and may also provide information in polymer analysis assays. For example, the UMI can be used to identify molecules that belong to the same lineage and therefore originate from the same initial molecule. In some embodiments, this information can be used to compensate for amplification variations and to detect and correct sequencing errors during analysis.

[0079] In some embodiments, the recording tag includes a spacer polymer. In certain embodiments, the recording tag includes a spacer at its end, e.g., the 3' end. As used herein, a reference to a spacer sequence in relation to a recording tag includes a spacer sequence identical to the spacer sequence associated with its homogeneous binder, or a spacer sequence complementary to the spacer sequence associated with its homogeneous binder. The end of the recording tag, e.g., the 3' spacer, allows for the transfer of homogeneous binder identification information from the coding tag to the recording tag during the first binding cycle (e.g., by annealing a complementary spacer sequence for primer extension or sticky end ligation). In one embodiment, the spacer sequence is about 1 to 20 base lengths, about 2 to 12 base lengths, or 5 to 10 base lengths. In some cases, the spacer sequence is about 2 to 5 base lengths. The length of the spacer may vary depending on factors such as temperature and reaction conditions for transferring coding tag information to the recording tag.

[0080] In some embodiments using spacer sequences, the recording tag associated with the polypeptide library shares a common spacer sequence. In other embodiments, the recording tag associated with the polypeptide library has a binding cycle-specific spacer sequence complementary to the binding cycle-specific spacer sequence of the adapter molecule. In some embodiments, the spacer sequence within the recording tag is designed to have minimal complementarity to other regions within the recording tag. In some cases, the spacer sequence of the recording tag needs to have minimal sequence complementarity to components present in the recording tag or coding tag, such as unique molecular identifiers, barcodes (e.g., compartment, partition, sample, spatial position), universal primer sequences, coding tag sequences, and cycle-specific sequences. In some embodiments, the spacer is designed based on a double-strand nucleic acid cleavage reagent (e.g., restriction enzyme) selected for use.

[0081] In certain embodiments, the recording tag includes a universal priming site, e.g., a forward or 5' universal priming site. The universal priming site is a nucleic acid sequence that can be used for priming and / or sequencing of a library amplification reaction. The universal priming site may include, but is not limited to, a priming site for PCR amplification, a flow cell adapter sequence that anneals to a complementary oligonucleotide on the flow cell surface (e.g., Illumina next-generation sequencing), a sequencing priming site, or a combination thereof. The universal priming site may be about 10 to about 60 bases long. In some embodiments, the universal priming site includes an Illumina P5 primer (AATGATACGGCGACCACCGA, SEQ ID NO: 1) or an Illumina P7 primer (CAAGCAGAAGACGGCATACGAGAT, SEQ ID NO: 2).

[0082] In certain embodiments, the recording tag includes a compartment tag. In some embodiments, the compartment tag is a component within the recording tag. In some embodiments, the recording tag may also include a barcode representing a compartment tag to which a barcode specific to the compartment is assigned, such as a droplet, microwell, or physical area on a support. The association of a compartment with a specific barcode can be achieved in any number of ways, such as by encapsulating barcoded beads in the compartment, by directly merging or adding barcoded droplets to the compartment, or by directly printing or injecting barcoded reagents into the compartment. Barcoded reagents in the compartment can be used to add a compartment-specific barcode to a polymer or fragment thereof within the compartment. Applied to protein fractionation into compartments, the barcode can be used to map analyzed peptides to their protein molecules of origin within the compartment. This can greatly facilitate the identification of proteins. Protein complexes can also be identified using compartment barcodes. In other embodiments, multiple compartments representing a subset of a population of compartments may be assigned a unique barcode representing the subset. In some embodiments, the recording tag includes a fractionation barcode containing identification information for the polymer within the fraction.

[0083] In some embodiments, one or more tags (e.g., compartment tags, partition barcodes, sample barcodes, fraction barcodes, etc.) further include a functional moiety that can react with multiple protein complexes, proteins, or internal amino acids, peptide backbone, or N-terminal amino acids on polypeptides. In some embodiments, the functional moiety is a click chemical moiety, an aldehyde, an azide / alkyne, or a maleimide / thiol, or an epoxide / nucleophile, an inverse electron-required Diels-Alder (iEDDA) group, or a moiety for a Staudinger reaction. In some specific embodiments, multiple compartment tags are formed by printing, spotting, inkjet printing, or a combination thereof onto compartments. In some embodiments, the tags are attached to polypeptides to link the tags to the polymer via polypeptide-polypeptide linkages. In some embodiments, the tag-attached polypeptide includes a protein ligase recognition sequence.

[0084] In certain embodiments, a peptide or polypeptide polymer may be immobilized on a support by an affinity capture reagent (and optionally covalently bonded), and a recording tag may be directly associated with the affinity capture reagent, or the polymer may be directly immobilized on the support by the recording tag.

[0085] In some embodiments, a polymer is attached to a bait nucleic acid to form a nucleic acid-polymer chimera. The immobilization method may include bringing the nucleic acid-polymer chimera close to the support by hybridizing the bait nucleic acid to a captured nucleic acid attached to the support, and covalently coupling the nucleic acid-polymer chimera to a solid support. In some cases, the nucleic acid-polymer chimera is indirectly coupled to the solid support, such as via a linker. In some embodiments, multiple nucleic acid-polymer chimeras are coupled on a solid support, and any adjacently coupled nucleic acid-polymer chimeras are separated from each other by an average distance of about 50 nm or more. In some cases, the bait nucleic acid includes a universal priming site or a portion thereof.

[0086] As shown in the exemplary format of Figure 1A, the peptide to be analyzed is attached to a bait nucleic acid that hybridizes to a capture nucleic acid hairpin immobilized on a support. The bait nucleic acid is ligated to a capture nucleic acid that contains a reactive coupling moiety for attachment to the support. In some examples, the bait or the capture nucleic acid, or the ligated bait and capture nucleic acid, can function as a recording tag from which information about the polypeptide can be transferred.

[0087] Figure 1A shows exemplary steps for preparing a recording tag using a capture nucleic acid that includes or is a nucleic acid hairpin having a recessed 5' phosphorylated end. In the second panel of Figure 1A, the bait nucleic acid with the attached peptide hybridizes with the capture nucleic acid hairpin and is ligated to the capture nucleic acid hairpin. The bait nucleic acid may include one or more barcodes as shown, and may also include restriction enzyme sites (or partial restriction enzyme sites). In the third panel of Figure 1A, a polymerase reaction is used to extend the 3' end of the capture nucleic acid hairpin to produce a double-stranded recording tag construct with the attached peptide. Once extension occurs using the bait nucleic acid as a template, the digestion site is double-stranded and can be recognized by an IIS-type restriction enzyme for cleavage, and the cleavage produces a recording tag having a 3' overhang (2-base pair sequence) with a recessed 5' phosphorylated end. This recording tag, including one or more barcodes, is available for information transfer from a coding tag. In some embodiments, the preparation of polymers for analysis using recording tags and immobilization on a support can be carried out using the method described in International Patent Application PCT / US2020 / 27840. The preparation and / or immobilization of polymers can be carried out before and / or separately from the information transfer step.

[0088] In some embodiments, the density or number of polymers with recording tags is controlled or titrated. In some examples, a desired interval, density, and / or quantity of recording tags in a sample can be titrated by providing a diluted or controlled number of recording tags. In some examples, a desired interval, density, and / or quantity of recording tags can be achieved by spiking competing or “dummy” competing molecules when providing, associating, and / or attaching the recording tags. In some cases, “dummy” competing molecules react in the same way as recording tags associated with or attached to polymers in the sample, but the competing molecules do not function as recording tags. In some specific examples, if the desired density is one functional recording tag per 1,000 sites available for attachment in the sample, the desired interval is achieved by using spikes in one functional recording tag for every 1,000 “dummy” competing molecules. In some examples, the ratio of functional recording tags is adjusted based on the reaction rate of the functional recording tags compared to the reaction rate of the competing molecules.

[0089] In some examples, labeling of polymers with recording tags is carried out using standard amine coupling chemistry. For example, e-amino groups (e.g., lysine residues) and N-terminal amino groups may be sensitive to labeling with amine-reactive coupling agents depending on the pH of the reaction (Mendoza et al., Mass Spectrom Rev(2009)28(5):785-815). In certain embodiments, the recording tag includes a reactive moiety (e.g., for a solid surface, a multifunctional linker, or for conjugation to the polymer), a linker, a universal priming sequence, a barcode (e.g., a compartment tag, a partition barcode, a sample barcode, a fraction barcode, or any combination thereof), an optional UMI, and a spacer (Sp) sequence to facilitate information transfer. In another embodiment, a protein can be first labeled with a universal DNA tag, and a barcode-Sp sequence (representing a sample, compartment, physical location on a slide, etc.) is later attached to the protein through an enzymatic or chemical coupling step. A universal DNA tag comprises a short sequence of nucleotides used to label a protein or polypeptide macromolecule and can be used as a barcode binding point (e.g., a compartment tag, a recording tag, etc.). For example, a recording tag may have a sequence at its end that is complementary to the universal DNA tag. In certain embodiments, the universal DNA tag is a universal priming sequence. When the universal DNA tag on a labeled protein hybridizes with the complementary sequence in a recording tag (e.g., bound to a bead), the annealed universal DNA tag can be extended via primer extension, transferring the recording tag information to the DNA-tagged protein. In certain embodiments, the protein is labeled with a universal DNA tag before protease digestion into a peptide. The universal DNA tag on the labeled peptide from the digest can then be converted into a useful and effective recording tag.

[0090] The recording tag may contain a reactive moiety to a congenerally reactive moiety present on a polymer, such as a protein (e.g., click chemical labeling, photoaffinity labeling). For example, the recording tag may contain an azide moiety for interacting with alkyne-derivative proteins, or it may contain a benzophenone for interacting with native proteins, etc. Upon binding of the target protein with a target protein-specific binder, the recording tag and the target protein are coupled via their corresponding reactivity. After labeling the target protein with the recording tag, the target protein-specific binder may be removed by digestion of a DNA capture probe linked to the target protein-specific binder. For example, the DNA capture probe may be designed to contain a uracil base and then targeted for digestion with a uracil-specific excision reagent (e.g., USER®). The target protein-specific binder can be dissociated from the target protein. In some embodiments, the recording tag can be linked to the polymer using other types of linkage other than hybridization. Suitable linkers can be attached to various positions on the recording tag, such as the 3' end in an internal position, or within a linker attached to the 5' end of the recording tag.

[0091] Information from one or more coding tags is transferred to a recording tag to generate an extended recording tag. In some embodiments, the extended recording tag includes a universal forward (or 5') priming array, information transferred from one or more coding tags, and a spacer array. In some embodiments, the extended recording tag includes a universal forward (or 5') priming array, any barcode and / or UMI (e.g., sample barcode, partition barcode, compartment barcode, or any combination thereof), one or more coding tags, a spacer array, and information transferred from a universal reverse (or 3') priming array.

[0092] B. Binder The methods described herein use binders configured to interact with polymers (e.g., polypeptides, peptides, proteins) to be analyzed. The assay may involve contacting multiple binders with multiple polymers. In some embodiments, the method may involve contacting a single polymer with a single binder, contacting multiple polymers with a single binder, or contacting multiple polymers with multiple binders. In some embodiments, the multiple binders may include a mixture of binders configured to bind to different target moieties.

[0093] The binder can be any molecule that can bind to a component or mechanism of a polypeptide (e.g., peptides, polypeptides, proteins, nucleic acids, carbohydrates, small molecules, etc.). The binder can be a naturally occurring, synthetically produced, or recombinantly expressed molecule. In some embodiments, the scaffold used to manipulate the binder can be derived from any species, e.g., human, non-human, or transgenic. The binder can bind to a portion of the target polymer or motif. The binder can bind to a single monomer or subunit of the polypeptide (e.g., a single amino acid) or to multiple linked subunits of the polypeptide (e.g., dipeptides, tripeptides, or higher-order peptides of longer polypeptide molecules).

[0094] In some examples, the binders include antibodies, antigen-binding antibody fragments, single-domain antibodies (sdAb), recombinant heavy-chain-only antibodies (VHH), single-chain antibodies (scFv), shark-derived variable domains (vNAR), Fv, Fab, Fab', F(ab')2, linear antibodies, diabodies, aptamers, peptide mimetic molecules, fusion proteins, reactive or non-reactive small molecules, or synthetic molecules.

[0095] In certain embodiments, the binder may be designed to form a covalent bond. The covalent bond may be conditional or designed to be advantageous when it binds to the correct site. For example, the target and its homologous binders may each be modified with reactive groups such that when the target-specific binder is bound to the target, a coupling reaction takes place to create a covalent bond between the two. Nonspecific binding of the binder to other sites lacking homologous reactive groups will not result in a covalent bond. In some embodiments, the target contains a ligand that can form a covalent bond to the binder. In some embodiments, the target contains a ligand group that can covalently bind to the binder. The covalent bond between the binder and its target may allow for the removal of nonspecifically bound binders using more rigorous washing, thus increasing the specificity of the assay. In some embodiments, this method includes a washing step after contacting the binder with a polymer to remove nonspecifically bound binders. The stringency of the washing step may be adjusted depending on the affinity of the binder to the target and / or the strength and stability of the formed complex.

[0096] In some embodiments, the binder is configured to provide specificity for the binder to the polymer. In certain embodiments, the binder may be a selective binder. As used herein, selective binding refers to the ability of a binder to preferentially bind to a particular ligand (e.g., an amino acid or a class of amino acids) compared to binding to different ligands (e.g., an amino acid or a class of amino acids). Selectivity is generally referred to as the equilibrium constant of the reaction of substitution of one ligand with another ligand in a complex with a binder. Typically, such selectivity is associated with the spatial geometry of the ligand, and / or the way and extent to which the ligand binds to the binder, for example, by hydrogen bonds, hydrophobic bonds, and van der Waals forces (non-covalent interactions), or by reversible or irreversible covalent bonds to the binder. It should also be understood that selectivity can be relative rather than absolute, and that different factors, including ligand concentration, can affect the same thing. Thus, in one example, the binder selectively binds to one of 20 standard amino acids. In some examples, the binder binds to an N-terminal amino acid residue, a C-terminal amino acid residue, or an internal amino acid residue.

[0097] In some embodiments, the binder is partially specific or selective. In some embodiments, the binder preferentially binds to one or more amino acids. In some examples, the binder can bind to or be able to bind to two or more of the 20 standard amino acids. For example, the binder may preferentially bind to amino acids A, C, and G over other amino acids. In some other embodiments, the binder may selectively or specifically bind to multiple amino acids. In some embodiments, the binder may also preferentially bind to one or more amino acids located at positions such as the second, third, fourth, fifth, etc., from the terminal amino acids. In some cases, the binder preferentially binds to specific terminal amino acids and the second-to-last amino acid. For example, the binder may preferentially bind to AA, AC, and AG, or the binder may preferentially bind to AA, CA, and GA. In some embodiments, the binder may exhibit flexibility and variability in target binding preference at some or all of the target positions. In some examples, the binder may preferentially bind to one or more specific target terminal amino acids and flexibly preferentially bind to the second-to-last target. In some other examples, the binder may preferentially target one or more specific target amino acids at the second-to-last amino acid position, and may flexibly preferentially target at terminal amino acid positions. In some embodiments, the binder is selective for targets including terminal amino acids and other components of the polymer. In some examples, the binder is selective for targets including terminal amino acids and at least a portion of the peptide backbone. In some specific examples, the binder is selective for targets including terminal amino acids and an amide peptide backbone. In some cases, the peptide backbone includes a native peptide backbone or post-translational modifications. In some embodiments, the binder exhibits allosteric linkage.

[0098] In some embodiments, this method involves contacting a mixture of binders with a mixture of polymers, and selectivity is required only with respect to other binders to which the target is exposed. It should also be understood that the selectivity of a binder does not need to be absolute with respect to a particular molecule, but may be absolute with respect to a part of the molecule. In some examples, the selectivity of a binder does not need to be absolute with respect to a particular amino acid, but may be selective with respect to a class of amino acids, such as amino acids having polar or nonpolar side chains, or having electrically charged (positively or negatively) side chains, or aromatic side chains, or side chains of a certain particular class or size. In some embodiments, the ability of a binder to selectively bind to a feature or component of a polymer is characterized by comparing the binding aptitudes of the binders. For example, the binding aptitude of a binder to a target can be compared with the binding aptitude of a binder to a different target, for example, a binder selective to a class of amino acids can be compared with a binder selective to a different class of amino acids. In some examples, a binder selective to nonpolar side chains is compared with a binder selective to polar side chains. In some embodiments, a selective binder for a feature, peptide component, or one or more amino acids exhibits at least 1, at least 2, at least 5, at least 10, at least 50, at least 100, or at least 500 times higher binding compared to a different selective binder for a feature, peptide component, or one or more amino acids.

[0099] In certain embodiments, the binder has high affinity and high selectivity for the target polymer, e.g., polypeptide. In particular, high binding affinity at low off-rate may be effective for hybridization of the adapter molecule to the coding tag. In certain embodiments, the binder has a Kd of less than about 500 nM, less than 200 nM, less than 100 nM, less than 50 nM, less than 10 nM, less than 5 nM, less than 1 nM, less than 0.5 nM, or less than 0.1 nM. In certain embodiments, the binder is added to the polypeptide at a concentration greater than 1, 5, 10, 100, or 1000 times its Kd to drive binding to completion. For example, the binding kinetics of an antibody to a single protein molecule is Chang This is described in et al., J Immunol Methods (2012) 378(1-2):102-115.

[0100] In certain embodiments, the binder may bind to a peptide, intervening amino acid, dipeptide (a sequence of two amino acids), tripeptide (a sequence of three amino acids), or the terminal amino acid of a higher-order peptide in a peptide molecule. In some embodiments, each binder in the binder library selectively binds to a specific amino acid, for example, one of 20 standard naturally occurring amino acids. Standard natural amino acids include alanine (A or Ala), cysteine ​​(C or Cys), aspartic acid (D or Asp), glutamic acid (E or Glu), phenylalanine (F or Phe), glycine (G or Gly), histidine (H or His), isoleucine (I or Ile), lysine (K or Lys), leucine (L or Leu), methionine (M or Met), asparagine (N or Asn), proline (P or Pro), glutamine (Q or Gln), arginine (R or Arg), serine (S or Ser), threonine (T or Thr), valine (V or Val), tryptophan (W or Trp), and tyrosine (Y or Tyr). In some embodiments, the binder is attached to an unmodified or natural (e.g., natural) amino acid. In some examples, the binder binds to unmodified or native dipeptides (sequences of two amino acids), tripeptides (sequences of three amino acids), or higher-order peptides of peptide molecules. The binder may be designed for high affinity to native or unmodified N-terminal amino acids (NTAAs), high specificity to native or unmodified NTAAs, or both. In some embodiments, the binder can be developed through the directed evolution of promising affinity scaffolds using phage display.

[0101] In certain embodiments, the binder may bind to post-translational modifications of amino acids. In some embodiments, the peptide comprises one or more post-translational modifications, which may be the same or different. The NTAA, CTAA, intervening amino acids, or combinations thereof of the peptide may be post-translational modified. Examples of post-translational modifications to amino acids include acylation, acetylation, alkylation (including methylation), biotinylation, butyrylation, carbamylation, carbonylation, deamidation, deiminoation, diphthamide formation, disulfide crosslinking, eliminylation, flavin attachment, formylation, γ-carboxylation, glutamylation, glycylation, glycosylation, glycieation, heme C attachment, hydroxylation, and hypsination. Examples of methylation include, but are not limited to, iodization, isoprenylation, lipidation, lipoylation, malonylation, methylation, myristolylation, oxidation, palmitoylation, pegylation, phosphopantethenylation, phosphorylation, prenylation, propionylation, retinilidensif base formation, S-glutathione addition, S-nitrosylation, S-sulfenylation, selenization, succinylation, sulfinization, ubiquitination, and C-terminal amidation (see also Seo and Lee, 2004, J. Biochem. Mol. Biol. 37:35-44).

[0102] In certain embodiments, lectins are used as binders for detecting the glycosylation status of proteins, polypeptides, or peptides. Lectins are carbohydrate-binding proteins that can selectively recognize glycan epitopes of free carbohydrates or glycoproteins. The list of lectins that recognize various glycosylation states (e.g., core fucose, sialic acid, N-acetyl-D-lactosamine, mannose, N-acetyl-glucosamine) includes A, AAA, AAL, ABA, ACA, ACG, ACL, AOL, ASA, BanLec, BC2L-A, BC2LCN, BPA, BPL, Calsepa, CGL2, CNL, Con, ConA, DBA, Discoidin, DSA, ECA, EEL, F17AG, Gal1, Gal1-S, Gal2, Gal3, Gal3C-S, Gal7-S, Gal9, GNA, GRFT, GS-I, GS-II, GSL-I, GSL-II, HHL, HIHA, HPA, I , II, Jacalin, LBA, LCA, LEA, LEL, Lentil, Lotus, LSL-N, LTL, MAA, MAH, MAL_I, Malectin, MOA, MPA, MPL, NPA, Orysata, PA-IIL, PA-IL, PALa, PHA-E, PHA-L, PHA-P, PHAE, PHAL, P Includes NA, PPL, PSA, PSL1a, PTL, PTL-I, PWM, RCA120, RS-Fuc, SAMB, SBA, SJA, SNA, SNA-I, SNA-II, SSA, STL, TJA-I, TJA-II, TxLCI, UDA, UEA-I, UEA-II, VFA, VVA, WFA, WGA (Zhang et al., 2016, MABS 8:524-535).

[0103] In some embodiments, the binder may bind to native or unmodified or unlabeled or native terminal amino acids. In some examples, the binder binds to unmodified or native dipeptides (sequences of two amino acids), tripeptides (sequences of three amino acids), or higher-order peptides of peptide molecules. The binder may be engineered for high affinity to modified NTAAs, high specificity to modified NTAAs, or both. In some embodiments, the binder may be developed through the directed evolution of promising affinity scaffolds using phage display.

[0104] By using the directed evolution of protein / enzyme scaffolds, it is possible to generate higher affinity, higher specificity binders that recognize the N-terminal amino acid in the context of N-terminal labeling. For example, Havranak et al. (U.S. Patent Publication 2014 / 0273004) describe manipulating aminoacyl-tRNA synthetase (aaRS) as a specific NTAA binder. The amino acid binding pocket of aaRS has an inherent ability to bind to congener amino acids, but generally exhibits low binding affinity and specificity. Furthermore, these native amino acid binders do not recognize N-terminal labels. By using the directed evolution of aaRS scaffolds, it is possible to generate higher affinity, higher specificity binders that recognize the N-terminal amino acid in the context of N-terminal labeling.

[0105] In certain embodiments, the binder may bind to modified or labeled terminal amino acids (e.g., functionalized or modified NTAAs). In some embodiments, the binder may bind to chemically or enzymatically modified terminal amino acids. Modified or labeled NTAAs include phenylisothiocyanate, PITC, 1-fluoro-2,4-dinitrobenzene (Sanger reagent, DNFB), benzyloxycarbonyl chloride or carbobenzoxy chloride (Cbz-Cl), N-(benzyloxycarbonyloxy)succinimide (Cbz-OSu or Cbz-O-NHS), dansilchloride (DNS-Cl or 1-dimethylaminonaphthalene-5-sulfonyl chloride), 4-sulfonyl-2-nitrofluorobenzene (SNFB), N-acetyl-isatoic anhydride, isatoic anhydride, 2-pyridinecarboxaldehyde, 2-formylphenylboronic acid, 2-acetylphenylboronic acid, 1-fluoro-2,4-dinitrobenzene, succinic anhydride, and 4-chloro-7-nitroben Zoflazan, pentafluorophenyl isothiocyanate, 4-(trifluoromethoxy)-phenyl isothiocyanate, 4-(trifluoromethyl)-phenyl isothiocyanate, 3-(carboxylic acid)-phenyl isothiocyanate, 3-(trifluoromethyl)-phenyl isothiocyanate, 1-naphthyl isothiocyanate, N-nitroimidazole-1-carboxymidoamide, N,N,A≦-bis(pivaloyl)-1H-pyrazole-1-carboxamidine, N,N,A≦-bis(benzyloxycarbonyl)-1H-pyrazole-1-carboxamidine, acetylating reagents, guanidinylating reagents, thioacylation reagents, thioacetylating reagents, or thiobenzylating reagents, or biheterocyclic methaneimine reagents may be functionalized. In some examples, the binder is bound to the labeled amino acid by contact with the reagent or by using the method described in International Patent Publication No. 2019 / 089846. In some cases, the binder binds to the amino acid labeled with the amine modification reagent.

[0106] The binder may bind to the N-terminal peptide, C-terminal peptide, or intervening peptide of a peptide, polypeptide, or protein molecule. The binder may bind to the N-terminal amino acid, C-terminal amino acid, or intervening amino acid of a peptide molecule. The binder may bind to the N-terminal or C-terminal diamino acid portion. The N-terminal diamino acid consists of the N-terminal amino acid and the second-to-last N-terminal amino acid. The C-terminal diamino acid is similarly defined for the C-terminus. In some embodiments, the binder binds to a chemically modified N-terminal amino acid residue or a chemically modified C-terminal amino acid residue. To increase the affinity of the binder to the small N-terminal amino acid (NTAA) of the peptide, the NTAA may be modified with an "immunogenic" hapten such as dinitrophenol (DNP). This can be implemented using a cycle sequencing approach with Sanger's reagent dinitrofluorobenzene (DNFB), which attaches the DNP group to the amine group of the NTAA. Commercially available anti-DNP antibodies have affinities in the low nM range (approximately 8 nM, LO-DNP-2) (Bilgicer et al., J Am Chem Soc (2009) 131(26):9361-9367), and therefore it is naturally possible to manipulate high-affinity NTAA conjugates for a large number of NTAAs modified with DNP (via DNFB) while simultaneously achieving good binding selectivity for specific NTAAs. In another example, NTAAs can be modified with sulfonylnitrophenol (SNP) using 4-sulfonyl-2-nitrofluorobenzene (SNFB). Similar affinity enhancements can also be achieved with alternative NTAA modifiers such as acetyl groups or amidinyl (guanidinyl) groups.

[0107] In certain embodiments, the binder may be an aptamer (e.g., a peptide aptamer, a DNA aptamer, or an RNA aptamer), a peptoid, an antibody or its specific binding fragment, an amino acid-binding protein or enzyme, an antibody-binding fragment, an antibody mimetic, a peptide, a peptide mimetic, a protein, or a polynucleotide (e.g., DNA, RNA, peptide nucleic acid (PNA), gPNA, cross-linked nucleic acid (BNA), xeno nucleic acid (XNA), glycerol nucleic acid (GNA), or threose nucleic acid (TNA), or a variant thereof).

[0108] As used herein, the terms antibody and antibody(plural) are used in a broad sense to include not only intact antibody molecules, such as immunoglobulin A, immunoglobulin G, immunoglobulin D, immunoglobulin E, and immunoglobulin M, but also any immunoreactive components of antibody molecules or parts thereof that bind immunospecifically to at least one epitope. Antibodies may occur naturally, be produced synthetically, or be recombinantly expressed. Antibodies may be fusion proteins. Antibodies may be antibody mimics. Examples of antibodies include, but are not limited to, Fab fragments, Fab' fragments, F(ab')2 fragments, single-chain antibody fragments (scFv), mini-antibodies, diabodies, cross-linked antibody fragments, Affibody®, nanobodies, single-domain antibodies, DVD-Ig molecules, alphabodies, affimers, affitins, cyclotides, and molecules. Immunoreactive products induced using antibody engineering or protein engineering techniques are also explicitly within the scope of the term antibody. Detailed descriptions of antibody and / or protein engineering, including related protocols, can be found, in particular, in J. Maynard and G. Georgiou, 2000, Ann. Rev. Biomed. Eng. 2:339-76, Antibody Engineering, R. Kontermann and S. Dubel, eds., Springer Lab Manual, Springer Verlag (2001), U.S. Patent No. 5,831,012, and S. Paul, Antibody Engineering Protocols, Humana Press (1995).

[0109] Similar to antibodies, nucleic acid and peptide aptamers that specifically recognize macromolecules, such as peptides or polypeptides, can be produced using known methods. Aptamers typically bind to target molecules with very high affinity in a highly specific and conformation-dependent manner, although aptamers with lower binding affinity can be selected if necessary. Aptamers have been shown to differentiate targets based on very small structural differences, such as the presence or absence of methyl or hydroxyl groups, and certain aptamers can distinguish between D-enantiomers and L-enantiomers. Aptamers that bind to small molecule targets, including drugs, metal ions, and organic dyes, peptides, biotin, and proteins (including, but not limited to, streptavidin, VEGF, and viral proteins), have been obtained. Aptamers have been shown to retain functional activity after biotinylation, fluorescein labeling, and when attached to glass surfaces and microspheres (see, e.g., Jayasena, 1999, Clin Chem 45:1628-50; Kusser, 2000, J. Biotechnol. 74:27-39; Colas, 2000, Curr Opin Chem Biol 4:54-9). Aptamers that specifically bind to arginine and AMP have also been described (see Patel and Suri, 2000, J. Biotech. 74:39-60). Oligonucleotide aptamers that bind to specific amino acids are disclosed in Gold et al. (1995, Ann. Rev. Biochem. 64:763-97). RNA aptamers that bind to amino acids have also been described (Ames and Breaker, 2011, RNA Biol. 8; 82-89; Mannironi et al., 2000, RNA 6: 520-27; Famulok, 1994, J. Am. Chem. Soc. 116: 1698-1706).

[0110] Binding agents can be created by modifying naturally occurring or synthetically produced proteins through genetic engineering to introduce one or more mutations into their amino acid sequences, thereby producing engineered proteins that bind to specific components or mechanisms of polypeptides (e.g., NTAAs, CTAAs, or post-translationally modified amino acids or peptides). For example, exopeptidases (e.g., aminopeptidases, carboxypeptidases, dipeptidylpeptidases, dipeptidylaminopeptidases), exoproteases, mutant exoproteases, mutant anticharin, mutant ClpS, antibodies, or tRNA synthetases can be modified to create binding agents that selectively bind to specific NTAAs. In another example, carboxypeptidases can be modified to create binding agents that selectively bind to specific CTAAs. The binders may also be designed or modified and utilized to specifically bind to modified NTAAs or modified CTAAs, for example, those with post-translational modifications (e.g., phosphorylated NTAA or phosphorylated CTAA) or those modified by labeling (e.g., PTC, 1-fluoro-2,4-dinitrobenzene (using Sanger reagent, DNFB), dansilk chloride (using DNS-Cl, or 1-dimethylaminonaphthalene-5-sulfonyl chloride), or thioacylation reagents, thioacetylation reagents, acetylation reagents, amidation (guanidinylation) reagents, or thiobenzylation reagents). Strategies for the directional evolution of proteins are known in the art (e.g., Yuan et al., 2005, Microbiol. Mol. Biol. Rev. 69:373-392) and include phage displays, ribosome displays, mRNA displays, CIS displays, CAD displays, emulsions, cell surface displays, yeast surface displays, bacterial surface displays, etc.

[0111] In another example, highly selectively engineered ClpS proteins have also been documented. Emili et al. describe the directional evolution of the E. coli ClpS protein via phage display, resulting in four distinct variants with the ability to selectively bind to NTAAs of aspartic acid, arginine, tryptophan, and leucine residues (U.S. Patent No. 9,566,335, the entire patent is incorporated by reference). In one embodiment, the binding portion of the binder comprises an evolutionarily conserved member or variant of the ClpS family of adapter proteins involved in the recognition and binding of the native N-terminal protein. (e.g., Schuenemann et al., (2009) EMBO Reports 10(5), Roman-Hernandez et al., (2009) PNAS 106(22):8888-93, Guo et al., (2002) JBC) See 277(48):46753-62, Wang et al., (2008) Molecular Cell 32:406-414). In some embodiments, amino acid residues corresponding to the ClpS hydrophobic binding pocket identified by Schuenemann et al. are modified to generate a binding site with desired selectivity.

[0112] In one embodiment, the binding portion includes a member of the UBR box recognition sequence family, or a variant of the UBR box recognition sequence family. The UBR recognition boxes are described in Tasaki et al., (2009), JBC 284(3):1884-95. For example, the binding portion may include UBR1, UBR2, or their variants, variants, or homologs.

[0113] In certain embodiments, the binder further comprises one or more detectable labels, such as fluorescent labels, in addition to the binding moiety. In some embodiments, the binder does not contain polynucleotides, such as coding tags. Optionally, the binder comprises synthetic or native antibodies. In some embodiments, the binder comprises an aptamer. In one embodiment, the binder comprises a polypeptide, such as a modified member of the ClpS family of adapter proteins, such as a variant of E. coli ClpS-binding polypeptide, and a detectable label. In one embodiment, the detectable label is optically detectable. In some embodiments, the detectable label comprises a fluorescent moiety, color-coded nanoparticles, quantum dots, or any combination thereof. In one embodiment, the label comprises a polystyrene dye containing a core dye such as FluoSphere®, Nile Red, fluorescein, or rhodamine; a derivatized rhodamine dye such as TAMRA; a phosphor; a polymethazine dye; a fluorescent phosphoramidite; TEXAS RED; green fluorescent protein; acridine; cyanine; cyanine 5 dye; cyanine 3 dye; 5-(2'-aminoethyl)-aminonaphthalene-1-sulfonic acid (EDANS); BODIPY; 120 ALEXA; or any derivative or modification of any of the above. In one embodiment, the detectable label is resistant to photobleaching while producing a large amount of signal (such as photons) at an intrinsic and easily detectable wavelength with a high signal-to-noise ratio.

[0114] In certain embodiments, antikarin is engineered for both high affinity and high specificity to labeled NTAAs (e.g., PTC, modified PTC, Cbz, DNP, SNP, acetyl, guanidinyl, aminoguanidinyl, diheterocyclic metanimine, etc.). Certain types of antikarin scaffolds have a shape suitable for binding single amino acids due to their β-barrel structure. The N-terminal amino acid (with or without modification) can fit into this "β-barrel" bucket and be recognized. Engineered high-affinity antikarin with novel binding activity has been described (review by Skerra, 2008, FEBS J.275:2677-2683). For example, antikarin with high-affinity binding (low nM) to fluorescein and digoxigenin has been designed (Gebauer et al., 2012, Methods Enzymol 503:157-188). The engineering of alternative scaffolding for new bonding functions has also been reviewed by Banta et al. (2013, Annu. Rev. Biomed. Eng. 15:93-113).

[0115] In some embodiments, the binder is directly or indirectly linked to the polymerizing domain. Thus, monomers, dimers, and higher-order (e.g., 3, 4, 5, or more) polymer polypeptides comprising one or more binders are provided herein. In some specific embodiments, the binder is a dimer. In some examples, two polypeptides of the present invention can be attached to each other covalently or non-covalently to form a dimer.

[0116] In some embodiments, the binder is derived from biological, naturally occurring, unnatural, or synthetic sources. In some examples, the binder is derived from novel protein design (Huang et al., (2016) 537(7620):320-327). In some examples, the binder has a structure, sequence, and / or activity designed from a first principle.

[0117] In some embodiments, binders that selectively bind to modified C-terminal amino acids (CTAAs) can be utilized. Carboxypeptidases are proteases that cleave / desorb terminal amino acids containing a free carboxyl group. Many carboxypeptidases exhibit amino acid preference; for example, carboxypeptidase B preferentially cleaves basic amino acids such as arginine and lysine. Carboxypeptidases can be modified to create binders that selectively bind to specific amino acids. In some embodiments, carboxypeptidases can be engineered to selectively bind to both the modified moiety and the α-carbon R group of a CTAA. Thus, engineered carboxypeptidases can specifically recognize 20 different CTAAs representing standard amino acids in the context of C-terminal labeling. Control of stepwise degradation of peptides from the C-terminus is achieved by using engineered carboxypeptidases that are active (e.g., binding activity or catalytic activity) only in the presence of a label. In one example, CTAAs can be modified with p-nitroanilide or 7-amino-4-methylcoumarinyl groups.

[0118] Other potential scaffolds that can be manipulated to produce binders for use in the methods described herein include anticarin, lipocalin, amino acid tRNA synthetase (aaRS), ClpS, Affilin®, Adnectin®, T cell receptor, zinc finger protein, thioredoxin, GST A1-1, DARPin, Affimer, Affitin, Alphabody, Avimer, Monobody, Antibody, Single-domain antibody, Nanobody, EETI-II, HPSTI, Intrabody, PHD-finger, V(NAR)LDTI, Ebibody, Ig(NAR), Nottin, Maxibody, Microbody, Neocardinostatin, pVIII, Tendamistat, VLR, Protein A scaffold, MTI-II, Ecotin, GCN4, Im9, Knitz domain, PBP Examples include transbody, tetranectin, WW domain, CBM4-2, DX-88, GFP, iMab, Ldl receptor domain A, Min-23, PDZ domain, avian pancreatic polypeptide, calybdotoxin / 10Fn3, domain antibody (Dab), a2p8 ankyrin repeat, insect defense A peptide, engineered AR protein, C-type lectin domain, staphylococcal nuclease, Src homology domain 3 (SH3), or Src homology domain 2 (SH2). For example, El-Gebali et al., (2019) Nucleic Acids Research 47:D427-D432 and See Finn et al., (2013) Nucleic Acids Res. 42 (Database issue): D222-D230. In some embodiments, the binder is derived from an enzyme that binds to one or more amino acids (e.g., aminopeptidases). In certain embodiments, the binder may be derived from antikalin or Clp protease adapter protein (ClpS).

[0119] In some cases, the binder can bind to post-translational modified amino acids. In some embodiments, detection of internal post-translational modified amino acids (e.g., phosphorylation, glycosylation, succinylation, ubiquitination, S-nitrosylation, methylation, N-acetylation, lipidation, etc.) is achieved before detection and elimination of terminal amino acids (e.g., NTAA or CTAA). In one example, a peptide is contacted with a binder for PTM modification, and information from the corresponding coding tag is transferred to a recording tag associated with the immobilized peptide. Once the detection and transfer of information related to amino acid modification is complete, the PTM modifying groups can be removed before detecting and transferring the coding tag information of the primary amino acid sequence using an N-terminal or C-terminal decomposition method. Thus, the resulting elongated nucleic acid indicates the presence of post-translational modifications in the peptide sequence, along with the primary amino acid sequence information, but not in order.

[0120] In some embodiments, the detection of internal post-modified amino acids can be performed simultaneously with the detection of the primary amino acid sequence. In one example, NTAA (or CTAA) is contacted with a binder specific to post-modified amino acids, either alone or as part of a binder library (e.g., a library consisting of 20 standard amino acids and binders for selected post-modified amino acids). A continuous cycle of terminal amino acid detachment and contact with the binder (or binder library) follows. Thus, the resulting elongated nucleic acid on a recording tag associated with the immobilized peptide indicates the presence and order of post-modifications in the context of the primary amino acid sequence.

[0121] In certain embodiments, polymers, such as polypeptides, are also contacted with non-homogeneous binders. As used herein, a non-homogeneous binder refers to a binder that is selective for a target different from the particular target under consideration (e.g., a feature or component of the polypeptide). For example, if the nNTAA is phenylalanine and the peptide is contacted with three binders selective for phenylalanine, tyrosine, and asparagine, the phenylalanine-selective binder becomes the first binder that can selectively bind to the nth NTAA (i.e., phenylalanine), while the other two binders become non-homogeneous binders of the peptide (because they are selective for NTAAs other than phenylalanine). However, the tyrosine and asparagine binders may be homogeneous binders of other peptides in the sample. Next, the nNTAA (phenylalanine) is cleaved from the peptide, thereby converting the n-1 amino acid of the peptide to an n-1 NTAA (e.g., tyrosine). Then, when the peptide is brought into contact with the same three binders, the binder selective for tyrosine becomes a second binder that can selectively bind to the n-1 NTAA (i.e., tyrosine), while the other two binders become non-homogeneous binders (because they are selective for NTAAs other than tyrosine).

[0122] Therefore, it should be understood that whether a drug is a binder or a non-homogeneous binder depends on the characteristics or properties of the specific polypeptide currently available for binding. Furthermore, when multiple polypeptides are analyzed in a multiplexing reaction, a binder for one polypeptide may be a non-homogeneous binder for another polypeptide, and vice versa. Therefore, it should be understood that the following description of binders is applicable to any type of binder described herein (i.e., both homogeneous and non-homogeneous binders).

[0123] In certain embodiments, the concentration of the binder in the solution is controlled to reduce the background and / or false-positive results of the assay. In some embodiments, the concentration of the binder may be any preferred concentration, e.g., about 0.0001 nM, about 0.001 nM, about 0.01 nM, about 0.1 nM, about 1 nM, about 2 nM, about 5 nM, about 10 nM, about 20 nM, about 50 nM, about 100 nM, about 200 nM, about 500 nM, or about 1000 nM. In other embodiments, the concentration of the soluble conjugate used in the assay is between approximately 0.0001 nM and approximately 0.001 nM, between approximately 0.001 nM and approximately 0.01 nM, between approximately 0.01 nM and approximately 0.1 nM, between approximately 0.1 nM and approximately 1 nM, between approximately 1 nM and approximately 2 nM, between approximately 2 nM and approximately 5 nM, between approximately 5 nM and approximately 10 nM, between approximately 10 nM and approximately 20 nM, between approximately 20 nM and approximately 50 nM, between approximately 50 nM and approximately 100 nM, between approximately 100 nM and approximately 200 nM, between approximately 200 nM and approximately 500 nM, between approximately 500 nM and approximately 1000 nM, or greater than approximately 1000 nM.

[0124] In some embodiments, the ratio of the soluble binder molecule to the immobilized polypeptide, e.g., polypeptide, is within any preferred range, e.g., about 0.00001:1, about 0.0001:1, about 0.001:1, about 0.01:1, about 0.1:1, about 1:1, about 2:1, about 5:1, about 10:1, about 15:1, about 20:1, about 25:1, about 30:1, about 35:1, about 40:1, about 45:1, about 50:1, about 55:1, about 60:1, about 65:1, about 70:1, about 75:1, about 80:1, about 85:1, about 90:1, about 95:1, about 100:1, about 10 4 :1, about 10 5 :1, about 10 6 The ratio may be 1, or more, or any ratio between the above ratios. A higher ratio of soluble binder molecules to immobilized polypeptides and / or nucleic acids can be used to complete the transfer of binding and / or coding tag information. This may be particularly useful for detecting and / or analyzing small amounts of polypeptides in a sample.

[0125] C. Coding Tags The described binders include, or are associated with, a coding tag containing identification information relating to the binder (e.g., representing or correlated with the binder). In some embodiments, the identification information from the coding tag includes information about the identity of the target bound by the binder. In some embodiments, the identification information from the coding tag includes information about the identity of one or more amino acids on the peptide bound by the binder, or are associated with it. In some cases, the coding tag includes a partial restriction enzyme recognition sequence used in the cleavage step. For example, the coding tag may include a sequence (or a portion thereof) that can be recognized by a double-stranded nucleic acid cleavage reagent (e.g., restriction enzyme). In some cases, the coding tag may include a single-stranded sequence that can be recognized by a double-stranded nucleic acid cleavage reagent (e.g., restriction enzyme).

[0126] A coding tag associated with a binder is a nucleic acid molecule of approximately 2 to 100 bases, containing any suitable length of polynucleotide, e.g., any integer between 2 and 100 (including 2 and 100), and containing identification information about the associated binder. A “coding tag” can also be fabricated from a “sequencing polymer” (see, e.g., Niu et al., 2013, Nat. Chem. 5:282-292; Roy et al., 2015, Nat. Commun. 6:7237, Lutz 2015, Macromolecules 48:4759-4767, each of which is incorporated in whole by reference). The coding tag may contain an encoder sequence or a sequence containing identification information, which may optionally have one spacer adjacent to one side or optionally have spacers adjacent to both sides. The coding tag may also consist of an optional UMI and / or an optional binding cycle-specific barcode. The coding tag may refer to a coding tag directly attached to the binder, a complementary sequence hybridized to a coding tag directly attached to the binder (for example, in the case of a double-stranded coding tag), or coding tag information present in the extended nucleic acid on the recording tag. In certain embodiments, the coding tag may further include a binding cycle-specific spacer or barcode, a unique molecular identifier, a universal priming site, or any combination thereof.

[0127] The coding tag may be a single-stranded molecule, a double-stranded molecule, or partially double-stranded. The coding tag may include a blunt end, a protruding end, or one of each. In some embodiments, the coding tag is partially double-stranded, which prevents annealing of the coding tag to the internal encoder and spacer array within the growing extension recording tag. In some embodiments, the coding tag may include a hairpin. In certain embodiments, the hairpin includes mutually complementary nucleic acid regions connected via nucleic acid strands. In some embodiments, the nucleic acid hairpin may also further include 3' and / or 5' single-stranded regions extending from the double-stranded stem segment. In some examples, the hairpin contains a single-stranded nucleic acid.

[0128] In some embodiments, the described binder includes a coding tag containing identification information relating to the binder. In some embodiments, the identification information from the coding tag includes information relating to the identity of the target bound by the binder. In some embodiments, the identification information from the coding tag includes information relating to the identity of one or more amino acids on the peptide bound by the binder.

[0129] A coding tag is a nucleic acid molecule of approximately 3 to 100 bases that provides identification information specific to its associated binder. Coding tags may include approximately 3 to 90 bases, approximately 3 to 80 bases, approximately 3 to 70 bases, approximately 3 to 60 bases, approximately 3 to 50 bases, approximately 3 to 40 bases, approximately 3 to 30 bases, approximately 3 to 20 bases, approximately 3 to 10 bases, or approximately 3 to 8 bases. In some embodiments, coding tags are approximately 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 25, 30, 35, 40, 55, 60, 65, 70, 75, 80, 85, 90, 95, or 100 bases in length. Coding tags may consist of DNA, RNA, polynucleotide analogs, or combinations thereof. Examples of polynucleotide analogs include PNA, gPNA, BNA, GNA, TNA, LNA, morpholinopolynucleotides, 2'-O-methylpolynucleotides, alkylribosyl-substituted polynucleotides, phosphorothioate polynucleotides, and 7-deazapurine analogs.

[0130] The binder can be attached to the coding tag in any preferred way and position so as not to disrupt the enzymatic reaction of the provided method (e.g., the function of the nucleic acid conjugation reagent, polymerase, and double-strand nucleic acid cleavage reagent). For example, the binder can be attached to the coding tag in any preferred way, as the coding tag is attached in a way that gives the coding tag an available 3' end. In some specific embodiments, the coding tag has a spacer which is a 3' overhang. In some embodiments, the binder is attached to the coding tag at the 5' end of the coding tag. In some embodiments, the binder is attached to the coding tag in the loop region. In some embodiments, the binder is attached to the coding tag at a position near the loop at the 5' end of the recoding tag. The coding tag can be conjugated to the binder directly or indirectly (e.g., via a linker) by any means known in the art, including covalent and non-covalent interactions. In some embodiments, the coding tag can be conjugated to the binder enzymatically or chemically. In some embodiments, the coding tag can be conjugated to the binder via ligation. In other embodiments, the coding tag is attached to the binder via an affinity binding pair (e.g., biotin and streptavidin). In some cases, the coding tag may be attached to the binder to a non-natural amino acid, such as via a covalent interaction with the non-natural amino acid.

[0131] In some embodiments, the binder is conjugated to the coding tag via a SpyCatcher-SpyTag interaction. The SpyTag peptide forms an irreversible covalent bond to the SpyCatcher protein via spontaneous isopeptide linkage, thereby providing a genetically encoded method that results in a peptide interaction resistant to force and harsh conditions (Zakeri et al., 2012, Proc. Natl. Acad. Sci. 109:E690-697, Li et al., 2014, J. Mol. Biol. 426:309-317). The binder may be expressed as a fusion protein containing the SpyCatcher protein. In some embodiments, the SpyCatcher protein is attached to the N-terminus or C-terminus of the binder. The SpyTag peptide may be coupled to the coding tag using standard conjugation chemistry (Hermanson, Bioconjugate Techniques, (2013) Academic Press).

[0132] In some embodiments, enzyme-based strategies are used to conjugate the binder to the coding tag. For example, the binder may be conjugated to the coding tag using formylglycine (FGly)-producing enzyme (FGE). In one example, a protein, e.g., SpyLigase, is used to conjugate the binder to the coding tag (Fierer et al.). al., Proc Natl Acad Sci USA.2014;111(13):E1176-E1181).

[0133] In other embodiments, the binder is conjugated to the coding tag via a SnoopTag-SnoopCatcher peptide-protein interaction. The SnoopTag peptide forms an isopeptide bond with the SnoopCatcher protein (Veggiani et al., Proc. Natl. Acad. Sci. USA, 2016, 113:1202-1207). The binder may be expressed as a fusion protein containing the SnoopCatcher protein. In some embodiments, the SnoopCatcher protein is added to the N-terminus or C-terminus of the binder. The SnoopTag peptide may be coupled to the coding tag using standard conjugation chemistry.

[0134] In yet another embodiment, the binder is conjugated to the coding tag via a HaloTag® protein fusion tag and its chemical ligand. HaloTag is a modified haloalkane dehalogenase designed to covalently bond to a synthetic ligand (HaloTag ligand) (Los et al., 2008, ACS Chem. Biol. 3:373-382). The synthetic ligand includes chloroalkane linkers attached to various useful molecules. A covalent bond is formed between HaloTag and the chloroalkane linker, which is highly specific, occurs rapidly under physiological conditions, and is essentially irreversible.

[0135] In some cases, the binder is attached to the coding tag by an enzyme, such as a saltase-mediated label (see, e.g., Antos et al., Curr Protoc Protein Sci. (2009) CHAPTER 15:Unit-15.3, International Patent Publication No. 2013 / 003555). The saltase enzyme catalyzes the transpeptide reaction (see, e.g., Falck et al, Antibodies (2018) 7(4):1-19). In some embodiments, the binder is modified with or conjugated to one or more N-terminal or C-terminal glycine residues.

[0136] In some embodiments, the binder is conjugated to the coding tag using a cysteine ​​bioconjugation method. In some embodiments, the binder is conjugated to the coding tag using π-clamp-mediated cysteine ​​bioconjugation (see, e.g., Zhang et al., Nat Chem. (2016) 8(2):120-128). In some cases, the binder is conjugated to the coding tag using 3-arylpropioronitrile (APN)-mediated tagging (see, e.g., Koniev et al., Bioconjug Chem. 2014;25(2):202-206).

[0137] In some embodiments, the set of coding tags used in a binding and information transfer cycle may include cycle information, such as using cycle-specific sequences. In one embodiment, the coding tag includes a binding cycle-specific sequence. In some embodiments, coding tags in a collection of binders share a common spacer sequence used in the assay (for example, an entire library of binders used in multiple binding cycle methods has a common spacer for their coding tags). In another embodiment, the coding tag consists of binding cycle tags that identify a particular binding cycle. In yet another embodiment, the coding tags in a library of binders have a binding cycle-specific spacer sequence. In some embodiments, the coding tag includes one binding cycle-specific spacer sequence. For example, the coding tag of a binder used in a first binding cycle includes a "cycle 1" specific spacer sequence, the coding tag of a binder used in a second binding cycle includes a "cycle 2" specific spacer sequence, and so on, up to "n" binding cycles. In further embodiments, the binding agent coding tag used in the first binding cycle includes a "cycle 1" specific spacer sequence and a "cycle 2" specific spacer sequence, and the binding agent coding tag used in the second binding cycle includes a "cycle 2" specific spacer sequence and a "cycle 3" specific spacer sequence, and so on, with up to "n" binding cycles. This embodiment is useful for subsequent PCR assembly of an unbound extension record tag after the binding cycles are complete. In some embodiments, the spacer sequence contains a sufficient number of bases to anneal to a complementary spacer sequence in the record tag or extension record tag in order to initiate a primer extension reaction or a sticky-end ligation reaction.

[0138] Cycle-specific spacer sequences can also be designed so that information transfer is conditional. For example, a first binding cycle can transfer information from a coding tag to a recording tag, and subsequent binding cycles can transfer information in a manner that depends on previously added spacer sequences. More specifically, the binding agent coding tag used in the first binding cycle includes a "cycle 1" specific spacer sequence and a "cycle 2" specific spacer sequence, and the binding agent coding tag used in the second binding cycle includes a "cycle 2" specific spacer sequence and a "cycle 3" specific spacer sequence, and so on, including up to "n" binding cycles. The binding agent coding tag from the first binding cycle can be annealed to the recording tag via a complementary cycle 1 specific spacer sequence. Once the coding tag information is transferred to the recording tag, the cycle 2 specific spacer sequence is positioned at the 3' end of the extended recording tag at the end of binding cycle 1. The binding agent coding tag from the second binding cycle can be annealed to the extended recording tag via a complementary cycle 2 specific spacer sequence. When coding tag information is transferred to the extended record tag, the cycle 3 specific spacer sequence is positioned at the 3' end of the extended record tag at the end of binding cycle 2, through "n" binding cycles. This embodiment provides that the transfer of binding information in a particular binding cycle between multiple binding cycles occurs only on the (extended) record tag that has experienced the previous binding cycle. In some cases, if the spacers added by the previous coding tag match at least some of the spacers of the second coding tag, the information is transferred from the second (or higher-order) coding tag to the extended record tag.

[0139] In some embodiments, the binder may not be able to bind to the homogeneous polymer. A "tracking" step can be used after each binding cycle, using an oligonucleotide containing a binding cycle-specific spacer, to keep the binding cycles synchronized even if a binding cycle fails. For example, if a homogeneous binder fails to bind to the polymer during binding cycle 1, a tracking step is added after binding cycle 1 using an oligonucleotide containing both a cycle 1-specific spacer, a cycle 2-specific spacer, and a "null" encoder sequence. The "null" encoder sequence may be an encoder sequence, or preferably a specific barcode that clearly identifies the "null" binding cycle. The "null" oligonucleotide can be annealed to the record tag via the cycle 1-specific spacer, and the cycle 2-specific spacer is transferred to the record tag. Thus, the binder from binding cycle 2 can be annealed to the extension record tag via the cycle 2-specific spacer despite the failure of the binding cycle 1 event. The "null" oligonucleotide marks binding cycle 1 as a failure of the binding event in the extension record tag.

[0140] In a preferred embodiment, a coupled cycle-specific encoder sequence is used for the coding tag. The coupled cycle-specific encoder sequence can be achieved by using a completely unique analyte (e.g., NTAA) coupled cycle encoder barcode, or by using a combination of an analyte (e.g., NTAA) encoder sequence bonded to a cycle-specific barcode. The advantage of using the combined approach is that the total number of barcodes that need to be designed is reduced. For a set of 20 analyte binders used over 10 cycles, only 20 analyte encoder sequence barcodes and 10 coupled cycle-specific barcodes need to be designed. In contrast, if the coupled cycles are directly embedded in the binder encoder sequence, a total of 200 independent encoder barcodes may need to be designed. The advantage of directly embedding coupled cycle information in the encoder sequence is that the overall length of the coding tag can be minimized when using error-corrected barcodes for nanopore readout. Using error-tolerant barcodes allows for the identification of high-precision barcodes using sequencing platforms and approaches that are more error-prone but offer other advantages such as faster analysis speeds, lower costs, and / or greater instrument portability. One such example is nanopore-based sequencing reads.

[0141] II. Polymer Analysis Assays Methods provided for the analysis of polymers, such as peptides, polypeptides, and proteins, which include the step of transferring information using a sequential encoding method described on a record tag, may include additional steps, processes, and reactions. In some embodiments, the polymer is a polypeptide, and a polypeptide analytical assay is performed. In some embodiments, the identity of the sequence (or a portion thereof) and / or protein is determined using a polypeptide analytical assay. In some examples, the polypeptide analytical assay includes evaluating at least a partial sequence or identity of the polypeptide using a preferred technique or procedure. For example, at least a partial sequence of the polypeptide can be evaluated by N-terminal amino acid analysis or C-terminal amino acid analysis. In some embodiments, at least a partial sequence of the polypeptide can be evaluated using a ProteoCode assay. In some cases, at least a partial sequence of a polypeptide can be evaluated by applying some of the techniques or procedures disclosed and / or claimed in U.S. Patent Publications 2019 / 0145982A1, 2020 / 0348308A1, 2020 / 0348307A1, and 2021 / 0208150A1.

[0142] In embodiments of the method for analyzing peptides or polypeptides, the method generally involves contacting a binder with at least one terminal amino acid (e.g., NTAA or CTAA) of the polypeptide, protein, or peptide, and upon contact, transferring information from a coding tag to a record tag associated with the polypeptide, protein, or peptide, thereby generating a primary extension record tag (encoding cycle). An exemplary encoding cycle by the method described herein is shown in Figure 1B. The binder shown in Figure 1B is attached by a linker to a coding tag containing identification information relating to the binder (binder barcode, BBC). In some embodiments, the binder may be attached to or conjugated to the coding tag in locations not depicted (e.g., the coding tag or other loop regions). For transferring the BBC information, polymerases, nucleic acid conjugation reagents (e.g., DNA ligases), and double-strand nucleic acid cleavage reagents (e.g., restriction enzymes) are provided. These reagents may then be added either as one enzyme at a time, or as two enzymes (polymerase and DNA ligase) followed by a restriction enzyme, or as a mixture of three enzymes. When the 5' end of the recording tag is ligated to the 3' end of the coding tag, polymerase extends the 3' (unligated) end of the recording tag to create a dsDNA molecule containing a 2-base pair spacer adjacent to each restriction enzyme site. After double-strand formation, the restriction enzyme binds and cleaves adjacent to its recognition site. After cleavage, the dsDNA on the recording tag contains a binder-specific barcode (BBC) and a 2nt3' overhang that functions as a spacer sequence in the next round of encoding. In some embodiments, after the information transfer (encoding) cycle, a portion of the polypeptide for analysis may be removed from the polypeptide (e.g., terminal amino acids or terminal dipeptides). The cycle of steps shown in Figure 1B may be repeated one or more times with additional binders and corresponding coding tags to further extend the recording tag.

[0143] In some further embodiments, the method includes labeling or modifying the polymer (e.g., a peptide) before or after contact with a binder. For example, the terminal amino acids of the polypeptide, protein, or peptide bound by the binder may be chemically labeled or modified terminal amino acids. In some further embodiments, the method further includes removing or eliminating terminal amino acids (e.g., NTAA or CTAA) from the polypeptide, protein, or peptide after the information transfer step. The terminal amino acids to be eliminated may be chemically labeled or modified terminal amino acids. Removal of NTAA by contact with an enzyme or chemical reagent converts the second-to-last amino acid of the polypeptide, protein, or peptide to a terminal amino acid. Polypeptide analysis may include one or more cycles of binding additional binders to terminal amino acids, transferring information from coding tags to an elongated nucleic acid, thereby generating a higher-order elongation record tag containing information about two or more binders, and periodically deleting terminal amino acids. Additional binding, information transfer, and removal occur with up to n amino acids as described above, generating an n-order elongated nucleic acid that collectively represents the polypeptide, protein, or peptide. Recording tags are used to record information collected from one or more binding events between one or more binders and the polymer being analyzed.

[0144] In some of the provided embodiments, the step involving NTAA in the described exemplary approach may instead be performed using a C-terminal amino acid (CTAA).

[0145] In some embodiments, the order of steps in the peptide or polypeptide sequencing assay process by degradation can be reversed or performed in various orders. For example, in some embodiments, terminal amino acid labeling can be performed before and / or after binding the polypeptide to the binder.

[0146] Provided herein is a method for analyzing a polymer, comprising: (a) providing a polymer and an associated recording tag conjugated to a support; (b) contacting the polymer with a binder capable of binding to the polymer and comprising a coding tag having identification information relating to the binder, thereby enabling binding between the polymer and the binder; (c) conjugating the 5' end of the recording tag to the 3' end of the coding tag with a nucleic acid conjugation reagent; (d) extending the recording tag using the coding tag as a template with a polymerase to produce a double-stranded extended recording tag; and (e) cleaving the double-stranded extended recording tag with a double-stranded nucleic acid cleavage reagent to produce a 3' overhang within the extended recording tag. In some cases, the binder is removed after step (e). In some embodiments, the method further includes analyzing the extended recoding tag. In some embodiments, the method further includes adding a universal priming site to the extended recording tag before analyzing the extended recording tag.

[0147] In some examples, step (a) is performed before steps (b), (c), (d), and (e). Steps (c), (d), and (e) can be performed in a stepwise manner by the nature of the design, although the reagents can be provided in a mixture. In some embodiments, step (b) is performed before steps (c) and (d), for example, by first providing the nucleic acid conjugation reagent and then providing the polymerase and double-strand nucleic acid cleavage reagent together. In some cases, the steps occur in the order of (c), (d), and (e). In some specific embodiments, the steps are performed in the order of (a), (b), (c), (d), and (e), with the optional repetition of steps (b), (c), (d), and (e) one or more times.

[0148] In some embodiments, the method further includes the step of removing a portion of the polypeptide, such as the terminal amino acids of the polypeptide, protein, or peptide (e.g., the N-terminal amino acid (NTAA)) to expose new terminal amino acids of the polypeptide, protein, or peptide. In some cases, a second cycle of steps (b), (c), (d), and (e) is repeated after the first cycle of steps (b), (c), (d), and (e) has been removed, after at least a portion of the polypeptide has been removed.

[0149] In some embodiments, the method involves treating a target polypeptide, protein, or peptide with a reagent for modifying the terminal amino acids of the polypeptide, protein, or peptide. In some embodiments, the reagent for modifying the terminal amino acids of the polypeptide includes a chemical agent or an enzymatic agent. In some embodiments, the target polypeptide, protein, or peptide is contacted with the reagent for modifying the terminal amino acids before step (b). In some embodiments, the target polypeptide, protein, or peptide is contacted with the reagent for modifying the terminal amino acids before removing the terminal amino acids.

[0150] In some embodiments, the method further includes removing the binder after transferring information from the coding AG to the recording tag. In some embodiments, the binder is removed after step (e). In some embodiments, the binder is removed before repeating step (b). In some embodiments, the binder removal is performed after transferring information from the coding tag to a recording tag associated with the polymer for analysis.

[0151] A. Samples and polymers In some embodiments, the analytical assay is performed against one or more polymers of unknown identity obtained from a sample. In some cases, the polymer is a mixture of molecules obtained from the sample. The polymer may be a large molecule composed of smaller subunits. In certain embodiments, the polymer is a protein, protein complex, polypeptide, peptide, nucleic acid molecule, carbohydrate, lipid, macrocyclic molecule, or chimeric polymer. The polymers (e.g., proteins, polypeptides, peptides) in the methods disclosed herein can be obtained from any suitable source or sample.

[0152] The methods disclosed herein can be used for analysis including the simultaneous (multiplexed) detection, identification, quantification, and / or sequencing of multiple polymers. As used herein, multiplexing refers to the analysis of multiple polymers (e.g., polypeptides) in the same assay. Multiple polymers may originate from the same sample or different samples. Multiple polymers may originate from the same or different subjects. Multiple polymers analyzed may be different polymers or the same polymer originating from different samples. Multiple polymers include two or more polymers, five or more polymers, ten or more polymers, fifty or more polymers, one hundred or more polymers, five hundred or more polymers, one thousand or more polymers, five thousand or more polymers, one hundred thousand or more polymers, one thousand thousand or more polymers, one thousand thousand or more polymers, one thousand thousand or more polymers, one thousand thousand or more polymers, one thousand thousand or more polymers, one thousand thousand or more polymers, one thousand thousand or more polymers, one thousand thousand or more polymers, or one thousand thousand or more polymers.

[0153] In some embodiments, macromolecules (e.g., proteins, polypeptides, or peptides) are obtained from a sample that is a biological sample. In some embodiments, the sample includes, but is not limited to, mammalian or human cells, yeast cells, and / or bacterial cells. In some embodiments, the sample includes cells derived from a sample obtained from a multicellular organism. For example, the sample may be isolated from an individual. In some embodiments, the sample may include a single cell type or multiple cell types. In some embodiments, the sample may be obtained from a mammalian organism or a human, for example, by puncture or other collection or sampling procedure. In some embodiments, the sample includes two or more cells.

[0154] In some embodiments, the biological sample may include whole cells and / or living cells and / or cell debris. In some examples, suitable sources or samples may include, but are not limited to, biological samples of virtually any organism, such as biopsy samples, cell cultures, cells (both primary cells and cultured cell lines), samples containing organelles or vesicles, tissues, and tissue extracts. For example, suitable sources or samples include: biopsies; fecal matter; body fluids (blood, whole blood, serum, plasma, urine, lymph, bile, aqueous humor, breast milk, earwax, chyle, atherosclerotic syrup, endolymph, perilymph, exudate, cerebrospinal fluid, interstitial fluid, aqueous or vitreous fluid, colostrum, sputum, amniotic fluid, saliva, anal and vaginal secretions, gastric acid, gastric juice, lymph, mucus (including nasal discharge and sputum), pericardial fluid, ascites, pleural fluid, pus, mucosal secretions, saliva, sebum (skin oil), sputum, synovial fluid, sweat and semen, transudate, vomit and mixtures thereof; exudate (e.g., fluid obtained from an abscess or any other site of infection or inflammation); mammalian-derived samples, preferably containing microbiome-containing samples; and human-derived samples, particularly preferably containing microbiome-containing samples. Examples of samples that may be used include, but are not limited to, fluids obtained from the joints of virtually all organisms (normal joints or joints affected by diseases such as rheumatoid arthritis, osteoarthritis, gout, or septic arthritis); environmental samples (air, agricultural, water, and soil samples); microbial samples including microbial biofilms and / or microbial communities, as well as samples derived from bacterial spores; tissue samples including tissue sections; research samples including extracellular fluid; extracellular supernatants from cell cultures; and cellular components including inclusion bodies within bacteria, mitochondria, and cellular periplasm. In some embodiments, biological samples include or are derived from body fluids, and the body fluids are obtained from mammals or humans. In some embodiments, samples include body fluids or cell cultures derived from body fluids.

[0155] In some embodiments, macromolecules (e.g., polypeptides and proteins) may be obtained and prepared from a single cell type or multiple cell types. In some embodiments, the sample comprises a population of cells. In some embodiments, the macromolecules (e.g., proteins, polypeptides, or peptides) may originate from cells or intracellular components, extracellular vesicles, organelles, or organized minor components thereof. The macromolecules (e.g., proteins, polypeptides, or peptides) may originate from organelles, such as mitochondria, nuclei, or vesicles. In one embodiment, one or more specific types of single cells or their subtypes may be isolated. In some embodiments, the sample may, but is not limited to, organelles (e.g., nucleus, Golgi apparatus, ribosomes, mitochondria, endoplasmic reticulum, chloroplasts, cell membranes, vesicles, etc.).

[0156] In certain embodiments, the macromolecule is or comprises a protein, protein complex, polypeptide, or peptide. Amino acid sequence information and post-translational modifications of the peptide, polypeptide, or protein are transduced into a nucleic acid-encoded library that can be analyzed via next-generation sequencing. The peptide may comprise L-amino acids, D-amino acids, or both. The peptide, polypeptide, protein, or protein complex may comprise standard naturally occurring amino acids, modified amino acids (e.g., post-translational modifications), amino acid analogs, amino acid mimes, or any combination thereof. In some embodiments, the peptide, polypeptide, or protein may be naturally occurring, synthetically produced, or recombinantly expressed. In any of the peptide embodiments described above, the peptide, polypeptide, protein, or protein complex may further comprise post-translational modifications. Standard natural amino acids include alanine (A or Ala), cysteine ​​(C or Cys), aspartic acid (D or Asp), glutamic acid (E or Glu), phenylalanine (F or Phe), glycine (G or Gly), histidine (H or His), isoleucine (I or Ile), lysine (K or Lys), leucine (L or Leu), methionine (M or Met), asparagine (N or Asn), proline (P or Pro), glutamine (Q or Gln), arginine (R or Arg), serine (S or Ser), threonine (T or Thr), valine (V or Val), tryptophan (W or Trp), and tyrosine (Y or Tyr). Non-standard amino acids include selenocysteine, pyrrolicin, and N-formylmethionine, β-amino acids, homoamino acids, proline and pyruvate derivatives, 3-substituted alanine derivatives, glycine derivatives, ring-substituted phenylalanine and tyrosine derivatives, linear core amino acids, and N-methyl amino acids.

[0157] Post-translational modifications (PTMs) of peptides, polypeptides, or proteins can be covalent or enzymatic modifications. Examples of PTMs include acylation, acetylation, alkylation (including methylation), biotinylation, butyrylation, carbamylation, carbonylation, deamidation, deiminoation, diphthamide formation, disulfide crosslinking, eliminylation, flavin attachment, formylation, gammacarboxylation, glutamylation, glycylation, glycosylation (e.g., N-linked, O-linked, C-linked phosphoglycosylation), glycietion, heme C attachment, and hydroxylation. Post-translational modifications include, but are not limited to, lylation, hypsination, iodation, isoprenylation, lipidation, lipoylation, malonylation, methylation, myristolylation, oxidation, palmitoylation, pegylation, phosphopantetheination, phosphorylation, propionylation, retinilidenci base formation, S-glutathionylation, S-nitrosylation, S-sulfenylation, selenization, succinylation, sulfination, ubiquitination, and C-terminal amidation. Post-translational modifications include modifications of the amino terminus and / or carboxyl terminus of peptides, polypeptides, or proteins. Modifications of terminal amino groups include, but are not limited to, desamino, N-lower alkyl, N-di-lower alkyl, and N-acyl modifications. Modifications of terminal carboxyl groups include, but are not limited to, amide, lower alkylamide, dialkylamide, and lower alkyl ester modifications (for example, lower alkyls are C1-C4 alkyls). Post-translational modifications include modifications of amino acids between the amino and carboxyl terms of peptides, polypeptides, or proteins (for example, those mentioned above, but not limited to those). Post-translational modifications can regulate the "biology" of intracellular proteins, such as their activity, structure, stability, or localization. For example, phosphorylation plays an important role in protein regulation, particularly in cellular signal transduction (Prabakaran et al.). (al., 2012, Wiley Interdiscip Rev Syst Biol Med 4:565-583). In another example, the addition of sugars to proteins, such as glycosylation, has been shown to promote protein folding, improve stability, alter regulatory functions, and target cell membranes by allowing lipids to attach to proteins. Post-translational modifications may also include peptide, polypeptide, or protein modifications to include one or more detectable labels.

[0158] In certain embodiments, peptides, polypeptides, or proteins can be fragmented. Peptides, polypeptides, or proteins can be fragmented by any means known in the art, including fragmentation by proteases or endopeptidases. In some embodiments, the fragmentation of peptides, polypeptides, or proteins is targeted by the use of specific proteases or endopeptidases. These specific proteases or endopeptidases bind to and cleave specific consensus sequences (e.g., TEV proteases). In other embodiments, the fragmentation of peptides, polypeptides, or proteins is not targeted by or randomly performed by nonspecific proteases or endopeptidases. Nonspecific proteases may bind to and cleave specific amino acid residues rather than consensus sequences (e.g., proteinase K is a nonspecific serine protease). In some embodiments, proteinases and endopeptidases known in the art and that can be used to cleave proteins or polypeptides into smaller peptide fragments include proteinase K, trypsin, chymotrypsin, pepsin, thermolysin, thrombin, factor Xa, furin, endopeptidase, papain, pepsin, subtilisin, elastase, enterokinase, Genenase® I, endoproteinase LysC, endoproteinase AspN, and endoproteinase GluC (Granvogl et al., 2007 Anal Bioanal Chem 389:991-1002). In certain embodiments, peptides, polypeptides, or proteins are fragmented by proteinase K, or optionally, a thermally unstable version of proteinase K, to allow for rapid inactivation. In some cases, proteinase K is stable with denaturing reagents such as urea and SDS, allowing for the digestion of completely denatured proteins. The fragmentation of proteins and polypeptides into peptides can be performed before or after the attachment of DNA tags or DNA recording tags.

[0159] Proteins can also be digested into peptide fragments using chemical reagents. These reagents may cleave at specific amino acid residues (for example, cyanogen bromide hydrolyzes the peptide bond at the C-terminus of a methionine residue). Examples of chemical reagents for fragmenting polypeptides or proteins into smaller peptides include cyanogen bromide (CNBr), hydroxylamine, hydrazine, formic acid, BNPS-skatole [2-(2-nitrophenylsulfenyl)-3-methylindole], iodosobenzoic acid, and ·NTCB+Ni(2-nitro-5-thiocyanobenzoic acid).

[0160] In a particular embodiment, following enzymatic or chemical cleavage, the resulting peptide fragments are of approximately the same desired length, for example, about 10 to 70 amino acids, about 10 to 60 amino acids, about 10 to 50 amino acids, about 10 to 40 amino acids, about 10 to 30 amino acids, about 20 to 70 amino acids, about 20 to 60 amino acids, about 20 to 50 amino acids, about 20 to 40 amino acids, about 20 to 30 amino acids, about 30 to 70 amino acids, about 30 to 60 amino acids, about 30 to 50 amino acids, or about 30 to 40 amino acids. The cleavage reaction can preferably be monitored in real time by spiking a protein or polypeptide sample with a short test FRET (fluorescence resonance energy transfer) peptide containing a peptide sequence with a proteinase or endopeptidase cleavage site. In intact FRET peptides, a fluorescent group and a quenching group are attached to one of the ends of the peptide sequence containing the cleavage site. Fluorescence is reduced due to fluorescence resonance energy transfer between the quenching agent and the fluorophore. When the test peptide is cleaved by a protease or endopeptidase, the quenching agent and fluorophore are separated, and fluorescence increases significantly. The cleavage reaction can be stopped when a specific fluorescence intensity is achieved, allowing for the attainment of a reproducible cleavage endpoint.

[0161] In some embodiments, a sample of macromolecules (e.g., peptides, polypeptides, or proteins) can undergo protein fractionation, in which proteins or peptides are separated by one or more properties such as cellular location, molecular weight, hydrophobicity, isoelectric point, or protein concentration. In some embodiments, a subset of macromolecules (e.g., proteins) in a sample is fractionated so that the subset of macromolecules is sorted from the rest of the sample. For example, a sample can undergo fractionation before being attached to a support. Alternatively, protein concentration can be used to select a specific protein or peptide (see, e.g., Whiteaker et al., (2007) Anal. Biochem. 362:44-54, the whole being incorporated by reference), or to select a specific post-translational modification (see, e.g., Huang et al., 2014. J. Chromatogr. A 1372:1-17, the whole being incorporated by reference). Alternatively, one or more specific classes of proteins, such as immunoglobulins, or immunoglobulin (Ig) isotypes, such as IgG, can be affinity-concentrated or selected for analysis. In the case of immunoglobulin molecules, analysis of the sequence and abundance or frequency of hypervariable sequences involved in affinity binding is particularly interesting because they change in response to disease progression or correlate with health, immunity, and / or disease phenotypes. Excessively abundant proteins can also be subtracted from samples using standard immunoaffinity methods. Removal of abundant proteins can be beneficial for plasma samples where more than 80% of the protein components are albumin and immunoglobulins. Several commercially available products are available for removing plasma samples with excessively abundant proteins, including depletion spin columns that remove the top 2–20 plasma proteins (Pierce, Agilent), or PROTIA and PROT20 (Sigma-Aldrich).

[0162] In certain embodiments, the dynamic range of a protein sample can be controlled by fractionating the protein sample using standard fractionation methods including electrophoresis and liquid chromatography (Zhou et al., 2012, Anal Chem 84(2):720-734), or by distributing the fractions into compartments (e.g., droplets) filled with a limited volume of protein-binding beads / resin (e.g., hydroxylated silica particles) (McCormick, 1989, Anal Biochem 181(1):66-74) to elute the bound proteins. Excess protein from each compartmentalized fraction is washed away. Examples of electrophoretic methods include capillary electrophoresis (CE), capillary isoelectrophoresis (CIEF), capillary isoelectrophoresis (CITP), free-flow electrophoresis, and gel eluate fractionation embedded electrophoresis (GELFrEE). Examples of liquid chromatography protein separation methods include reversed-phase chromatography (RP), ion exchange (IE), size exclusion (SE), and hydrophilic interactions. Examples of compartment partitions include emulsions, droplets, microwells, and physically separated regions on a flat substrate. Exemplary protein-binding beads / resins include silica nanoparticles induced by phenolic or hydroxyl groups (e.g., StrataClean Resin from Agilent Technologies, RapidClean from LabTech, etc.). By limiting the binding capacity of the beads / resin, highly abundant proteins eluting a given fraction are only partially bound to the beads, and excess proteins are removed.

[0163] In some embodiments, partition barcodes are used, which involve assigning unique barcodes to subsamplings of polymers from a population of polymers in a sample. These partition barcodes may consist of identical barcodes resulting from partitioning of polymers within compartments labeled with the same barcode (e.g., a barcoded bead population where multiple beads share the same barcode). The use of physical compartments effectively subsamples the original sample to provide partition barcode assignments. For example, a set of beads labeled with 10,000 different partition barcodes is provided. Furthermore, assume that in a given assay, a population of 1 million beads is used in the assay. On average, there are 100 beads per compartment barcode (Poisson distribution). Furthermore, assume that the beads capture aggregates of 10 million polymers. On average, there are 10 polymers per bead, 100 compartments per compartment barcode, and effectively 1,000 polymers per partition barcode (consisting of 100 compartment barcodes for 100 different physical compartments).

[0164] In another embodiment, single-molecule fractionation and fractional barcoding of polypeptides are achieved by labeling the polypeptide with a DNAUMI tag (e.g., a recording tag) that can be amplified (chemically or enzymatically) at the N-terminus, C-terminus, or both. The DNA tag is attached to the polypeptide body (internal amino acids) via nonspecific photolabeling or specific chemical attachment to reactive amino acids such as lysine. Information from the recording tag attached to the peptide terminus is transferred to the DNA tag via an enzymatic emulsion PCR (Williams et al., Nat Methods, (2006) 3(7):545-550; Schutze et al., Anal Biochem. (2011) 410(1):155-157) or an emulsion in vitro transcription / reverse transcription (IVT / RT) step. In preferred embodiments, nanoemulsions are used such that, on average, there is less than one polypeptide per emulsion droplet of 50 nm to 1000 nm in size (Nishikawa et al., J Nucleic Acids. (2012) 2012:923214, Gupta et al., Soft Matter. (2016) 12(11):2826-41, Sole et al., Langmuir (2006, 22(20):8326-8332). Furthermore, all components of PCR are contained in an aqueous emulsion mixture comprising primers, dNTPs, Mg2+, polymerase, and PCR buffer. When using IVT / RT, the recording tag is designed using a T7 / SP6 RNA polymerase promoter sequence to produce a transcript that hybridizes to a DNA tag attached to the polypeptide body (Ryckelynck et al. al., RNA. (2015) 21(3):458-469). Reverse transcriptase (RT) copies information from hybridized RNA molecules to DNA tags. Thus, emulsion PCR or IVT / RT can be used to effectively transfer information from terminal recording tags to multiple DNA tags attached to the polypeptide body.

[0165] In some embodiments, a sample of a polymer target (e.g., a peptide, polypeptide, or protein) may be processed into a physical area or volume, such as a compartment. Various processing and / or labeling steps can be performed on the sample. In some embodiments, the compartment separates or isolates a subset of polymers from the polymer sample. In some examples, the compartment may be an aqueous compartment (e.g., a microfluidic droplet), a solid compartment (e.g., a picotiter well or microtiter well on a plate, tube, vial, or beads), or a separated area on a surface. In some cases, the compartment may contain one or more beads on which the polymers can be immobilized. In some embodiments, the polymers within the compartment are labeled with a compartment tag containing a barcode. For example, polymers in one compartment may be labeled with the same barcode, or polymers in multiple compartments may be labeled with the same barcode. For example, see Valihrach et al., Int J Mol Sci. 2018 Mar 11;19(3).pii:E807. Encapsulation of cellular contents via gelation in beads is a useful approach to single cell analysis (Tamminen et al., Front Microbiol (2015) 6:195, Spencer et al., ISME J (2016) 10(2):427-436). By barcoding single-cell droplets, all components from a single cell can be labeled with the same identifier (Klein et al., Cell (2015) 161(5):1187-1201, Zilionis et al., Nat Protoc (2017) 12(1):44-73, International Patent Publication No. 2016 / 130704).The barcoding of compartments can be achieved in various ways, including direct incorporation of a unique barcode into each droplet, by droplet bonding (Bio-Rad Laboratories), introduction of barcoded beads into droplets (10×Genomics), or by combined barcoding of the components of the droplets after encapsulation and gelation, as well as by split-pool combined barcoding, as described by Gunderson et al. (see International Patent Publication 2016 / 130704, the entire text of which is incorporated by reference). Similar combined labeling schemes can also be applied to the nucleus (Vitak et al.). al., Nat Methods (2017) 14(3):302-308).

[0166] In some embodiments, the polymers are bonded to a support before contact with a binder. In some cases, it is desirable to use a support with a large carrying capacity to immobilize a large number of polymers in a sample. In some embodiments, it is preferable to immobilize the polymers using a three-dimensional support (e.g., a porous matrix or beads). In some examples, the preparation involves bonding the polymers to nucleic acid molecules or oligonucleotides before or after immobilization. In some embodiments, multiple polymers are attached to a support before contact with a binder.

[0167] The support may include, but is not limited to, any solid or porous support, including beads, microbeads, arrays, glass surfaces, silicon surfaces, plastic surfaces, filters, membranes, PTFE membranes, nylon, microtiter wells, ELISA plates, spin interference disks, nitrocellulose membranes, nitrocellulose-based polymer surfaces, nanoparticles, or microspheres. Materials for the support may include, but are not limited to, acrylamide, agarose, cellulose, dextran, nitrocellulose, glass, gold, quartz, polystyrene, polyethylene vinyl acetate, polypropylene, polyester, polymethacrylate, polyacrylate, polyethylene, polyethylene oxide, polysilicate, polycarbonate, polyvinyl alcohol (PVA), Teflon®, fluorocarbon, nylon, silicone rubber, silica, polyanhydride, polyglycolic acid, polyvinyl chloride, polylactic acid, polyorthoester, functionalized silane, polypropyl fumerate, collagen, glycosaminoglycans, polyamino acids, or any combination thereof. In certain embodiments, the support is beads, such as polystyrene beads, polymer beads, polyacrylate beads, agarose beads, cellulose beads, dextran beads, acrylamide beads, solid core beads, porous beads, paramagnetic beads, glass beads, silica-based beads, or controlled porous beads, or any combination thereof. In some specific embodiments, the support is porous agarose beads.

[0168] In some embodiments, the support may include any suitable solid material, including porous and non-porous materials, that can directly or indirectly associate polymers, such as polypeptides, by any means known in the Art, including covalent and non-covalent interactions, or any combination thereof. The support may be two-dimensional (e.g., planar) or three-dimensional (e.g., gel matrix or beads). The support may be any support surface, including, but not limited to, beads, microbeads, arrays, glass surfaces, silicon surfaces, plastic surfaces, filters, membranes, PTFE membranes, PTFE films, nitrocellulose membranes, nitrocellulose-based polymer surfaces, nylon, microtiter wells, ELISA plates, spin interference disks, polymer matrices, nanoparticles, or microspheres. Support materials include, but are not limited to, acrylamide, agarose, cellulose, dextran, nitrocellulose, glass, gold, quartz, polystyrene, polyethylene vinyl acetate, polypropylene, polyester, polymethacrylate, polyacrylate, polyethylene, polyethylene oxide, polysilicate, polycarbonate, polyvinyl alcohol (PVA), Teflon®, fluorocarbon, nylon, silicone rubber, polyanhydride, polyglycolic acid, polyvinyl chloride, polylactic acid, polyorthoester, functionalized silane, polypropyl fumerate, collagen, glycosaminoglycan, polyamino acid, dextran, or any combination thereof. The support further includes molded polymers such as thin films, membranes, bottles, dishes, fibers, woven fibers, and tubes, particles, beads, microspheres, fine particles, or any combination thereof. For example, if the solid surface is a bead, the beads may include, but are not limited to, ceramic beads, polystyrene beads, polymer beads, polyacrylate beads, methylstyrene beads, agarose beads, cellulose beads, dextran beads, acrylamide beads, solid core beads, porous beads, paramagnetic beads, glass beads, or controlled porous beads, silica-based beads, or any combination thereof. The beads may be spherical or irregular in shape.The beads or support may be porous. The size of the beads may range from nanometers, e.g., 100 nm to several millimeters, e.g., 1 mm. In certain embodiments, the size of the beads is in the range of about 0.2 microns to about 200 microns, or about 0.5 microns to about 5 microns. In some embodiments, the beads may have a diameter of about 1, 1.5, 2, 2.5, 2.8, 3, 3.5, 4, 4.5, 5, 5.5, 6, 6.5, 7, 7.5, 8, 8.5, 9, 9.5, 10, 10.5, 15, or 20 μm. In certain embodiments, the “bead” support may refer to individual beads or a group of beads. In some embodiments, the solid surface is made up of nanoparticles. In certain embodiments, the size of the nanoparticles is in the range of diameters from about 1 nm to about 500 nm, for example, between about 1 nm and about 20 nm, between about 1 nm and about 50 nm, between about 1 nm and about 100 nm, between about 10 nm and about 50 nm, between about 10 nm and about 100 nm, between about 10 nm and about 200 nm, between about 50 nm and about 100 nm, between about 50 nm and about 150 nm, between about 50 nm and about 200 nm, between about 100 nm and about 200 nm, or between about 200 nm and about 500 nm. In some embodiments, the nanoparticles may have diameters of about 10 nm, about 50 nm, about 100 nm, about 150 nm, about 200 nm, about 300 nm, or about 500 nm. In some embodiments, the nanoparticles have a diameter of less than about 200 nm.

[0169] Various reactions can be used to attach polymers to a support (e.g., a solid support or a porous support). Polymers can be attached to the support directly or indirectly. In some cases, polymers are attached to the support via nucleic acids. An exemplary reaction is the copper-catalyzed reaction of azides and alkynes to form triazoles (Huisgen). Examples of various substitution reactions include 1,3-dipolar ring addition, strain-promoted azide-alkyne ring addition (SPAAC), diene-dienophile reactions (Diels-Alder), strain-promoted alkyne-nitrone ring addition, strained alkenes with azides, tetrazines, or tetrazoles, alkenes with azide[3+2] ring addition, alkenes with tetrazines and reverse electron-required Diels-Alder (IEDDA) reactions (e.g., m-tetrazine (mTet) or phenyltetrazine (pTet) and trans-cyclooctene (TCO), or pTet and alkenes), alkenes with tetrazoles, Staudinger ligation of azides and phosphines, and substitution of leaving groups by nucleophilic attack on electrophilic atoms (Horisawa 2014, Knall, Hollauf et al. 2014). Exemplary substitution reactions include reactions of amines with activated esters, N-hydroxysuccinimide esters, isocyanates, isothiocyanates, aldehydes, epoxides, and the like. In some embodiments, iEDDA click chemistry is used to immobilize polymers (e.g., polypeptides) onto a support because it is rapid and achieves high yields at low input concentrations. In another embodiment, m-tetrazine is used in the iEDDA click chemistry reaction instead of tetrazine because it has improved binding stability. In yet another embodiment, phenyltetrazine (pTet) is used in the iEDDA click chemistry reaction. In one case, the polypeptide is labeled with an alkyne-NHS ester (acetylene-PEG-NHS ester) reagent or a bifunctional click chemistry reagent such as alkyne-benzophenone to produce an alkyne-labeled polypeptide. In some embodiments, the alkyne may also be a strained alkyne such as cyclooctine, including dibenzocyclooctyl (DBCO).

[0170] In certain embodiments where multiple polymers are immobilized on the same support, the polymers can be appropriately spaced apart to accommodate analytical steps used to evaluate a target. For example, it may be advantageous to optimally space the polymers to allow nucleic acid-based methods for evaluating and sequencing proteins to be performed. In some cases, the spacing of polymers on the support is determined considering that information transfer from adapter molecules hybridized to a binding agent coding tag attached to one immobilized polymer may reach adjacent polymers.

[0171] In some embodiments, the surface of the support is passivated (blocked). A “passivated” surface refers to a surface treated with an outer layer of material. Methods for passivating surfaces include standard methods from the literature of fluorescence single-molecule analysis, such as polymer-like polyethylene glycol (PEG) (Pan et al., 2015, Phys. Biol. 12:045006), polysiloxane (e.g., Pluronic F-127), star-shaped polymer (e.g., star-shaped PEG) (Groll et al., 2010, Methods Enzymol. 472:1-18), hydrophobic dichlorodimethylsilane (DDS) + self-assembled Tween®-20 (Hua et al., 2014, Nat. Methods 11:1233-1236), diamond-like carbon (DLC), DLC + PEG (Stavis et al., 2011, Proc. Natl. Acad. Sci. USA) This includes passivating the surface with a compound (108:983-988) and an amphoteric moiety (e.g., U.S. Patent Application Publication No. 2006 / 0183863). In addition to covalent surface modification, many passivators can also be used, including surfactants such as Tween®-20, polysiloxanes in solution (Pluronic series), polyvinyl alcohol (PVA), and proteins such as BSA and casein. Alternatively, the density of a polymer (e.g., a protein, polypeptide, or peptide) can be titrated on or within the volume of a solid substrate by spiking a competing substance or "dummy" reactive molecule when immobilizing the protein, polypeptide, or peptide on the solid substrate.

[0172] To control the spacing of immobilized targets on a support, the density of functional coupling groups (e.g., TCO or carboxyl groups (COOH)) for target attachment can be titrated on the substrate surface. In some embodiments, multiple target molecules (e.g., polymers) are spaced on or within the volume of a support (e.g., a porous support) such that adjacent molecules are spaced at distances of approximately 50 nm to 500 nm, or approximately 50 nm to 400 nm, or approximately 50 nm to 300 nm, or approximately 50 nm to 200 nm, or approximately 50 nm to 100 nm. In some embodiments, multiple molecules are spaced on the surface of the support at average distances of at least 50 nm, at least 60 nm, at least 70 nm, at least 80 nm, at least 90 nm, at least 100 nm, at least 150 nm, at least 200 nm, at least 250 nm, at least 300 nm, at least 350 nm, at least 400 nm, at least 450 nm, or at least 500 nm. In some embodiments, multiple molecules are separated on the surface of the support by an average distance of at least 50 nm. In some embodiments, molecules are separated on the surface or within the volume of the support empirically such that the relative frequency of intermolecular events to intramolecular events (e.g., information transfer) is <1:10, <1:100, <1:1,000, or <1:10,000.

[0173] In some embodiments, the multiple polymers are approximately 50-100 nm, approximately 50-250 nm, approximately 50-500 nm, approximately 50-750 nm, approximately 50-1,000 nm, approximately 50-1,500 nm, approximately 50-2,000 nm, approximately 100-250 nm, approximately 100-500 nm, approximately 200-500 nm, approximately 300-500 nm, and approximately 100-1,000 nm. The molecules are coupled on a support separated by an average distance between two adjacent molecules in the range of approximately 500-600 nm, 500-700 nm, 500-800 nm, 500-900 nm, 500-1000 nm, 500-2000 nm, 500-5000 nm, 1000-5000 nm, or 3000-5000 nm.

[0174] In some embodiments, the appropriate spacing of the polymers on the support is achieved by titrating the ratio of available attachment molecules on the substrate surface. In some examples, the substrate surface (e.g., bead surface) is functionalized with carboxyl groups (COOH) that are treated with an activator (e.g., the activator is EDC and sulfo-NHS). In some examples, the substrate surface (e.g., bead surface) contains NHS moieties. In some embodiments, a mixture of mPEG n -NH2 and NH2-PEG n -mTet is added to the activated beads (where n is any numerical value such as 1 to 100). The ratio between mPEG3-NH2 (not available for coupling) and NH2-PEG 24 -mTet (available for coupling) is titrated to produce an appropriate density of functional moieties available for attaching the polypeptide to the substrate surface. In certain embodiments, the average spacing between coupling moieties (e.g., NH2-PEG4-mTet) on the solid surface is at least 50 nm, at least 100 nm, at least 250 nm, or at least 500 nm. In some specific embodiments, the ratio of NH2-PEG n -mTet to mPEG3-NH2 is about 1:1000 or greater, about 1:10,000 or greater, about 1:100,000 or greater, or about 1:1,000,000 or greater. In some further embodiments, the recording tag attaches to NH2-PEG n -mTet. In some embodiments, the spacing of the polymers on the support is achieved by controlling the concentration and / or number of available COOH or other functional groups on the support.

[0175] B. Cleavage In some embodiments, the methods provided further include removing a portion of the polymer after information transfer. The removal of a portion of the polymer exposes a new portion of the polymer for analysis, which can be bound, for example, by a binder provided in the next cycle of analysis. If a polypeptide is being analyzed, the method may include removing a portion of the polypeptide containing one or more amino acids (e.g., terminal amino acids). In embodiments of methods for analyzing peptides or polypeptides using a degradation approach, a first binder is contacted and bound to the nNTAA of the peptide of n amino acids, transferring coding tag information to the nucleic acid associated with the peptide, thereby generating a primary elongated nucleic acid (e.g., on a recording tag), and the nNTAA is detached. Removal of the n labeled NTAA by contact with an enzyme or chemical reagent converts the n-1 amino acids of the peptide to the N-terminal amino acid, which is referred to herein as the n-1 NTAA. A second binder is brought into contact with the peptide or polypeptide and bound to an n-1 NTAA, transferring information from the coding tag to the primary elongated nucleic acid, thereby generating a secondary elongated nucleic acid (for example, to produce a linked n-th order elongated nucleic acid representing the peptide). The detachment of n-1 labeled NTAAs converts n-2 amino acids of the peptide or polypeptide to the N-terminal amino acid, which is referred to herein as the n-2 NTAA. Additional binding, information transfer, and detachment occur with up to n amino acids as described above, collectively producing an n-th order elongated nucleic acid representing the peptide or polypeptide or n distinct elongated nucleic acids. As used herein, n "order" as used with respect to a binder, coding tag, or elongated nucleic acid refers to the nth binding cycle in which the binder and its associated coding tag are used, or the nth binding cycle in which the elongated nucleic acid is created (for example, on a recording tag). In some embodiments, the steps involving NTAAs in the exemplary approaches described may instead be performed using C-terminal amino acids (CTAAs).

[0176] In certain embodiments relating to the analysis of peptides or polypeptides, terminal amino acids are removed or cleaved from the peptide or polypeptide to expose new terminal amino acids. The portion of the polypeptide to be removed may be a labeled (e.g., chemically labeled or modified) portion of the polypeptide. In some embodiments, the terminal amino acids are NTAAs. In other embodiments, the terminal amino acids are CTAAs. Cleavage of terminal amino acids can be achieved by any number of known techniques, including chemical and enzymatic cleavage. In some embodiments, an engineered enzyme or reagent that catalyzes or facilitates the removal of a labeled (e.g., chemically labeled or modified) N-terminal amino acid is used. In some embodiments, the terminal amino acids are removed or eliminated using one of the methods described in International Patent Publication 2019 / 089846, and International Patent Publication PCT / US2020 / 29969 and PCT / US2020 / 24521.

[0177] In some embodiments, the reagents for removing terminal amino acids include carboxypeptidase or aminopeptidase or variants, mutants, or modified proteins thereof; hydrolase or variants, mutants, or modified proteins thereof; mild Edman degradation reagent; Edmanase enzyme; anhydrous TFA, a base; or any combination thereof. In some embodiments, mild Edman degradation is performed using dichloroic acid or monochloroic acid, or mild Edman degradation is performed using TFA, TCA, or DCA, or mild Edman degradation is performed using triethylamine, triethanolamine, or triethylammonium acetate (Et3NHOAc).

[0178] In some cases, the reagent for removing amino acids contains a base. In some embodiments, the base is a hydroxide, alkylated amine, cyclic amine, carbonate buffer, trisodium phosphate buffer, or metal salt. In some examples, the hydroxide is sodium hydroxide. Alkylated amines are selected from methylamine, ethylamine, propylamine, dimethylamine, diethylamine, dipropylamine, trimethylamine, triethylamine, tripropylamine, cyclohexylamine, benzylamine, aniline, diphenylamine, N,N-diisopropylethylamine (DIPEA), and lithium diisopropylamide (LDA). Cyclic amines are selected from pyridine, pyrimidine, imidazole, pyrrole, indole, piperidine, prolysine, 1,8-diazabicyclo[5.4.0]undeca-7-ene (DBU), and 1,5-diazabicyclo[4.3.0]non-5-ene (DBN). The carbonate buffer contains sodium carbonate, potassium carbonate, calcium carbonate, sodium bicarbonate, potassium bicarbonate, or calcium bicarbonate. The metal salt contains silver, or the metal salt is AgClO4.

[0179] In some embodiments, the methods provided involve cleaving the N-terminal amino acid when modified with groups such as PTC, modified PTC, Cbz, DNP, SNP, acetyl, guanidinyl, or diheterocyclic methanimines. Enzymatic cleavage of NTAA can be achieved by aminopeptidases or other peptidases. Aminopeptidases exist naturally as monomeric and multimeric enzymes and may be metal or ATP-dependent. Natural aminopeptidases have very limited specificity and generally progressively cleave the N-terminal amino acid, cleaving amino acids one after another. In the methods described herein, the aminopeptidase (e.g., metalloenzymatic aminopeptidase) can be manipulated to have specific binding or catalytic activity to NTAA only when modified with an N-terminal label. Thus, the aminopeptidase cleaves only a single amino acid at a time from the N-terminus, allowing control of the degradation cycle. In some embodiments, the modified aminopeptidase is selective for the N-terminal label but non-selective with respect to the identity of the amino acid residue. In other embodiments, the modified aminopeptidase is selective for both amino acid residue identity and N-terminal labeling. Manipulated aminopeptidase variants that bind to and cleave individual groups or subgroups of labeled (biotinylated) NTAAs have been described (see PCT Publication 2010 / 065322).

[0180] Engineered aminopeptidase variants that bind to and cleave individual groups or subgroups of labeled (biotinylated) NTAAs have been described (PCT Publication 2010 / 065322, the entire work of which is incorporated by reference). Aminopeptidases are enzymes that cleave amino acids from the N-terminus of proteins or peptides. Natural aminopeptidases have very limited specificity and generally progressively cleave the N-terminal amino acid, cleaving amino acids one after another (Kishor et al., 2015, Anal. Biochem. 488:6-8). However, residue-specific aminopeptidases have been identified (Eriquez et al., J. Clin. Microbiol. 1980, 12:667-71, Wilce et al., 1998, Proc. Natl. Acad. Sci. USA 95:3472-3477, Liao et al., 2004, Prot. Sci. 13:1802-10). Control of stepwise degradation of peptides from the N-terminus can be achieved by using engineered aminopeptidases that are active only in the presence of a label (e.g., binding activity or catalytic activity). In certain embodiments, the aminopeptidase may be engineered to be nonspecific, and as a result, it simply recognizes the labeled N-terminus rather than selectively recognizing one particular amino acid over another. In yet another embodiment, cyclic cleavage is achieved by cleaving acetylated NTAAs using engineered acylpeptide hydrolases (APHs). In yet another embodiment, amidation (guanidinylation) of NTAA is used to enable gentle cleavage of NaOH-labeled NTAA (Hamada, (2016) Bioorg Med Chem Lett 26(7):1690-1695).

[0181] In some embodiments, the method further comprises contacting the polypeptide with a proline aminopeptidase under conditions suitable for cleaving the N-terminal proline prior to step (b). In some examples, the proline aminopeptidase (PAP) is an enzyme capable of specifically cleaving the N-terminal proline from a polypeptide. PAP enzymes that cleave the N-terminal proline are also called proline iminopeptidases (PIPs). Known monomeric PAPs include family members from B. coagulans, L. delbrueckii, N. gonorrhoeae, F. meningosepticum, S. marcescens, T. acidophilum, and L. plantarum (MEROPS S33.001) (Nakajima et al., J Bacteriol. (2006) 188(4):1599-606, Kitazono et al., Bacteriol (1992) 174(24):7919-7925). Known multimeric PAPs include D. hansenii (Bolumar et al., (2003) 86(1-2):141-151) and similar homologs from other species (Basten et al., Mol Genet Genomics (2005) 272(6):673-679). Either natural or engineered variants / mutants of the PAP can be used.

[0182] Regarding embodiments relating to CTAA binders, methods for cleaving CTAA from peptides or polypeptides are also known in the art. For example, U.S. Patent No. 6,046,053 discloses a method of reacting a peptide or protein with an alkyl acid anhydride to convert the carboxyl terminus to an oxazolone, and then releasing the C-terminal amino acid by reaction with an acid and an alcohol or ester. Enzymatic cleavage of CTAA can also be achieved by carboxypeptidases. Some carboxypeptidases exhibit amino acid preference; for example, carboxypeptidase B preferentially cleaves basic amino acids such as arginine and lysine. As described above, carboxypeptidases can also be modified in the same way as aminopeptidases to manipulate carboxypeptidases that specifically bind to CTAA having a C-terminal label. Thus, carboxypeptidases cleave only a single amino acid from the C-terminus at a time, allowing control of the degradation cycle. In some embodiments, the modified carboxypeptidase is selective for the C-terminal label but non-selective with respect to the identity of the amino acid residue. In another embodiment, the modified carboxypeptidase is selective for both amino acid residue identity and C-terminal labeling.

[0183] C. Analysis In some embodiments, the decompression record tag generated from performing the provided method includes information transferred from one or more coding tags. In some embodiments, the decompression record tag includes a set of information transferred from multiple coding tags indicating the sequence of binding by a binder. This decompression record tag can be analyzed by any preferred method. In some embodiments, the decompression record tag (or part thereof) is amplified at least before determining the sequence of coding tags within the decompression record tag. In some embodiments, the decompression record tag (or part thereof) is emitted at least before determining the sequence of coding tags within the decompression record tag.

[0184] The length of the final extended record tag produced by the method described herein depends on several factors, including the length of the coding tag (e.g., barcode and spacer) and the length of the nucleic acid (e.g., optionally, optionally, any unique molecular identifier, spacer, universal priming site, barcode, or a combination thereof). After the final tag information is transferred to the extended nucleic acid (e.g., from any coding tag), the tag can be capped by adding a universal reverse priming site via ligation, primer extension, or other methods known in the art. In some embodiments, a universal forward priming site in the nucleic acid (e.g., on the record tag) is compatible with a universal reverse priming site to be added to the final extended nucleic acid. In some embodiments, the universal reverse priming site is an Illumina P7 primer or an Illumina P5 primer. Depending on the sense of the nucleic acid strand from which the identification information from the coding tag is transferred, sense or antisense P7 can be added. The extended nucleic acid library can be cleaved or amplified directly from a support (e.g., beads) and used in conventional next-generation sequencing assays and protocols.

[0185] In some embodiments, the primer extension reaction is performed on a library of single-stranded extended nucleic acids (e.g., extended on a recording tag) to copy its complementary strand. In some embodiments, the peptide sequencing assay (e.g., a ProteoCode assay) includes several chemical and enzymatic steps in a periodic progression.

[0186] Extended nucleic acid recording tags can be processed and analyzed using various nucleic acid sequencing methods. In some embodiments, an extended recording tag containing information from one or more coding tags and any other nucleic acid components is processed and analyzed. In some embodiments, a set of extended recording tags can be concatenated. In some embodiments, the extended recording tags can be amplified before sequencing.

[0187] Examples of sequencing methods include, but are not limited to, next-generation sequencing methods such as chain termination sequencing (Sanger sequencing); synthesis sequencing, ligation sequencing, hybridization sequencing, Polony sequencing, ion semiconductor sequencing, and pyrosequencing; and third-generation sequencing methods such as single-molecule real-time sequencing, nanopore-based sequencing, duplex-interrupted sequencing, and direct DNA imaging using advanced microscopy.

[0188] Suitable sequencing methods for use in the present invention include, but are not limited to, hybridization sequencing, synthesis techniques (e.g., HiSeq® and Solexa®, Illumina), SMRT® (Single Molecule Real-Time) technology (Pacific Biosciences), true single-molecule sequencing (e.g., HeliScope®, Helicos Biosciences), large-scale parallel next-generation sequencing (e.g., SOLiD®, Applied Biosciences, Solexa and HiSeq®, Illumina), massively parallel semiconductor sequencing (e.g., Ion Torrent), pyrosequencing techniques (e.g., GS FLX and GS Junior Systems, Roche / 454), and nanopore sequencing (e.g., Oxford Nanopore Technologies).

[0189] Nucleic acid libraries (e.g., elongated nucleic acids) can be amplified in various ways. A nucleic acid library (e.g., a recording tag containing information from one or more coding tags) can undergo exponential amplification, for example, via PCR or emulsion PCR. Emulsion PCR is known to yield more uniform amplification (Hori, Fukano et al., Biochem Biophys Res Commun (2007) 352(2):323-328). Alternatively, a nucleic acid library (e.g., elongated nucleic acids) can undergo linear amplification, for example, via in vitro transcription of template DNA using T7 RNA polymerase. A nucleic acid library (e.g., elongated nucleic acids) can be amplified using primers compatible with the universal forward priming and universal reverse priming sites it contains. A nucleic acid library (e.g., a recording tag) can also be amplified using tailed primers to add sequences to the 5' end, 3' end, or either end of the elongated nucleic acid. Sequences that can be added to the ends of the extended nucleic acid include library-specific index sequences, and multiple libraries, adapter sequences, read primer sequences, or other sequences can be multiplexed in a single sequencing run to make the extended nucleic acid library compatible with the sequencing platform. An example of library amplification in preparation for next-generation sequencing is as follows: Using an extended nucleic acid library eluted from approximately 1 mg of beads (approximately 10 ng), 200 μM dNTPs, 1 μM forward and reverse amplification primers, and 0.5 μl (1 U) of Phusion Hot Start enzyme (New England Biolabs), a 20 μl PCR reaction volume is set up and subjected to 20 cycles of 98°C for 30 seconds, followed by 98°C for 10 seconds, 60°C for 30 seconds, 72°C for 30 seconds, and 72°C for 7 minutes, and then held at 4°C.

[0190] In certain embodiments, a library of nucleic acids (e.g., elongated nucleic acids) can undergo targeted enrichment before, during, or after amplification. In some embodiments, targeted enrichment can be used to selectively capture or amplify elongated nucleic acids representing a target macromolecule (e.g., polypeptide) from a library of elongated nucleic acids before sequencing. In some embodiments, targeted enrichment for protein sequencing is esoteric due to the high cost and difficulty of producing highly specific binders for target proteins. In some cases, antibodies are notoriously nonspecific and difficult to scale production across thousands of proteins. In some embodiments, the methods of this disclosure circumvent this problem by converting the protein code into a nucleic acid code, which can then utilize a wide range of targeted DNA enrichment strategies available for DNA libraries. In some cases, the peptide of interest can be enriched in a sample by enriching the corresponding elongated nucleic acid. Methods for targeted enrichment are known in the art and include hybrid capture assays, PCR-based assays such as TruSeq custom amplicons (Illumina), and padlock probes (also known as molecular inversion probes) (see Mamanova et al., (2010) Nature Methods 7:111-118, Bodi et al., J. Biomol. Tech. (2013) 24:73-86, Ballester et al., (2016) Expert Review of Molecular Diagnostics 357-372, Mertes et al., (2011) Brief Funct. Genomics 10:374-386, Nilsson et al., (1994) Science 265:2085-8, each of which is incorporated herein by reference in whole).

[0191] In one embodiment, a library of nucleic acids (e.g., extension record tags) is enriched via a hybrid capture-based assay. In the hybrid capture-based assay, the library of extension nucleic acids is hybridized to a target-specific oligonucleotide labeled with an affinity tag (e.g., biotin). The extension nucleic acids hybridized to the target-specific oligonucleotide are “pulled down” via the affinity tag using an affinity ligand (e.g., streptavidin-coated beads) to wash away background (non-specific) extension nucleic acids. The enriched extension nucleic acids (e.g., extension nucleic acids) are then obtained for positive enrichment (e.g., elution from beads). In some embodiments, oligonucleotides complementary to the corresponding extension nucleic acid library expression of the peptide of interest can be used in the hybrid capture assay. In some embodiments, sequential rounds or enrichments can also be performed using the same or different bait sets.

[0192] In another embodiment, enriched fractions of library elements representing a subset of polypeptides can be selected and modularized using primer extension and ligation-based amplification enrichment (e.g., AmpliSeq, PCR, TruSeq TSCA). The degree of primer extension, ligation, or amplification can also be controlled using competing oligonucleotides. In its simplest implementation, this can be achieved by having a mixture of target-specific primers containing a universal primer tail and competing primers lacking a 5' universal primer tail. After initial primer extension, only primers with the 5' universal primer sequence can be amplified. The ratio of primers with and without the universal primer sequence controls the proportion of targets amplified. In other embodiments, the proportion of library elements undergoing primer extension, ligation, or amplification can be controlled by including primers that hybridize but are not extensional.

[0193] Targeted enrichment methods can also be used in negative selection mode to selectively remove elongated nucleic acids from the library before sequencing. Examples of undesirable elongated nucleic acids that can be removed include those representing excessively abundant polypeptide species, such as proteins, albumins, and immunoglobulins.

[0194] The enriched fraction of a specific locus can also be modulated by using a competing oligonucleotide bait in the hybrid capture step that hybridizes to the target but lacks the biotin moiety. The competing oligonucleotide bait competes for hybridization to the target with the standard biotinylation bait, effectively modulating the proportion of the target pulled down during enrichment. The 10-order-of-magnitude dynamic range of protein expression can be compressed by several orders of magnitude using this competitive suppression approach, especially for species with excessively abundant proteins such as albumin. Thus, compared to standard hybrid capture, the proportion of captured library elements for a particular locus can be downregulated to enrichment from 100% to 0%.

[0195] Furthermore, library canonicalization techniques can be used to remove excessively abundant species from extended nucleic acid libraries. This approach is ideal for libraries of defined lengths derived from peptides produced by site-specific protease digestion, such as trypsin, LysC, and GluC. In one example, canonicalization can be achieved by denaturing a double-stranded library and re-annealing the library elements. Due to the quadratic rate constant of bimolecular hybridization kinetics, abundant library elements re-anneal more rapidly than less abundant elements (Bochman, Paeschke et al. 2012). ssDNA library elements can be chromatographed on a hydroxyapatite column (VanderNoot, et al., 2012, Biotechniques). Abundant dsDNA library elements can be separated using methods known in the art, such as treating the library with a bispecific nuclease (DSN) from king crab (Shagin et al., (2002) Genome Res. 12:1935-42) that disrupts dsDNA library elements (53:373-380).

[0196] Fractionation, enrichment, and subtraction of polypeptides and / or the resulting extended nucleic acid library before attachment to a support can save sequencing reads and improve the measurement of less abundant species.

[0197] In some embodiments, a library of nucleic acids (e.g., elongated nucleic acids) is ligated or end-complementary PCR to create long DNA molecules, each containing multiple different elongation recorder tags, elongation coding tags, or ditags (Du et al., (2003) BioTechniques 35:66-72, Muecke et al., (2008) Structure 16:837-841, U.S. Patent No. 5,834,252, each of which is incorporated by reference). This embodiment is preferred for nanopore sequencing, in which long strands of DNA are analyzed by a nanopore sequencing device.

[0198] In some embodiments, direct single-molecule analysis is performed on nucleic acids (e.g., elongated nucleic acids) (see, e.g., Harris et al., (2008) Science 320:106-109). Nucleic acids (e.g., elongated nucleic acids) can be analyzed directly on a support such as a flow cell or beads suitable for loading onto a flow cell surface (optionally, microcell patterned), and the flow cell or beads can be integrated with a single-molecule sequencer or single-molecule decoding instrument. For single-molecule decoding, hybridization of several rounds of fluorescently labeled decoding oligonucleotides (Gunderson et al., (2004) Genome Res. 14:970-7) can be used to verify both the identity and sequence of coding tags within the elongated nucleic acid (e.g., on a recording tag). In some embodiments, the binder can be labeled with cycle-specific coding tags as described above (see also Gunderson et al., (2004) Genome Res. 14:970-7).

[0199] Following sequencing of a nucleic acid library (e.g., extended nucleic acids), the resulting sequences, if used, are folded by UMI, associated with their corresponding polypeptides, and aligned to the entire proteome. The resulting sequences can also be folded by their compartment tags and associated with their corresponding compartment proteomes, which, in certain embodiments, contain only a single or very limited number of protein molecules. Both protein identification and quantification can be readily derived from this digital peptide information.

[0200] The methods disclosed herein can be used for analysis including the simultaneous (multiplexed) detection, quantification, and / or sequencing of multiple polymers. As used herein, multiplexing refers to the analysis of multiple polymers (e.g., polypeptides) in the same assay. Multiple polymers may originate from the same sample or different samples. Multiple polymers may originate from the same or different subjects. Multiple polymers analyzed may be different polymers or the same polymer originating from different samples. Multiple polymers include two or more polymers, five or more polymers, ten or more polymers, fifty or more polymers, one hundred or more polymers, five hundred or more polymers, one thousand or more polymers, five thousand or more polymers, one hundred thousand or more polymers, one thousand thousand or more polymers, one thousand thousand or more polymers, one thousand thousand or more polymers, one thousand thousand or more polymers, one thousand thousand or more polymers, one thousand thousand or more polymers, one thousand thousand or more polymers, or one thousand thousand or more polymers.

[0201] III. Kits and manufactured products Provided herein are kits and products comprising components for pre-forming polymer analysis assays. In some embodiments, the kit includes a binder containing a coding tag containing identification information relating to the binder. In some embodiments, the kit includes a set of binders, each containing a coding tag. In some embodiments, the binder is configured to bind to a polymer associated with a recording tag, and reagents for transferring information from the coding tag to the recording tag are also provided. The kit also provides nucleic acid conjugation reagents, polymerases, and double-strand nucleic acid cleavage reagents. In some embodiments, the nucleic acid conjugation reagents, polymerases, and double-strand nucleic acid cleavage reagents are provided as a mixture. For example, the nucleic acid conjugation reagent is a chemical ligation reagent or an enzymatic ligation reagent. In some cases, the double-strand nucleic acid cleavage reagent is a restriction enzyme.

[0202] In some embodiments, the kit further contains other reagents for processing and analyzing target polymers (e.g., proteins, polypeptides, or peptides). The kit and products may contain one or more of the reagents and components used in the methods described in Sections I and II. In some embodiments, the kit includes reagents for preparing a sample, such as for preparing a polymer from a sample and conjugating it to a support. In some embodiments, the kit optionally includes instructions for performing a polymer analysis assay. In some embodiments, the kit includes one or more of the following components: a binder, a nucleic acid conjugation reagent, a polymerase, a double-strand nucleic acid cleavage reagent, a solid support, a recording tag, a reagent for transferring information, a sequencing reagent, and / or any necessary buffers. The recording tag, binder, and coding tag may have specific structures and characteristics as described in Sections IA, IB, and IC, respectively.

[0203] In one embodiment, the components provided herein are used to prepare a reaction mixture. In a preferred embodiment, the reaction mixture is a solution. In a preferred embodiment, the reaction mixture comprises one or more of the following: a binder and associated coding tag, a solid support, a recording tag, a nucleic acid conjugation reagent, a polymerase, a double-strand nucleic acid cleavage reagent, a sequencing reagent, a buffer, and / or other additives used to transfer information.

[0204] In another embodiment, disclosed herein is a kit for performing a polymer analysis assay comprising a library of binders, each binder comprising or associated with a coding tag. In some embodiments, the coding tag comprises identification information relating to the binding site of the binder. In some examples, the binding site can bind to one or more N-terminal, internal, or C-terminal amino acids of a target peptide or polypeptide, or to one or more N-terminal, internal, or C-terminal amino acids of a peptide modified by a functionalization / modification reagent.

[0205] In some embodiments, the kits and products further comprise multiple nucleic acid molecules or oligonucleotides. In some embodiments, the kit comprises multiple barcodes. The barcodes may include compartment barcodes, partition barcodes, sample barcodes, fraction barcodes, or any combination thereof. In some cases, the barcodes include unique molecular identifiers (UMIs). In some examples, the barcodes comprise DNA molecules, DNA with pseudocomplementary bases, RNA molecules, BNA molecules, XNA molecules, LNA molecules, PNA molecules, γPNA molecules, non-nucleic acid sequencing polymers, such as polysaccharides, polypeptides, peptides, or polyamides, or combinations thereof. In some embodiments, the barcodes are configured to adhere to polymers for analysis within the sample, such as proteins, or to nucleic acid components associated with polymers.

[0206] In some embodiments, the kit further comprises reagents for processing polymers, such as proteins. Any combination of protein fractionation, concentration, and subtraction methods can be performed. For example, the reagents can be used to fragment or digest proteins. In some cases, the kit includes reagents and components for fractionation, isolation, subtraction, and concentration of proteins. In some examples, the kit further comprises proteases such as trypsin, LysN, or LysC. In some embodiments, the kit comprises a support for immobilizing one or more targets, and reagents for immobilizing the targets on the support. In some embodiments, the kit further comprises reagents for removing the N-terminal amino acids (NTAAs) of polypeptides, such as chemical agents or enzymes. In some embodiments, the kit further comprises reagents for modifying the terminal amino acids of polypeptides, such as chemical agents or enzymes.

[0207] In some embodiments, the kit also includes one or more buffers or reaction fluids necessary to perform any of the steps of a polymer analysis assay. Buffers, including wash buffers, reaction buffers, and binding buffers, elution buffers, etc., are known to those skilled in the art. In some embodiments, the kit further includes buffers and other components accompanying the other reagents described herein. The reagents, buffers, and other components may be supplied in vials (such as sealed vials), containers, ampoules, bottles, jars, flexible packaging (e.g., sealed Mylar or plastic bags), etc. Any component of the kit can be sterile and / or sealed.

[0208] In some embodiments, the kit includes one or more reagents for nucleic acid sequencing analysis. In some examples, the reagents for sequence analysis are intended for use in synthesis sequencing, ligation sequencing, single-molecule sequencing, single-molecule fluorescence sequencing, hybridization sequencing, Polony sequencing, ion semiconductor sequencing, pyrosequencing, single-molecule real-time sequencing, nanopore-based sequencing, or direct imaging of DNA using advanced microscopy, or any combination thereof.

[0209] In addition to the components described above, the kit may further include instructions for using the kit components to carry out the method, i.e., instructions for sample preparation, processing and / or analysis. The kits described herein may also include other materials desirable from a commercial and user perspective, including other buffers, diluents, filters, syringes and accompanying documentation with instructions for carrying out any of the methods described herein.

[0210] Any of the kit components described above, and any molecules, molecular complexes or conjugates, reagents (e.g., chemical or biological reagents), drugs, structures (e.g., supports, surfaces, particles, or beads), reaction intermediates, reaction products, binding complexes, or any other products disclosed and / or used in exemplary kits and methods, may be provided separately or in any preferred combination to form a kit.

[0211] IV. Exemplary Embodiments The embodiments provided include the following: 1. A method for analyzing polymers, (a) Providing a polymer and an associated recording tag bonded to a support, (b) A step of bringing a polymer into contact with a binder capable of bonding to the polymer, the binder comprising a coding tag having identification information relating to the binder, thereby enabling bonding between the polymer and the binder, (c) A step of joining the 5' end of the recording tag to the 3' end of the coding tag using a nucleic acid conjugation reagent, (d) A step of using polymerase to extend the recording tag using the coding tag as a template to generate a double-stranded extended recording tag, (e) A step of cleaving the double-stranded extension record tag with a double-stranded nucleic acid cleavage reagent to generate a 3' overhang within the extension record tag, A method comprising the steps of transferring information from a coding tag to a recording tag to generate an expanded recording tag. 2. The method according to Embodiment 1, wherein in step (d), the double-strand extension record tag includes a recognition sequence that can be recognized by a double-strand nucleic acid cleavage reagent. 3. The method according to Embodiment 1 or Embodiment 2, wherein the cleavage in step (e) releases a binder from the polymer. 4. The method according to any one of Embodiments 1 to 3, wherein the recording tag and the coding tag include nucleic acids. 5. The method according to any one of Embodiments 1 to 4, wherein the extension record tag includes a nucleic acid hairpin. 6. The method according to any one of Embodiments 1 to 5, wherein steps (a), (b), (c), (d), and (e) are performed sequentially. 7. The method according to any one of Embodiments 1 to 6, wherein steps (b), (c), (d), and (e) are repeated one or more times in a periodic manner. 8. The method according to any one of Embodiments 1 to 7, wherein the polymer is a lipid, carbohydrate, protein, polypeptide, or peptide. 9. The method according to Embodiment 8, wherein the peptide is obtained by fragmenting a protein from a biological sample. 10. The method according to any one of Embodiments 1 to 9, wherein the polymer for analysis is not a nucleic acid. 11. The method according to any one of embodiments 7 to 10, further comprising removing a portion of the polymer before repeating step (b). 12. The method according to any one of embodiments 8 to 11, further comprising removing the N-terminal amino acid (NTAA) of the polypeptide to expose a new NTAA of the polypeptide before repeating step (b). 13. The method according to any one of Embodiments 7 to 12, wherein in step (e), the 3' overhang of the extension record tag generated by the double-stranded nucleic acid cleavage reagent is available for hybridization with a second coding tag when step (b) is repeated. 14. The method according to any one of embodiments 1 to 13, further comprising removing the binder. 15. The method according to Embodiment 14, wherein the binder is removed after the information of the coding tag has been transferred to the recording tag. 16. The method according to Embodiment 14 or Embodiment 15, wherein the binder is removed before repeating step (b). 17. The method according to any one of Embodiments 1 to 16, wherein the method comprises, in step (b), contacting a plurality of polymers with a single binder or a plurality of binders. 18. The method according to any one of Embodiments 8 to 17, further comprising treating the polypeptide with a reagent for modifying the terminal amino acids of the polypeptide. 19. The method according to Embodiment 18, wherein the reagent for modifying the terminal amino acids of the polypeptide comprises a chemical agent or an enzyme. 20. The method according to Embodiment 18 or Embodiment 19, wherein the polypeptide is treated with a reagent for modifying the terminal amino acids of the polypeptide before step (b). 21. The method according to any one of Embodiments 1 to 20, wherein the recording tag associated with the polymer includes a double-stranded region. 22. The method according to Embodiment 21, wherein the recording tag associated with the polymer includes a nucleic acid hairpin. 23. The method according to any one of Embodiments 1 to 22, wherein the recording tag associated with the polymer has a 3' overhang. 24. The method according to any one of Embodiments 1 to 23, wherein the recording tag includes a barcode. 25. The method according to any one of Embodiments 1 to 24, wherein the recording tag includes a unique molecular identifier (UMI). 26. The method according to Embodiment 24 or Embodiment 25, wherein step (a) includes providing a barcode and / or UMI to a recording tag. 27. The method according to Embodiment 24 or Embodiment 26, wherein the barcode is a sample barcode, a fraction barcode, a spatial barcode, and / or a compartment tag. 28. The method according to Embodiment 24 or Embodiment 25, wherein step (a) includes providing a barcode and / or UMI to a recording tag using ligation and / or extension. 29. The method according to any one of embodiments 23 to 28, wherein step (a) includes cutting the record tag to generate a 3' overhang. 30. The method according to any one of embodiments 23 to 29, wherein the 3' overhang of the recording tag is generated by extension and / or cleavage with a double-strand nucleic acid cleavage reagent. 31. The method according to any one of Embodiments 1 to 30, wherein the recording tag includes a universal priming portion. 32. The method according to Embodiment 31, wherein the universal priming region includes a priming region for amplification, sequencing, or both. 33. The method according to any one of Embodiments 1 to 32, wherein steps (c), (d), and (e) are performed as a one-pot reaction. 34. The method according to Embodiment 33, wherein the nucleic acid conjugation reagent, polymerase, and double-strand nucleic acid cleavage reagent of steps (c), (d), and (e) are each provided as a mixture. 35. The method according to any one of embodiments 1 to 32, wherein steps (c), (d), and (e) are performed sequentially and separately. 36. The method according to any one of Embodiments 1 to 32, wherein the nucleic acid conjugation reagent of step (c) and the polymerase of step (d) are provided simultaneously. 37. The method according to any one of Embodiments 1 to 32, wherein the polymerase of step (d) and the double-strand nucleic acid cleavage reagent of step (e) are provided simultaneously. 38. The method according to any one of Embodiments 1 to 37, wherein the coding tag comprises a partial restriction enzyme recognition sequence. 39. The method according to Embodiment 38, wherein the partially restriction enzyme recognition sequence of the coding tag is single-stranded. 40. The method according to any one of Embodiments 1 to 39, wherein the coding tag includes a barcode and / or a unique molecular identifier (UMI). 41. The method according to any one of Embodiments 1 to 40, wherein the recording tag and the coding tag each include a spacer. 42. The method according to Embodiment 41, wherein the spacer is a nucleic acid molecule with 10 bases or less, 9 bases or less, 8 bases or less, 7 bases or less, 6 bases or less, 5 bases or less, 4 bases or less, 3 bases or less, or 2 bases or less. 43. The method according to Embodiment 41 or Embodiment 42, wherein the recording tag or extended recording tag includes a spacer which is a cleavage site overhang generated by a double-stranded nucleic acid cleavage reagent. 44. The method according to any one of embodiments 41 to 43, wherein the spacer is a cycle-specific spacer or a cycle-alternating spacer. 45. The method according to embodiment 44, wherein the method self-terminates after one cycle of information transfer due to an incompatible overhang. 46. ​​The method according to any one of embodiments 41 to 45, wherein if the spacers added by the previous coding tag match at least a portion of the spacers of the second coding tag, information is transferred from the second coding tag to the decompression record tag. 47. The method according to any one of Embodiments 1 to 47, wherein the recording tag and coding tag do not include spacers. 48. The method according to any one of Embodiments 1 to 47, wherein the coding tag has a 3' end available for the reaction. 49. The method according to Embodiment 48, wherein the binder is attached to the 5' end of the coding tag. 50. The method according to any one of Embodiments 1 to 49, wherein the method includes one or more washing steps. 51. The method according to Embodiment 50, wherein the washing step is performed before step (c). 52. The method according to Embodiment 50, wherein the washing step is performed before step (d). 53. The method according to Embodiment 50, wherein the washing step is performed before step (e). 54. The method according to any one of Embodiments 1 to 53, wherein the nucleic acid conjugation reagent is a chemical ligation reagent or an enzyme ligation reagent. 55. The method according to Embodiment 54, wherein the enzyme nucleic acid conjugation reagent is a ligase. 56. The method according to any one of Embodiments 1 to 55, wherein the double-strand nucleic acid cleavage reagent is a restriction enzyme. 57. The method according to Embodiment 55, wherein the restriction enzyme is an IIS-type restriction enzyme. 58. IIS-type restriction enzymes, 5'...GCAGTGNN...3' The method according to Embodiment 57, which recognizes a series of bases including 3'...CGTCACNN...5'. 59. The method according to Embodiment 57 or Embodiment 58, wherein the restriction enzyme is Nb.BtsI or BtsI-v2 or a derivative thereof. 60. The method according to any one of Embodiments 1 to 59, wherein the ligation, extension, and cutting in steps (c), (d), and (e) occur in a stepwise manner, respectively. 61. The method according to any one of Embodiments 1 to 60, wherein after the recording tag is conjugated to the coding tag by a nucleic acid conjugation reagent, the binder does not remain bound to the polymer. 62. The method according to any one of Embodiments 1 to 61, wherein the final information transfer cycle is performed to provide a capping sequence to the extended record tag, and optionally the capping sequence includes a universal priming region for amplification, sequencing, or both. 63. The method according to Embodiment 62, wherein the capping sequence is provided with a coding tag associated with a binder that can bind to universal features of a polymer. 64. The method according to Embodiment 63, wherein the universal feature is the chemical modification of the N-terminal amino acid of the polypeptide. 65. The method according to any one of embodiments 8 to 64, wherein the decompressed recoding tag includes a set of information transferred from multiple coding tags. 66. The method according to any one of embodiments 1 to 65, further comprising analyzing one or more of the elongation record tags. 67. The method according to embodiment 66, wherein one or more of the extension record tags are amplified before analysis. 68. The method of Embodiment 66 or Embodiment 67, wherein the analysis of the extension record tag includes a nucleic acid sequencing method. 69. The method according to Embodiment 68, wherein the nucleic acid sequencing method is synthesis sequencing, ligation sequencing, hybridization sequencing, Polony sequencing, ion semiconductor sequencing, pyrosequencing, single-molecule real-time sequencing, nanopore-based sequencing, or direct imaging of DNA using an advanced microscope. 70. The method according to any one of Embodiments 1 to 69, wherein the binder is a polypeptide or a protein. 71. The method according to Embodiment 70, wherein the binder is an aminopeptidase or a variant, mutant, or modified protein thereof, aminoacyl-tRNA synthetase or a variant, mutant, or modified protein thereof, antikalin or a variant, mutant, or modified protein thereof, ClpS, ClpS2 or a variant, mutant, or modified protein thereof, UBR box protein or a variant, mutant, or modified protein thereof, or a modified small molecule that binds to an amino acid, i.e., vancomycin or a variant, mutant, or modified molecule thereof, or an antibody or its binding fragment, or any combination thereof. 72. The method according to any one of Embodiments 8 to 71, wherein the binder is bound to a single amino acid residue, dipeptide, tripeptide, or post-translational modification of a polypeptide polymer. 73. The method according to embodiment 72, wherein the binder is configured to bind to the N-terminal amino acid residue of the polypeptide. 74. The method according to Embodiment 73, wherein the binder is configured to bind to a chemically modified or labeled N-terminal amino acid residue of the polypeptide. 75. A kit for polymer analysis, A binder including a coding tag, wherein the coding tag includes identification information relating to the binder, nucleic acid conjugation reagent, polymerase and A double-strand nucleic acid cleavage reagent is included, A kit in which a binder is configured to bind to a polymer associated with a recording tag, and identification information from a coding tag is configured for transfer from the coding tag to the recording tag associated with the polymer. 76. The kit according to Embodiment 75, wherein the polymer is a lipid or a carbohydrate. 77. The kit according to Embodiment 75, wherein the polymer is a protein, polypeptide, or peptide. 78. The kit according to Embodiment 77, further comprising a reagent for removing the N-terminal amino acid (NTAA) of a polypeptide. 79. The kit according to Embodiment 78, wherein the reagent for removing terminal amino acids of a polypeptide comprises a chemical agent or an enzyme. 80. The kit according to any one of embodiments 77 to 79, further comprising a reagent for modifying the terminal amino acids of a polypeptide. 81. The kit according to Embodiment 80, wherein the reagent for modifying the terminal amino acids of a polypeptide comprises a chemical agent or an enzyme. 82. A kit according to any one of embodiments 75 to 81, wherein the recording tag associated with the polymer includes a double-stranded region. 83. The kit according to Embodiment 82, wherein the recording tag associated with the polymer includes a nucleic acid hairpin. 84. A kit according to any one of embodiments 75 to 83, wherein the recording tag associated with the polymer has a 3' overhang. 85. A kit according to any one of embodiments 75 to 84, wherein the recording tag includes a barcode. 86. The kit according to Embodiment 85, wherein the barcode is a sample barcode, a fraction barcode, a spatial barcode, and / or a compartment tag. 87. A kit according to any one of embodiments 75 to 86, wherein the recording tag includes a unique molecular identifier (UMI). 88. A kit according to any one of embodiments 75 to 87, wherein the recording tag includes a universal priming portion. 89. The kit according to Embodiment 88, wherein the universal priming site includes a priming site for amplification, sequencing, or both. 90. The kit according to any one of Embodiments 75 to 89, wherein the kit comprises a mixture comprising a nucleic acid conjugation reagent, a polymerase, and a double-strand nucleic acid cleavage reagent. 91. A kit according to any one of embodiments 75 to 90, wherein the coding tag includes a partial restriction enzyme recognition sequence. 92. The kit according to Embodiment 91, wherein the partially restriction enzyme recognition sequence of the coding tag is single-stranded. 93. A kit according to any one of embodiments 75 to 92, wherein the coding tag includes a barcode and / or a unique molecular identifier (UMI). 94. A kit according to any one of embodiments 75 to 93, wherein the recording tag and the coding tag each include a spacer. 95. The kit according to Embodiment 94, wherein the spacer is a nucleic acid molecule with 10 bases or less, 9 bases or less, 8 bases or less, 7 bases or less, 6 bases or less, 5 bases or less, 4 bases or less, 3 bases or less, or 2 bases or less. 96. The kit according to Embodiment 94 or Embodiment 95, wherein the spacer is a cycle-specific spacer or a cycle-alternating spacer. 97. A kit according to any one of embodiments 75 to 93, wherein the recording tag and coding tag do not include spacers. 98. The kit according to any one of embodiments 75 to 97, wherein the coding tag has a 3' end available for the reaction. 99. The kit according to any one of Embodiments 75 to 98, wherein the enzyme nucleic acid conjugation reagent is a ligase. 100. A kit according to any one of Embodiments 75 to 99, wherein the double-strand nucleic acid cleavage reagent is a restriction enzyme. 101. The kit according to Embodiment 100, wherein the restriction enzyme is an IIS-type restriction enzyme. 102. IIS-type restriction enzymes, 5'...GCAGTGNN...3' A kit according to Embodiment 101 that recognizes a series of bases including 3'...CGTCACNN...5'. 103. The kit according to Embodiment 101 or Embodiment 102, wherein the restriction enzyme is Nb.BtsI or BtsI-v2 or a derivative thereof. 104. The kit according to any one of embodiments 75 to 103, further comprising a binder capable of binding to universal features of a polymer, including a capping sequence. 105. A kit according to any one of embodiments 75 to 104, wherein the kit comprises a mixture containing multiple binders. 106. The kit according to any one of embodiments 75 to 105, wherein the binder is a polypeptide or a protein. 107. The kit according to Embodiment 106, wherein the binder is an aminopeptidase or a variant, mutant, or modified protein thereof, an aminoacyl-tRNA synthetase or a variant, mutant, or modified protein thereof, anticharin or a variant, mutant, or modified protein thereof, ClpS, ClpS2 or a variant, mutant, or modified protein thereof, a UBR box protein or a variant, mutant, or modified protein thereof, or a modified small molecule that binds to an amino acid, i.e., vancomycin or a variant, mutant, or modified molecule thereof, or an antibody or its binding fragment, or any combination thereof. 108. The kit according to any one of Embodiments 77 to 107, wherein the binder binds to a single amino acid residue, dipeptide, tripeptide, or post-translational modification of the polypeptide polymer. 109. The kit according to Embodiment 108, wherein the binder is configured to bind to the N-terminal amino acid residue of the polypeptide. 110. The kit according to any one of Embodiments 80 to 109, wherein the binder is configured to bind to an N-terminal amino acid residue of a polypeptide, which is modified by a reagent to modify the terminal amino acids of the polypeptide. 111. A kit according to any one of embodiments 75 to 110, further comprising one or more washing buffers. 112. A kit according to any one of embodiments 75 to 111, further comprising a support for immobilizing polymers and / or recording tags. 113. The kit according to Embodiment 112, wherein the support is a three-dimensional support (e.g., a porous matrix or beads). 114. The kit according to Embodiment 113, wherein the support comprises polystyrene beads, polyacrylate beads, polymer beads, agarose beads, cellulose beads, dextran beads, acrylamide beads, solid core beads, porous beads, paramagnetic beads, glass beads, controlled porous beads, silica-based beads, or any combination thereof. [Examples]

[0212] V. Examples The following examples are provided to illustrate, but are not limited to, the methods, compositions, and uses provided herein. Specific aspects of the present invention, including, but not limited to, embodiments of the Proteocode® polypeptide sequencing assay, information transfer between coding tags and recording tags, methods for preparing nucleotide-polypeptide conjugates, methods for attaching nucleotide-polypeptide conjugates to a support, methods for generating barcodes, methods for preparing specific binders that recognize the N-terminal amino acids of polypeptides, reagents, and methods for modifying and / or removing N-terminal amino acids from polypeptides, and methods for analyzing extended recording tags, are disclosed in previously published applications US2019 / 0145982A1, US2020 / 0348308A1, US2020 / 0348307A1, US2021 / 0208150A1, and WO2020 / 223000, the entire contents of which are incorporated herein by reference.

[0213] Example 1. Evaluation of analyte immobilization using nucleic acid hybridization and bonding to a solid support. This example illustrates an exemplary method for attaching (immobilizing) nucleic acid-polypeptide conjugates to a solid support.

[0214] In the hybridization-based immobilization method, nucleic acid-polypeptide complexes were hybridized and ligated onto chemically immobilized hairpin-captured DNA on magnetic beads. The captured nucleic acids were conjugated to the beads using trans-cyclooctene (TCO) and methyltetrazine (mTet) based click chemistry. Short hairpin-captured nucleic acids (16-base pair stem, 5-base loop, 24-base 5' overhang) modified with TCO were reacted with mTet-coated magnetic beads. Phosphorylated nucleic acid-polypeptide complexes (10 nM) were annealed to the hairpin DNA attached to the beads with 5× SSC and 0.02% SDS, and incubated at 37°C for 30 minutes. The beads were washed once with PBST and resuspended in 1× Quick ligation solution (New England Biolabs, USA) containing T4 DNA ligase. After incubation at 25°C for 30 minutes, the beads were washed twice with PBST and resuspended in 50 μL of PBST. Immobilized total nucleic acid-peptide conjugates containing aminoFA-terminal peptide (FAGVAMPGAEDDVVGSGSK, SEQ ID NO: 3), aminoAFA-terminal peptide (AFAGVAMPGAEDDVVGSGSK, SEQ ID NO: 4), and aminoAA-terminal peptide (AAGVAMPGAEDDVVGSGSK, SEQ ID NO: 5) were quantified by qPCR using a specific set of primers. For comparison, peptides were immobilized on beads using a non-hybridization-based method without a ligation step. The non-hybridization-based method was performed by incubating DNA-tagged peptides modified with 30 μM TCO, containing aminoFA-terminal peptide, aminoAFA-terminal peptide, and aminoAA-terminal peptide, with mTet-coated magnetic beads overnight at 25°C.

[0215] As shown in Table 1, similar Ct values ​​were observed for the non-hybridization preparation method with a graft density of 1:100,000 and the hybridization-based preparation method with a graft density of 1:10,000. The loading amount of DNA-tagged peptide in the hybridization-based preparation method was 1 / 3000 compared to the non-hybridization preparation method. In general, it was observed that the hybridization-based immobilization method required less starting material. [Table 1]

[0216] Example 2. Exemplary serial encoding assay. In this example, two exemplary binders were used to specifically bind to nucleic acid-polypeptide conjugates immobilized on a solid support. One binder binds to the N-terminal phenylalanine residue of the polypeptide (F-binder, 31-F), and the other binder binds to the N-terminal leucine residue of the polypeptide (L-binder, 44-L). Both binders were manipulated from a lipocalin scaffold by directed evolution, as specifically described in Example 1 of US2021 / 0208150A1. The binders were conjugated to corresponding nucleic acid coding tags containing barcodes with identification information relating to the binder. The coding tags specific to each binder were attached to SpyTag via a PEG linker, and the resulting fusions were reacted with a binder-SpyCatcher fusion protein via a SpyTag-SpyCatcher interaction, essentially as described in US2021 / 0208150A1.

[0217] For the encoding assay, two test polymers (FSGVARGDVRGGK(azide), hereafter referred to as F-peptide, SEQ ID NO: 6, and LAESAFSGVARGDVRGGK(azide), hereafter referred to as L-peptide, SEQ ID NO: 7) were conjugated to immobilized bead-attached captured DNA (SEQ ID NO: 8) via an alkyne-azide reaction (essentially as described in Example 1). The DNA-polypeptide conjugate (20 nM) was annealed to the captured DNA attached to the beads with 5 × SSC and 0.02% SDS, and incubated at 37°C for 30 minutes. The beads were washed once with PBST and resuspended in 1 × Quick ligation solution (New England Biolabs, USA) containing T4 DNA ligase. After incubation at 25°C for 30 minutes, the beads were washed with PBST, twice with 0.1 M NaOH + 0.1% Tween® 20, and twice with PBST. The recording tag in the conjugate contains a barcode in the polymer, a 2nt overhang complementary region, a type II restriction enzyme binding region, and an adjacent region. After hybridization and ligation, the sample was treated with Krenow fragments (3'->5' exo-) (MCLAB, USA), BtsI-V2 (0.5 units / uL, New England Biolabs, USA), dNTPs (125uM each), and CutSmart buffer (50mM potassium acetate, 20mM tris acetate, 10mM magnesium acetate, 100μg / ml BSA, pH 7.9, New England Biolabs, USA). The samples were incubated with Biolabs (USA) at 25°C for 30 minutes and washed with PBST, two washes of 0.1M NaOH + 0.1% Tween® 20, and two washes of PBST. As a result, a 2nt overhang was formed at 3' (see Figure 1A).

[0218] The beads were treated with 18 mM PMI reagent (pyrazolemethanimine, 4-trifluoromethyl-1H-pyrazole) in DMA and a MOPS mixture (60% DMA and 40% MOPS, pH 7.6) at 25°C for 30 minutes to modify the N-terminus of the immobilized peptide, and washed three times with PBST. The coding tags attached to the F-binding and L-binding agents each formed a loop with an 8 bp double helix and a 2 nt overhang at 3', which is complementary to the 3' overhang of the recording tag on the bead. The coding tag contains a unique barcode for identifying the binding agent and also has a BtsI-V2 binding sequence and a 2 nt complementary overhang region for the next binding cycle. Two binders (100 nM each) were incubated with beads at 25°C for 15 minutes in quick Ligase buffer (66 mM Tris-HCl, 10 mM MgCl2, 1 mM dithiothreitol, 1 mM ATP, 7.5% polyethylene glycol (PEG6000), pH 7.6) at 0.125 units / µL of Krenow fragments (3'->5' exo-) (MCLAB, USA), dNTP mixture (125 µM each), 0.5 units / µL of BtsI-V2, and different concentrations of T4 DNA ligase (New England Biolabs, USA). Subsequently, the mixtures were washed once with PBST, twice with 0.1 M NaOH + 0.1% Tween® 20, and twice with PBST. During the merging process, the record tag is extended, the barcode information from the coding tag is transferred to the record tag, and an extended record tag is formed (see Figure 1B).

[0219] Following the binding cycle, capping was performed to introduce primer sites for downstream PCR to amplify the extension record tag for analysis. Capping oligos containing loop DNA with a 2nt 3' overhang complementary to the 3' overhang of the extension or non-extension record tag were introduced. Instead of performing separate capping steps, longer coding tags containing complementary primer sequences can be used to introduce the primer sites for downstream PCR during the extension reaction. 400 nM capping oligos were incubated with beads in quick Ligase buffer (66 mM Tris-HCl, 10 mM MgCl2, 1 mM dithiothreitol, 1 mM ATP, 7.5% polyethylene glycol (PEG6000), pH 7.6) in the presence of 12.5 units / uL T4 DNA ligase for 15 minutes at 25°C. Samples were washed with one PBST, two 0.1 M NaOH + 0.1% Tween® 20, and two PBST washes. The assay extension record tags were subjected to PCR amplification and analyzed by next-generation sequencing (NGS) to reveal barcode information related to the binder that interacted with the polymer.

[0220] Figure 2 shows exemplary encoding results generated by the described encoding method. In the ligation step, three different concentrations of T4 DNA ligase (0.125, 1.25, and 12.5 units / µl) were tested. After ligation, the Krenow fragment was added for the extension step, and the BtsI-V2 enzyme was added for the cleavage step. The fraction of encoded record tags (the ratio of extended record tags to the total amount of record tags on the beads (extended and unextended)) was evaluated by NGS sequencing, and specific encoding results for both binders are shown.

[0221] Example 3. Encoding assay for other types of polymers. The data shown in Figure 2 represents the analysis of the extension record tag, confirming the successful transfer of barcode information from the coding tag to the record tag after the F-binding or L-binding agent binds to the peptide polymer. The described technique can be employed for other types of polymers, such as lipids, carbohydrates, or macrocyclic molecules. To perform the encoding assay, such polymers must be immobilized on a solid support (e.g., beads) and associated with nucleic acid record tags. The encoding step remains the same regardless of the type of immobilized polymer. Association with the record tag can be direct (e.g., covalent binding) or indirect (e.g., association via a solid support). In the latter case, the record tag must co-localize with or be in close proximity to the polymer during the encoding assay. Binding agents can be selected to specifically bind to components of the polymer. Each binding agent must be conjugated to a corresponding nucleic acid coding tag containing a barcode with identification information about the binding agent. During encoding, the barcode information is transferred to the record tag associated with the polymer, an extension record tag is generated, and the binding history of the polymer is recorded within the extension record tag. The binding cycle can be repeated multiple times, either separately or in a mixture, using different binders that interact with the polymer. Below, representative methods known in the art are disclosed that can be used to adapt the disclosed encoding assays for different types of polymers, such as carbohydrates, lipids, or macrocyclic molecules.

[0222] First, exemplary binders are known that can specifically bind to components of carbohydrates, lipids, or macrocyclic molecules. For example, lectins are carbohydrate-binding proteins that can selectively recognize glycan epitopes of free carbohydrates or glycoproteins and can be used as specific binders for carbohydrate-containing macromolecules. Importantly, there are known lectins that recognize different components of carbohydrates, such as mannose-binding lectins, galactose / N-acetylglucosamine-binding lectins, sialic acid / N-acetylglucosamine-binding lectins, and fucose-binding lectins (disclosed, e.g., in WO2012 / 049285A1). Lipid-binding proteins are also well known and can be used as binders (see, e.g., Bernlohr DA, et al., Intracellular lipid-binding proteins and their genes. Annu Rev Nutr. 1997;17:277-303). Lipid-binding antibodies are commonly known and can be used as binders for lipid-containing macromolecules (see, for example, Alving CR. Antibodies to lipids and liposomes: immunology and safety. J Liposome Res. 2006;16(3):157-66). Furthermore, proteins that specifically bind to macrocyclic molecules are also known (see, for example, Villar EA, et al., How proteins bind macrocycles. Nat Chem). Biol.2014 Sep;10(9):723-31; Hunter TM, et al., Protein recognition of macrocycles:binding of anti-HIV metallocyclams to lysozyme.Proc Natl Acad Sci US A.2005 Feb 15;102(7):2288-92).

[0223] Secondly, an exemplary carbohydrate detection encoding assay can be carried out using methods known in the art, as follows:

[0224] Approach I. Reductive Amination (Based on Yang SJ, Zhang H. Glycan analysis by reversible reaction to hydrazide beads and mass spectrometry. Anal Chem. 2012; 84(5):2232-2238). (a) Generates an immobilized, record tag-attached carbohydrate conjugate. The carbohydrate is oxidized with sodium periodate to produce an aldehyde. An amine-terminated DNA recording tag is conjugated, and the resulting imine is reduced using sodium borocyanohydride to produce a carbohydrate recording tag conjugate. Preferably, a hydrazide, alkoxyamine, or similarly reactive DNA can be used to produce a more stable reaction product (e.g., a hydrazone) that does not require a reducing agent. The DNA-coupled carbohydrate is immobilized on a solid support via the DNA recording tag, as described in Example 2. (b) By utilizing the aforementioned SpyCatcher-concanavalin A (ConA) fusion, a lectin-DNA coding tag (binding agent coding tag) conjugate is generated. The coding tag contains a barcode containing identification information related to ConA. (c) As described in Example 2, the carbohydrate is analyzed to determine whether it contains a component that binds to ConA by transferring barcode information from the lectin-related coding tag to the recording tag.

[0225] Approach II. Diazo coupling (based on Matsuura K, et al., Facile synthesis of stable and lectin-recognizable DNA-carbohydrate conjugates via diazo coupling. Bioconjug Chem. 2000 Mar-Apr;11(2):202-11). In Approach II, step (a) (immobilization of the record tag-attached carbohydrate conjugate) can be carried out as follows: 1) Amination of the carbohydrate with ammonium bicarbonate in water to produce a β-glycosylamine; 2) Convert the amine to a carboxylate derivative having a nitrophenyl functional group and an amide. Hydrogenate the nitro group on a palladium catalyst and treat with NaNO2 and HCl to provide the diazo compound. Steps (b) and (c) are the same as in Approach I.

[0226] Thirdly, exemplary lipid detection encoding assays can be carried out using methods known in the art as follows:

[0227] Approach I. Fatty Acids (Hiroshi Miwa, High-performance liquid chromatographic determination of free fatty acids and esterified fatty acids in biological materials as their 2-nitrophenylhydrazides, Analytica Chimica Acta, Volume 465, Issues 1-2, 2002, Pages 237-255, ISSN 0003-2670). (a) Extract fatty acids from a biological source and activate the carboxylic acids by EDC / CDI chemistry. Couple with amine-terminated or hydrazide-terminated DNA recording tags to produce recording tag-attached lipid conjugates. Immobilize the DNA-coupled lipids onto a solid support via the DNA recording tags as described in Example 2.

[0228] Approach II. Reactive lipid (based on X. Wei & H. Yin (2015)) Covalent modification of DNA by α,β unsaturated aldehydes derived from lipid peroxidation: Recent progress and challenges, Free Radical Research, 49:7, 905-917). (a) Obtain a reactive lipid substrate such as malondialdehyde (MDA) or 4-hydroxynonenal (HNE), and couple the hydrazide-terminated DNA recording tag to the reactive lipid species. Alternatively, couple the amine-terminated DNA recording tag to the aldehyde on the reactive lipid, and reduce the resulting imine with sodium cyanoborohydride. The next step in both approaches is to generate a binder-DNA coding tag conjugate by utilizing the aforementioned SpyCatcher binder fusion. The coding tag contains a barcode with identification information relating to the binder. Fatty acid-binding proteins (FABPs), other lipid-binding proteins, or lipid-binding antibodies can be used as binders. Finally, as described in Example 2 of this application, the barcode information is transferred from the binder-associated coding tag to a recording tag, and thus analyzed to determine whether the lipid contains a component that binds to the binder.

[0229] Fourth, McElhiney J, et al., Rapid isolation of a single-chain antibody against the cyanobacterial toxin microcystin-LR by phage display and its use in the immunoaffinity concentration of microcystins from Based on water.Appl Environ Microbiol.2002 Nov;68(11):5288-95, an exemplary macrocycle (microcystin) detection encoding assay can be performed as follows using methods known to those skilled in the art. (a) React the dehydroalanine of microcystin with 2-mercaptoethylamine to generate a primary amine, and then couple the DNA recording tag to the primary amine using an amine-reactive DNA recording tag (e.g., NHS-DNA derivative) to generate a DNA recording tag-conjugated microcystin. (b) Generate a single-stranded antibody-SpyCatcher binder that recognizes microcystin. The production of the single-stranded antibody is described in McElhiney J, et al. 2002. Couple a DNA coding tag to a SpyTag (the coding tag contains a barcode with identification information regarding the single-stranded antibody), and then react with SpyCatcher to generate a binder coding tag conjugate as described above. (c) Transfer the barcode information from the single-stranded antibody-related coding tag to the recording tag as described in Example 2, and thus analyze whether the polymer contains microcystin. [Table 2] The present invention provides, for example, the following items. (Item 1) A method for analyzing a polymer, comprising: (a) providing a polymer and a related recording tag conjugated to a support; (b) contacting the polymer with a binder capable of binding to the polymer, the binder comprising a coding tag having identification information regarding the binder, to enable binding between the polymer and the binder; (c) conjugating the 5'-end of the recording tag to the 3'-end of the coding tag by a nucleic acid conjugation reagent; (d) extending the recording tag using the coding tag as a template by a polymerase to generate a double-stranded extended recording tag; (e) cleaving the double-stranded extended recording tag with a double-stranded nucleic acid cleavage reagent to generate a 3' overhang within the extended recording tag. A method comprising transferring information from the coding tag to the recording tag to generate the extended recording tag. (Item 2) The method according to item 1, wherein in step (d), the double-stranded extended recording tag comprises a recognition sequence that can be recognized by the double-stranded nucleic acid cleavage reagent. (Item 3) The method according to item 1, wherein the cleavage in step (e) releases the binder from the polymer. (Item 4) The method according to item 1, wherein steps (b), (c), (d), and (e) are repeated continuously one or more times in a periodic manner. (Item 5) The method according to item 1, wherein the polymer is a polypeptide. (Item 6) The method according to item 5, wherein the polypeptide is obtained by fragmenting proteins from a biological sample. (Item 7) The method according to item 1, wherein the polymer for analysis is not a nucleic acid. (Item 8) The method according to item 4, further comprising removing a portion of the polymer before repeating step (b). (Item 9) The method according to item 5, further comprising removing the N-terminal amino acid (NTAA) of the polypeptide to expose a new NTAA of the polypeptide before repeating step (b). (Item 10) The method according to item 4, wherein in step (e), the 3' overhang of the extended recording tag generated by the double-stranded nucleic acid cleavage reagent is available for hybridization with a second coding tag when step (b) is repeated. (Item 11) The method according to item 1, wherein the method comprises contacting a plurality of polymers with a single binder or a plurality of binders in step (b). (Item 12) The method of item 5, further comprising treating the polypeptide with a reagent for modifying the terminal amino acids of the polypeptide before step (b). (Item 13) The method according to item 1, wherein the recording tag associated with the polymer includes a nucleic acid hairpin. (Item 14) The method according to item 1, wherein the nucleic acid conjugation reagent of step (c), the polymerase of step (d), and the double-strand nucleic acid cleavage reagent of step (e) are provided simultaneously. (Item 15) The method according to item 1, wherein the double-strand nucleic acid cleavage reagent is an IIS-type restriction enzyme. (Item 16) The method according to item 1, further comprising analyzing one or more of the extension record tags, wherein the analysis of the extension record tags includes a nucleic acid sequencing method. (Item 17) A kit for polymer analysis, A binder including a coding tag, wherein the coding tag includes identification information relating to the binder, nucleic acid conjugation reagent, polymerase and A double-strand nucleic acid cleavage reagent is included, A kit wherein the binder is configured to bind to a polymer associated with a recording tag, and the identification information from the coding tag is configured for transfer from the coding tag to the recording tag associated with the polymer. (Item 18) The kit according to item 17, wherein the polymer is a polypeptide. (Item 19) The kit according to item 18, further comprising a reagent for removing the N-terminal amino acid (NTAA) of the polypeptide. (Item 20) The kit according to item 17, wherein the kit comprises a mixture containing the nucleic acid conjugation reagent, the polymerase, and the double-strand nucleic acid cleavage reagent. (Item 21) The kit according to item 17, further comprising the polymer and / or a support for immobilizing the recording tag.

Claims

1. A method for analyzing polymers, (a) Providing a polymer and a related recording tag bonded to a support, (b) A step of bringing the polymer into contact with a binder, wherein the binder includes a coding tag having identification information relating to the binder, and the binder binds to the polymer. (c) The step of covalently joining the 5' end of the recording tag to the 3' end of the coding tag using a nucleic acid conjugation reagent, (d) A step of using polymerase to extend the recording tag using the coding tag as a template to generate a double-stranded extended recording tag, (e) The step of cleaving both strands of the double-stranded extension record tag with a double-stranded nucleic acid cleavage reagent to generate a 3' overhang within the extension record tag, A method comprising the steps of transferring information from the coding tag to the record tag to generate the extended record tag.

2. The method according to claim 1, wherein in step (d), the double-stranded nucleic acid cleavage reagent recognizes the recognition sequence of the double-stranded extension record tag.

3. The method according to claim 1, wherein the cutting in step (e) releases the binder from the polymer.

4. The method according to claim 1, wherein steps (b), (c), (d), and (e) are repeated one or more times in a periodic manner.

5. The method according to claim 1, wherein the polymer is a polypeptide.

6. The method according to claim 5, wherein the polypeptide is obtained by fragmenting a protein from a biological sample.

7. The method according to claim 1, wherein the length of the 3' overhang in the extension record tag is less than 6 base pairs.

8. The method according to claim 4, further comprising removing a portion of the polymer before repeating step (b).

9. The method of claim 5, further comprising removing the N-terminal amino acid (NTAA) of the polypeptide to expose a new NTAA of the polypeptide before repeating step (b).

10. The method according to claim 4, wherein in step (e), the 3' overhang of the extension record tag generated by the double-stranded nucleic acid cleavage reagent is available for hybridizing with a second coding tag when step (b) is repeated.

11. The method of claim 5, further comprising treating the polypeptide with a reagent for modifying the terminal amino acids of the polypeptide before step (b).

12. The method according to claim 1, wherein the recording tag associated with the polymer includes a nucleic acid hairpin.

13. The method according to claim 1, wherein the nucleic acid conjugation reagent of step (c), the polymerase of step (d), and the double-strand nucleic acid cleavage reagent of step (e) are provided simultaneously.

14. The method according to claim 1, wherein the double-strand nucleic acid cleavage reagent is an IIS-type restriction enzyme.

15. The method according to claim 1, further comprising analyzing one or more of the extension record tags, wherein the analysis of the extension record tags includes a nucleic acid sequencing method.

16. A kit for polymer analysis, A binder including a coding tag, wherein the coding tag includes identification information relating to the binder, (i) Nucleic acid conjugation reagent, (ii) polymerase, and (iii) Double-strand nucleic acid cleavage reagent A mixture containing, A kit wherein the binder is bound to a polymer associated with a recording tag, and the identification information from the coding tag is configured for transfer from the coding tag to the recording tag associated with the polymer.

17. The kit according to claim 16, wherein the polymer is a polypeptide.

18. The kit according to claim 17, further comprising a reagent for removing the N-terminal amino acid (NTAA) of the polypeptide.

19. The kit according to claim 16, wherein the double-strand nucleic acid cleavage reagent is a restriction enzyme.

20. The kit according to claim 16, further comprising the polymer and / or a support for immobilizing the recording tag.