Amino acid binding agents and uses thereof

By developing a variety of binding agents from the binding agent library, the problems of insufficient sensitivity and selectivity in the detection of amino acids in the existing technology have been solved, and high specificity and high affinity binding to a variety of amino acids has been achieved, thereby improving the detection efficiency.

CN122122195APending Publication Date: 2026-05-29GLYPHIC BIOTECHNOLOGIES INC
View PDF 6 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
GLYPHIC BIOTECHNOLOGIES INC
Filing Date
2024-09-20
Publication Date
2026-05-29

AI Technical Summary

Technical Problem

Existing binding agents lack sufficient sensitivity and selectivity when detecting single monomeric amino acids, making it difficult to efficiently identify and specifically bind multiple amino acids or their derivatives.

Method used

A binding agent library containing various binding agents, such as antibodies, antibody fragments, and engineered scFv, was developed. This library can specifically recognize and bind to various amino acids or their derivatives. Affinity maturation was performed using methods such as yeast surface display and phage display. The binding agent library contains polyclonal and monoclonal antibodies, and detection was performed using ELISA, surface plasmon resonance, and biofilm interferometry.

Benefits of technology

It achieves highly specific and high-affinity binding to a variety of amino acids or their derivatives, and can recognize and bind to specific amino acids at low concentrations, thus improving the sensitivity and selectivity of detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122122195A_ABST
    Figure CN122122195A_ABST
Patent Text Reader

Abstract

Provided herein are binding agents that selectively bind to individual monomer types (e.g., amino acid types) of a polymeric analyte (e.g., a peptide). The present disclosure provides methods for producing and, optionally, engineering binding agents. Such binding agents can be used for a variety of applications, including single-molecule peptide or protein sequencing.
Need to check novelty before this filing date? Find Prior Art

Description

Cross-referencing

[0001] This application claims the benefit of U.S. Provisional Patent Application No. 63 / 584,382, filed September 21, 2023; U.S. Provisional Patent Application No. 63 / 588,867, filed October 9, 2023; and U.S. Provisional Patent Application No. 63 / 589,391, filed October 11, 2023, each of which is incorporated herein by reference in its entirety. Statement regarding federally funded research or development

[0002] This invention was developed with the support of the U.S. government, pursuant to authorization numbers HG012960 and GM148298 granted by the National Institutes of Health. The government holds certain rights to this invention. sequence list

[0003] This application includes a sequence list, which has been submitted electronically in XML format and is incorporated herein by reference in its entirety. The XML copy was created on September 16, 2024, named 60652-708_601_SL.xml, and is 106,896 bytes in size. Background Technology

[0004] Characterization of polymeric analytes (such as proteins) is essential for understanding biology and pathology, as well as for the development of diagnostics and therapies. De novo sequencing of polymeric analytes (e.g., proteins) is necessary for protein discovery, determination of structure-function relationships, and identification of biomarker targets.

[0005] Binders that identify and bind to polymer analytes can be used to identify the presence of polymer analytes in a sample. However, current binders are limited in their ability to sensitively and selectively detect individual monomer types (e.g., amino acid types). Therefore, new binders are needed for detecting each monomer type of a given polymer analyte. Summary of the Invention

[0006] This document recognizes the need for novel binders capable of recognizing and binding to individual monomers of a polymer analyte. In some embodiments, this disclosure provides binders capable of recognizing and binding to a single amino acid or a derivative thereof. The derivatives may include modified amino acids, such as chemically modified or enzymatically modified amino acids.

[0007] In one aspect, this document provides a binding agent that binds to a type of protein amino acid or a derivative thereof with higher specificity or higher affinity than with all other types of protein amino acids or derivatives thereof, wherein the type of amino acid is not tryptophan. In other aspects, this document provides a binding agent that binds to a type of protein amino acid or a derivative thereof in a group with higher specificity or higher affinity than with all other types of protein amino acids or derivatives thereof, wherein the type of amino acid is not tryptophan.

[0008] In some embodiments, the group of two or more protein amino acids or their derivatives comprises at least 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 different protein amino acids or their derivatives. In some embodiments, the group of two or more protein amino acids or their derivatives comprises naturally occurring protein amino acids. In some embodiments, the group of two or more protein amino acids or their derivatives comprises 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 standard protein amino acids or their derivatives. In some embodiments, the group of two or more protein amino acids or their derivatives consists of 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19 or 20 standard protein amino acids or their derivatives.

[0009] In some embodiments, the binding agent is part of a binding agent library, wherein the binding agent library contains multiple binding agents that collectively recognize and specifically bind to at least five different types of protein amino acids. In some embodiments, the binding agent is part of a binding agent library, wherein the binding agent library contains multiple binding agents that collectively recognize and specifically bind to all 20 types of standard protein amino acids. In some embodiments, the binding agents of the binding agent library can recognize and specifically bind to at least one post-translational modified amino acid. In some embodiments, the binding agent library contains multiple polyclonal antibodies. In some embodiments, the binding agent library contains multiple monoclonal antibodies.

[0010] In some embodiments, the conjugate comprises an antibody, an antibody fragment, a single-chain variable fragment (scFv), or a nanobody. In some embodiments, the conjugate comprises an antibody. In some embodiments, the antibody is immunoglobulin G (IgG). In some embodiments, the antibody is generated by immunizing an animal with an immunogenically effective amount of a composition comprising a carrier protein conjugated to an amino acid or an amino acid derivative. In some embodiments, the composition comprises a carrier protein containing a polymer linker conjugated to an amino acid or an amino acid derivative. In some embodiments, the polymer linker comprises polyethylene glycol (PEG). In some embodiments, the antibody is a monoclonal antibody. In some embodiments, the conjugate comprises an engineered scFv. In some embodiments, the engineered scFv is derived from IgG. In some embodiments, the engineered scFv is engineered using directed evolution. In some embodiments, the engineered scFv undergoes affinity maturation using yeast surface display, phage display, ribosome display, or continuous evolution. In some embodiments, the engineered scFv exhibits improved specificity for a type of protein amino acid or its derivative by using yeast surface display.

[0011] In some embodiments, the derivative or a derivative thereof comprises a chemically modified amino acid. In some embodiments, the chemically modified amino acid comprises a hydantoin, phenylthiocarbamoyl, or anilinethiazolinone derivative of an amino acid. In some embodiments, the chemically modified amino acid comprises an amino acid coupled to a linker. In some embodiments, the linker is coupled to the amino acid at the N-terminus. In some embodiments, the linker coupled to the chemically modified amino acid comprises a hydantoin, phenylthiocarbamoyl, or anilinethiazolinone moiety. In some embodiments, the linker is coupled to the amino acid at the C-terminus. In some embodiments, the linker comprises a linker to a nucleic acid molecule. In some embodiments, the linker comprises an amino acid reactive group and a linker to a nucleic acid molecule.

[0012] In some embodiments, an indirect enzyme-linked immunosorbent assay (ELISA) is used to measure the specific binding of the binding agent to one type of protein amino acid or a derivative thereof, rather than to all other types of protein amino acids or their derivatives.

[0013] In some embodiments, surface plasmon resonance is used to measure the specific binding of a binding agent to one type of protein amino acid or a derivative thereof, rather than to all other types of protein amino acids or their derivatives.

[0014] In some embodiments, biomembrane interferometry is used to measure the specific binding of the binding agent to one type of protein amino acid or its derivative, rather than to all other types of protein amino acids or their derivatives.

[0015] In some embodiments, the binder comprises a detectable label. In some embodiments, the detectable label comprises a nucleic acid molecule, a fluorophore, a mass tag, or a protein. In some embodiments, the detectable label comprises a protein, wherein the protein contains an additional antibody or an additional antibody fragment. In some embodiments, the detectable label comprises a nucleic acid molecule, wherein the nucleic acid molecule contains a barcode sequence. In some embodiments, the detectable label is conjugated to the binder using a portion selected from the group consisting of: biotin, dethiobiotin, avidin, streptavidin, neutral avidin, SpyCatcher, SpyTag, SNAP tag, click chemistry portion, and cysteine ​​tag. In some embodiments, the binder comprises atypical amino acids, wherein the detectable label is conjugated to the binder via atypical amino acids. In some embodiments, the atypical amino acids include click chemistry portions.

[0016] In some embodiments, the binder specifically binds to phenylalanine or a derivative thereof, but not to all other types of protein amino acids or their derivatives. In some embodiments, the binder specifically binds to phenylalanine or a derivative thereof, but not to all other types of protein amino acids or their derivatives in the group. In some embodiments, the binder specifically binds to leucine or a derivative thereof, but not to all other types of protein amino acids or their derivatives. In some embodiments, the binder specifically binds to leucine or a derivative thereof, but not to all other types of protein amino acids or their derivatives in the group. In some embodiments, the binder specifically binds to valine or a derivative thereof, but not to all other types of protein amino acids or their derivatives. In some embodiments, the binder specifically binds to valine or a derivative thereof, but not to all other types of protein amino acids or their derivatives in the group. In some embodiments, the binder specifically binds to tyrosine or a derivative thereof, but not to all other types of protein amino acids or their derivatives in the group. In some embodiments, the binder specifically binds to tyrosine or a derivative thereof, but not to all other types of protein amino acids or their derivatives in the group. In some embodiments, the binder specifically binds to proline or a derivative thereof, rather than to all other types of protein amino acids or their derivatives.

[0017] In some implementation schemes, the binder is derived from mice.

[0018] On the other hand, this article provides an antibody fragment that binds specifically to one type of protein amino acid or its derivative in a group with higher specificity or higher affinity than to all other types of protein amino acids or their derivatives in the group of two or more protein amino acids or their derivatives.

[0019] In some embodiments, the group of two or more protein amino acids or their derivatives comprises at least 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 different protein amino acids or their derivatives. In some embodiments, the group of two or more protein amino acids or their derivatives comprises naturally occurring protein amino acids. In some embodiments, the group of two or more protein amino acids or their derivatives comprises 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 standard protein amino acids or their derivatives. In some embodiments, the group of two or more protein amino acids or their derivatives consists of 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19 or 20 standard protein amino acids or their derivatives.

[0020] In some embodiments, the antibody fragment is part of an antibody or antibody fragment library, wherein the antibody or antibody fragment library contains multiple antibodies or antibody fragments that collectively recognize and specifically bind to at least five different types of protein amino acids or derivatives thereof. In some embodiments, the antibody fragment is part of an antibody or antibody fragment library, wherein the antibody or antibody fragment library contains multiple antibodies or antibody fragments that collectively recognize and specifically bind to all 20 types of standard protein amino acids or derivatives thereof. In some embodiments, the binding agent of the antibody or antibody fragment library comprises multiple polyclonal antibodies. In some embodiments, the antibody or antibody fragment library comprises multiple monoclonal antibodies.

[0021] In some embodiments, the antibody fragment is derived from an antibody. In some embodiments, the antibody is immunoglobulin G (IgG). In some embodiments, the antibody is generated by immunizing an animal with an immunogenically effective amount of a composition comprising a carrier protein conjugated to an amino acid or an amino acid derivative. In some embodiments, the composition comprises a carrier protein containing a polymer linker conjugated to an amino acid or an amino acid derivative. In some embodiments, the polymer linker comprises polyethylene glycol (PEG). In some embodiments, the antibody is a monoclonal antibody.

[0022] In some embodiments, the antibody fragment comprises an engineered scFv. In some embodiments, the engineered scFv is derived from IgG. In some embodiments, the engineered scFv is engineered using directed evolution. In some embodiments, the engineered scFv undergoes affinity maturation using yeast surface display, phage display, ribosome display, or continuous evolution. In some embodiments, the engineered scFv exhibits improved specificity for a type of protein amino acid or a derivative thereof by using yeast surface display.

[0023] In some embodiments, the derivative or derivative thereof comprises a chemically modified amino acid. In some embodiments, the chemically modified amino acid comprises a hydantoin, phenylthiocarbamoyl, or anilinethiazolinone derivative of an amino acid. In some embodiments, the chemically modified amino acid comprises an amino acid coupled to a linker. In some embodiments, the linker is coupled to the amino acid at the N-terminus. In some embodiments, the linker coupled to the chemically modified amino acid comprises a hydantoin, phenylthiocarbamoyl, or anilinethiazolinone moiety. In some embodiments, the linker is coupled to the amino acid at the C-terminus. In some embodiments, the linker comprises a linker to a nucleic acid molecule. In some embodiments, the linker comprises an amino acid reactive group and a linker to a nucleic acid molecule. In some embodiments, an indirect enzyme-linked immunosorbent assay (ELISA) is used to measure the specific binding of an antibody fragment to one type of protein amino acid or a derivative thereof, rather than to all other types of protein amino acids or derivatives thereof in the group. In some embodiments, surface plasmon resonance is used to measure the specific binding of an antibody fragment to one type of protein amino acid or a derivative thereof, rather than to all other types of protein amino acids or derivatives thereof in the group. In some embodiments, biomembrane interferometry is used to measure the specific binding of an antibody fragment to one type of protein amino acid or a derivative thereof, rather than to all other types of protein amino acids or derivatives thereof in the group. In some embodiments, the antibody fragment includes a detectable marker. In some embodiments, the detectable marker comprises a nucleic acid molecule, a fluorophore, a mass tag, or a protein. In some embodiments, the detectable marker comprises a protein containing an additional antibody or an additional antibody fragment. In some embodiments, the detectable marker comprises a nucleic acid molecule containing a barcode sequence. In some embodiments, the detectable marker is conjugated to an antibody or antibody fragment using a portion selected from the group consisting of: biotin, dethiobiotin, avidin, streptavidin, neutral avidin, SpyCatcher, SpyTag, SNAP tag, click chemistry portion, and cysteine ​​tag. In some embodiments, the antibody fragment comprises atypical amino acids, wherein the detectable marker is conjugated to the antibody or antibody fragment via atypical amino acids. In some embodiments, the atypical amino acids include a click chemistry portion.

[0024] In some embodiments, the antibody fragment specifically binds to phenylalanine or a derivative thereof, but not to all other types of protein amino acids or their derivatives in the group. In some embodiments, the antibody fragment specifically binds to leucine or a derivative thereof, but not to all other types of protein amino acids or their derivatives in the group. In some embodiments, the antibody fragment specifically binds to valine or a derivative thereof, but not to all other types of protein amino acids or their derivatives in the group. In some embodiments, the antibody fragment specifically binds to tyrosine or a derivative thereof, but not to all other types of protein amino acids or their derivatives in the group. In some embodiments, the antibody fragment specifically binds to proline or a derivative thereof, but not to all other types of protein amino acids or their derivatives in the group.

[0025] On the other hand, this paper provides a binder with an equilibrium dissociation constant (K0) of less than 1 nanomolar concentration (1 nm). d () binds to amino acids or their derivatives, as determined by biofilm layer interferometry (BLI) or surface plasmon resonance (SPR).

[0026] In some implementations, the binder comprises an antibody or an antibody fragment.

[0027] In some implementations, the binder specifically binds to a type of amino acid or a derivative thereof.

[0028] In some embodiments, the dissociation constant of the binder for amino acids or their derivatives is smaller than that of the binder for other types of amino acids or their derivatives.

[0029] In some embodiments, the binding agent is part of a binding agent library, wherein the binding agent library contains multiple binding agents that collectively recognize and specifically bind to at least five different types of protein amino acids. In some embodiments, the binding agent is part of a binding agent library, wherein the binding agent contains multiple binding agents that collectively recognize and specifically bind to all 20 types of standard protein amino acids. In some embodiments, the binding agent library contains multiple polyclonal antibodies. In some embodiments, the binding agent library contains multiple monoclonal antibodies.

[0030] In some implementations, the binder comprises immunoglobulin G (IgG).

[0031] In some embodiments, the binder is generated by immunizing animals with an immunogenically effective amount of a composition comprising a carrier protein coupled to an amino acid or an amino acid derivative. In some embodiments, the composition comprises a carrier protein containing a polymeric linker coupled to an amino acid or an amino acid derivative. In some embodiments, the polymeric linker comprises polyethylene glycol (PEG).

[0032] In some implementations, the binding agent is a monoclonal antibody.

[0033] In some embodiments, the binder comprises engineered scFv. In some embodiments, the engineered scFv is derived from IgG. In some embodiments, the engineered scFv is engineered using directed evolution. In some embodiments, the engineered scFv undergoes affinity maturation using yeast surface display, phage display, ribosome display, or sequential evolution. In some embodiments, the engineered scFv exhibits improved specificity for a type of protein amino acid or its derivatives through the use of yeast surface display.

[0034] In some embodiments, the derivative or a derivative thereof comprises a chemically modified amino acid. In some embodiments, the chemically modified amino acid comprises a hydantoin, phenylthiocarbamoyl, or anilinethiazolinone derivative of an amino acid. In some embodiments, the chemically modified amino acid comprises an amino acid coupled to a linker. In some embodiments, the linker is coupled to the amino acid at the N-terminus. In some embodiments, the linker coupled to the chemically modified amino acid comprises a hydantoin, phenylthiocarbamoyl, or anilinethiazolinone moiety. In some embodiments, the linker is coupled to the amino acid at the C-terminus. In some embodiments, the linker comprises a linker to a nucleic acid molecule. In some embodiments, the linker comprises an amino acid reactive group and a linker to a nucleic acid molecule.

[0035] In some embodiments, indirect enzyme-linked immunosorbent assay (ELISA) is used to measure the binding of the binder to the amino acid or its derivative. In some embodiments, surface plasmon resonance is used to measure the binding of the binder to the amino acid or its derivative. In some embodiments, biofilm interferometry is used to measure the binding of the binder to the amino acid or its derivative.

[0036] In some embodiments, the binder comprises a detectable label. In some embodiments, the detectable label comprises a nucleic acid molecule, a fluorophore, a mass tag, or a protein. In some embodiments, the detectable label comprises a protein, wherein the protein contains an additional antibody or an additional antibody fragment. In some embodiments, the detectable label comprises a nucleic acid molecule, wherein the nucleic acid molecule contains a barcode sequence. In some embodiments, the detectable label is conjugated to the binder using a portion selected from the group consisting of: biotin, dethiobiotin, avidin, streptavidin, neutral avidin, SpyCatcher, SpyTag, SNAP tag, click chemistry portion, and cysteine ​​tag. In some embodiments, the binder comprises atypical amino acids, wherein the detectable label is conjugated to the binder via atypical amino acids. In some embodiments, the atypical amino acids include click chemistry portions.

[0037] In some embodiments, the binder specifically binds to phenylalanine or a derivative thereof, but not to all other types of standard protein amino acids or their derivatives. In some embodiments, the binder specifically binds to leucine or a derivative thereof, but not to all other types of standard protein amino acids or their derivatives. In some embodiments, the binder specifically binds to valine or a derivative thereof, but not to all other types of standard protein amino acids or their derivatives. In some embodiments, the binder specifically binds to tyrosine or a derivative thereof, but not to all other types of standard protein amino acids or their derivatives. In some embodiments, the binder specifically binds to proline or a derivative thereof, but not to all other types of standard protein amino acids or their derivatives.

[0038] On the other hand, this article provides an antibody fragment that specifically binds to free amino acids or their derivatives, wherein the free amino acids are not directly coupled to the peptide.

[0039] In some embodiments, the free amino acid is a cleaved amino acid. In some embodiments, the cleaved amino acid is cleaved using a chemical stimulus. In some embodiments, the chemical stimulus includes the use of isothiocyanates. In some embodiments, the chemical stimulus includes the use of acids.

[0040] In some embodiments, the derivative or a derivative thereof comprises a chemically modified amino acid. In some embodiments, the chemically modified amino acid comprises a hydantoin, phenylthiocarbamoyl, or anilinethiazolinone derivative of an amino acid. In some embodiments, the chemically modified amino acid comprises an amino acid coupled to a linker. In some embodiments, the linker is coupled to the amino acid at the N-terminus. In some embodiments, the linker coupled to the chemically modified amino acid comprises a hydantoin, phenylthiocarbamoyl, or anilinethiazolinone moiety. In some embodiments, the linker is coupled to the amino acid at the C-terminus. In some embodiments, the linker comprises a linker to a nucleic acid molecule. In some embodiments, the linker comprises an amino acid reactive group and a linker to a nucleic acid molecule.

[0041] In some embodiments, the antibody fragment is part of an antibody or antibody fragment library, wherein the antibody or antibody fragment library contains multiple antibodies or antibody fragments that collectively recognize and specifically bind to at least five different types of protein amino acids. In some embodiments, the antibody fragment is part of an antibody or antibody fragment library, wherein the antibody or antibody fragment library contains multiple antibodies or antibody fragments that collectively recognize and specifically bind to all 20 types of standard protein amino acids. In some embodiments, the binding agent of the antibody or antibody fragment library comprises multiple polyclonal antibodies. In some embodiments, the antibody or antibody fragment library comprises multiple monoclonal antibodies.

[0042] In some embodiments, the antibody fragment is derived from an antibody. In some embodiments, the antibody is immunoglobulin G (IgG). In some embodiments, the antibody is generated by immunizing an animal with an immunogenically effective amount of a composition comprising a carrier protein conjugated to an amino acid or an amino acid derivative. In some embodiments, the composition comprises a carrier protein containing a polymer linker conjugated to an amino acid or an amino acid derivative. In some embodiments, the polymer linker comprises polyethylene glycol (PEG). In some embodiments, the antibody is a monoclonal antibody.

[0043] In some embodiments, the antibody fragment comprises an engineered scFv. In some embodiments, the engineered scFv is derived from IgG. In some embodiments, the engineered scFv is engineered using directed evolution. In some embodiments, the engineered scFv undergoes affinity maturation using yeast surface display, phage display, ribosome display, or continuous evolution. In some embodiments, the engineered scFv exhibits improved specificity for free amino acids or their derivatives through the use of yeast surface display.

[0044] In some embodiments, the derivative or a derivative thereof comprises a chemically modified amino acid. In some embodiments, the chemically modified amino acid comprises a hydantoin, phenylthiocarbamoyl, or anilinethiazolinone derivative of an amino acid. In some embodiments, the chemically modified amino acid comprises an amino acid coupled to a linker. In some embodiments, the linker is coupled to the amino acid at the N-terminus. In some embodiments, the linker coupled to the chemically modified amino acid comprises a hydantoin, phenylthiocarbamoyl, or anilinethiazolinone moiety. In some embodiments, the linker is coupled to the amino acid at the C-terminus. In some embodiments, the linker comprises a linker to a nucleic acid molecule. In some embodiments, the linker comprises an amino acid reactive group and a linker to a nucleic acid molecule.

[0045] In some embodiments, indirect enzyme-linked immunosorbent assay (ELISA) is used to measure the specific binding of antibody fragments to free amino acids or their derivatives. In some embodiments, surface plasmon resonance is used to measure the specific binding of antibody fragments to free amino acids or their derivatives. In some embodiments, biofilm interferometry is used to measure the specific binding of antibody fragments to free amino acids or their derivatives.

[0046] In some embodiments, the antibody fragment includes a detectable marker. In some embodiments, the detectable marker comprises a nucleic acid molecule, a fluorophore, a mass tag, or a protein. In some embodiments, the detectable marker comprises a protein, wherein the protein contains an additional antibody or an additional antibody fragment. In some embodiments, the detectable marker comprises a nucleic acid molecule, wherein the nucleic acid molecule contains a barcode sequence. In some embodiments, the detectable marker is conjugated to the antibody or antibody fragment using a portion selected from the group consisting of: biotin, dethiobiotin, avidin, streptavidin, neutral avidin, SpyCatcher, SpyTag, SNAP tag, click chemistry portion, and cysteine ​​tag. In some embodiments, the antibody fragment includes atypical amino acids, wherein the detectable marker is conjugated to the antibody or antibody fragment via atypical amino acids. In some embodiments, the atypical amino acids include a click chemistry portion.

[0047] In some embodiments, the antibody fragment specifically binds to phenylalanine or a derivative thereof, but not to all other types of standard protein amino acids or their derivatives. In some embodiments, the antibody fragment specifically binds to leucine or a derivative thereof, but not to all other types of standard protein amino acids or their derivatives. In some embodiments, the antibody fragment specifically binds to valine or a derivative thereof, but not to all other types of standard protein amino acids or their derivatives. In some embodiments, the antibody fragment specifically binds to tyrosine or a derivative thereof, but not to all other types of standard protein amino acids or their derivatives. In some embodiments, the antibody fragment specifically binds to proline or a derivative thereof, but not to all other types of standard protein amino acids or their derivatives.

[0048] On the other hand, this article discloses an antibody or antibody fragment that binds to isoleucine and valine amino acids or derivatives thereof, wherein the antibody or antibody fragment binds to isoleucine with a greater affinity than when binding to valine.

[0049] In some implementations, ELISA is used to measure affinity.

[0050] In another aspect, this article provides an antibody or antibody fragment that binds to phenylalanine, tyrosine, and tryptophan amino acids or their derivatives, wherein the antibody or antibody fragment binds to phenylalanine with a greater affinity than it binds to tyrosine or tryptophan.

[0051] In some implementations, ELISA is used to measure affinity.

[0052] On the other hand, this paper discloses a method for identifying amino acids, the method comprising: (a) providing (i) a plurality of peptides containing a plurality of amino acids, wherein the plurality of amino acids comprises N different types of amino acids, and (ii) M different types of antibodies or antibody fragments, wherein the antibodies or antibody fragments of the M different types of antibodies or antibody fragments can bind to more than one type of amino acid, wherein M and N are integers greater than 1; (b) recording the interactions between the antibodies or antibody fragments of the M different types of antibodies or antibody fragments and the amino acids or derivatives thereof in the plurality of amino acids; (c) repeating (a) and (b) at least once to generate binding patterns of the M different types of antibodies or antibody fragments; and (d) using the binding patterns to determine the identity of each amino acid in the N different types of amino acids.

[0053] In some implementations, (d) includes determining which of the M different types of antibodies or antibody fragments bind to amino acids to produce a binding pattern.

[0054] In another aspect, this article provides a method for identifying amino acids, the method comprising: (a) providing (i) a peptide comprising multiple amino acids, and (ii) an antibody or antibody fragment capable of recognizing at least two amino acids; (b) recording the interaction between the antibody or antibody fragment and the amino acids of the peptide or derivatives thereof; (c) repeating (a) and (b) N times to generate a binding pattern between the antibody or antibody fragment and the amino acid; and (d) using the binding pattern to determine the identity of the amino acid or derivative thereof.

[0055] In some implementations, during (a), multiple peptides are provided, and wherein the multiple peptides are subjected to (b), (c) and (d).

[0056] In some implementations, the amino acid is an N-terminal amino acid.

[0057] In some implementations, the antibody or antibody fragment is part of an antibody or antibody fragment library.

[0058] In some implementations, (d) includes determining the ratio of the number of recorded binding events between the antibody or antibody fragment and an amino acid to N.

[0059] On the other hand, this document discloses an antibody, wherein the antibody is generated by a method comprising: (a) immunizing an animal with an immunogenically effective amount of a composition comprising a carrier protein coupled to a polymeric linker molecule and an amino acid or a derivative thereof; and (b) obtaining the antibody.

[0060] In some implementations, the polymer connector includes polyethylene glycol (PEG).

[0061] In some implementations, the carrier protein includes bovine serum albumin (BSA).

[0062] In some implementations, the polymer linker is coupled to a carrier protein and an amino acid or a derivative thereof.

[0063] In some embodiments, the derivative comprises a chemically modified amino acid. In some embodiments, the chemically modified amino acid comprises a hydantoin, phenylthiocarbamoyl, or anilinethiazolinone derivative of the amino acid. In some embodiments, the chemically modified amino acid comprises an amino acid coupled to a linker. In some embodiments, the linker is coupled to the amino acid at the N-terminus. In some embodiments, the linker coupled to the chemically modified amino acid comprises a hydantoin, phenylthiocarbamoyl, or anilinethiazolinone moiety. In some embodiments, the linker is coupled to the amino acid at the C-terminus. In some embodiments, the linker comprises a linker to a nucleic acid molecule. In some embodiments, the linker comprises an amino acid reactive group and a linker to a nucleic acid molecule.

[0064] In some embodiments, the method further includes obtaining a polyclonal mixture of antibodies from an animal, wherein the polyclonal mixture contains antibodies. In some embodiments, the antibodies are monoclonal antibodies, wherein the method further includes generating monoclonal antibodies from the polyclonal mixture of antibodies. In some embodiments, this generation includes using a hybridoma.

[0065] In some embodiments, the method further includes generating antibodies using phage panning. In some embodiments, the method further includes obtaining a first antibody from an animal, derivatizing the first antibody into one or more single-chain fragments (scFvs), phage panning the one or more scFvs, and generating antibodies from the one or more scFvs.

[0066] On the other hand, this article provides an antibody or antibody fragment comprising a heavy chain (V) containing three complementarity-determining regions (CDRs). H ) and light chains containing three CDRs (V L ), where: V H CDR #1 has an amino acid sequence selected from the group consisting of: GYAFTSYK (SEQ ID NO: 1); GFNINDYH (SEQ ID NO: 2); GYTFTDYY (SEQ ID NO: 3); DFNIQDYY (SEQ ID NO: 4); GFNIKD (SEQ ID NO: 5); GFTFSSYA (SEQ ID NO: 6); GYSFTDYT (SEQ ID NO: 7); GYTFTIYW (SEQ ID NO: 8); GYTSTIYW (SEQ ID NO: 9); GYYITSGYY (SEQ ID NO: 10); and GFTFSSFG (SEQ ID NO: 11); V H CDR #2 has an amino acid sequence selected from the group consisting of: IDPYNGGT (SEQ ID NO: 12); IDPENGVT (SEQ ID NO: 13); INPNGGT (SEQ ID NO: 14); IDPENGDT (SEQ ID NO: 15); ISSGGSYT (SEQ ID NO: 16); INPYNGFT (SEQ ID NO: 17); IIPTTGYT (SEQ ID NO: 18); INPTGYS (SEQ ID NO: 19); ISHDGNN (SEQ ID NO: 20); and ISSGSTNI (SEQ ID NO: 21); V HCDR #3 has an amino acid sequence selected from the group consisting of: ARSGYSFNY (SEQ ID NO: 22); NGINYEGHFDY (SEQ ID NO: 23); ANWRRSGWFAY (SEQ ID NO: 24); NAVRYANYPRAMDY (SEQ ID NO: 25); NAMRYGNYPRPMDY (SEQ ID NO: 26); ARPGDGYYKYFDV (SEQ ID NO: 27); AIWLRRENFDY (SEQ ID NO: 28); ALWLRRENFDY (SEQ ID NO: 29); EMPDYSGYHDY (SEQ ID NO: 30); SMPDYSAFHDY (SEQ ID NO: 31); ESFIRTTYKVPVDY (SEQ ID NO: 32); and ARFDYDRGAFAY (SEQ ID NO: 33); V L CDR #1 has an amino acid sequence selected from the group consisting of: QGISGN (SEQ ID NO: 34); QSLLNSRIRKNH (SEQ ID NO: 35); QDINSY (SEQ ID NO: 36); QSLLNSRTRKNH (SEQ ID NO: 37); QSLLNSRTRKNS (SEQ ID NO: 38); SVSY (SEQ ID NO: 39); QTIVHSNGNTY (SEQ ID NO: 40); QSLLNSRTRKNN (SEQ ID NO: 41); QSLFNSRTRKNY (SEQ ID NO: 42); ETVDYYGTSL (SEQ ID NO: 43); and QSLLKNSNQKNY (SEQ ID NO: 44); V L CDR #2 has an amino acid sequence selected from the following groups: HGT; WAS; RAN; ATS; KVS; AAS; or V LCDR #3 has an amino acid sequence selected from the group consisting of: VQYAQFPYT (SEQ ID NO: 51); KQSYNLRT (SEQ ID NO: 52); LQYDEFPPT (SEQ ID NO: 53); KQSYNLPT (SEQ ID NO: 54); QQWRSNPPT (SEQ ID NO: 55); FQGSHVPYT (SEQ ID NO: 56); KQSFNLLT (SEQ ID NO: 57); KQAYNLQT (SEQ ID NO: 58); QQSRKIPWT (SEQ ID NO: 59); and QQFYNYLT (SEQ ID NO: 60).

[0067] On the other hand, this article provides an antibody or antibody fragment comprising a heavy chain sequence selected from Table 7.

[0068] On the other hand, this article provides an antibody or antibody fragment comprising a light chain sequence selected from Table 8.

[0069] On the other hand, this document provides a nucleic acid composition comprising a first nucleic acid encoding a variable heavy chain region comprising amino acid residues of Table 7 and a second nucleic acid encoding a variable light chain region comprising amino acid residues of Table 8.

[0070] On the other hand, this article provides an antibody or antibody fragment comprising a variable heavy chain sequence selected from Table 9.

[0071] On the other hand, this article provides an antibody or antibody fragment comprising a variable light chain sequence selected from Table 10.

[0072] On the other hand, this document provides a nucleic acid composition comprising a first nucleic acid encoding a heavy chain region comprising amino acid residues of Table 7 or Table 9 and a second nucleic acid encoding a light chain region comprising amino acid residues of Table 8 or Table 10.

[0073] On the other hand, this article provides a kit containing multiple antibodies or antibody fragments, wherein the multiple antibodies or antibody fragments bind specifically or semi-specifically to at least 5 different types of protein amino acids or their derivatives.

[0074] On the other hand, this article provides a method for processing a polymer analyte comprising multiple monomers, the method comprising: (a) providing a polymer analyte, a capture fraction, and a polymerizable molecule; (b) coupling one of the monomers to the capture fraction to produce a monomer-capture fraction complex; (c) contacting the monomer-capture fraction complex with a binder that is coupled to a barcode molecule; and (d) using a nuclease to couple the barcode molecule to the polymerizable molecule or the capture fraction.

[0075] In some implementations, the nuclease is a CRISPR-associated (Cas) enzyme or a guide editor enzyme.

[0076] In some embodiments, the monomer-capture partial complex of (c) comprises a cleaved monomer-capture partial complex, and further includes cleaving the monomer-capture partial complex of (b) prior to (c) to produce a cleaved monomer-capture partial complex.

[0077] In some embodiments, the method further includes, after (a)-(d), (e) decoupling the monomer from the captured portion. In some embodiments, the method further includes repeating (a)-(e) to sequence the polymer analyte.

[0078] In some implementations, the method further includes repeating (a)-(d).

[0079] In some implementations, polymerizable molecules are coupled with polymeric analytes.

[0080] In some implementations, polymerizable molecules associate with polymeric analytes.

[0081] In some implementations, polymerizable molecules include nucleic acid barcode molecules containing barcode sequences.

[0082] In some implementations, the polymer analyte includes peptides.

[0083] In some implementations, the capture fraction and the polymerizable molecule each comprise nucleic acid molecules.

[0084] In some embodiments, the polymerizable molecule comprises a tandem array of CRISPR-Cas9 target sites. In some embodiments, the tandem array of CRISPR-Cas9 target sites comprises multiple target sites, including a first target site which is active, and multiple target sites other than the first target site which are inactive. In some embodiments, the multiple target sites other than the first target site are truncated and therefore inactive. In some embodiments, the multiple target sites other than the first target site are truncated at the 5' end of each of the multiple target sites. In some embodiments, the barcode molecule comprises a pegRNA molecule containing a barcode sequence. In some embodiments, the pegRNA also comprises a key sequence configured to activate a second target site among the multiple target sites, wherein the second target site is located near the first target site. In some embodiments, (d) includes conjugating the pegRNA to the first target site using a Cas or guide editor enzyme, thereby inserting the barcode sequence and the key sequence near the first target site. In some embodiments, the insertion causes inactivation of the first target site and activation of the second target site.

[0085] In some embodiments, the polymerizable molecule comprises a prototype spacer adjacent motif (PAM) and a first CRISPR-Cas9 target site. In some embodiments, the barcode molecule comprises a pegRNA molecule containing a barcode sequence and a propagation sequence containing a second CRISPR-Cas9 target site. In some embodiments, (d) includes using Cas or a guide editor enzyme to couple the pegRNA between the first CRISPR-Cas9 target site and the PAM, thereby inserting the barcode sequence and the propagation sequence near the first CRISPR-Cas9 target site. In some embodiments, the insertion causes inactivation of the first CRISPR-Cas9 target site and activation of the second CRISPR-Cas9 target site.

[0086] In some embodiments, the polymer analyte includes a peptide, and therein multiple monomers comprising multiple amino acids of the peptide.

[0087] In some implementations, the binder comprises an antibody or an antibody fragment.

[0088] In some implementations, the captured portion or polymerizable molecule is coupled to the substrate.

[0089] In some implementations, the captured portion and polymerizable molecules are coupled to the substrate.

[0090] On the other hand, this article discloses a composition comprising an antibody or antibody fragment conjugated to a guide RNA (gRNA).

[0091] In some embodiments, the gRNA comprises a barcode sequence. In some embodiments, the barcode sequence identifies an antibody or antibody fragment. In some embodiments, the gRNA also comprises a key sequence configured to insert into a tandem array of CRISPR-Cas9 target sites containing a first target site and a second target site. In some embodiments, the key sequence is configured to insert between the first target site and the second target site. In some embodiments, the gRNA also comprises a propagation sequence containing a CRISPR-Cas9 target site. In some embodiments, the propagation sequence is configured to insert into or adjacent to a polymerizable molecule containing an additional CRISPR-Cas9 target site and a PAM. In some embodiments, the propagation sequence is configured to insert between an additional CRISPR-Cas9 target site and a PAM. In some embodiments, the antibody or antibody fragment specifically binds to monomers of a polymer analyte comprising multiple monomers. In some embodiments, the polymer analyte comprises a peptide, and the antibody or antibody fragment binds to an amino acid or a derivative thereof. In some embodiments, the antibody or antibody fragment specifically binds to one of 20 standard protein amino acids or derivatives thereof. In some embodiments, the derivatives include chemically modified amino acids. In some embodiments, the chemically modified amino acids include phenylthiocarbamoyl, hydantoin, or aniline thiazolinone derivatives of amino acids.

[0092] On the other hand, this article provides a bispecific binding agent comprising a first binding portion and a second binding portion; wherein the first binding portion binds to a first small molecule; wherein the second binding portion binds to a second small molecule; and wherein the first small molecule is different from the second small molecule.

[0093] In some embodiments, the first small molecule comprises an amino acid or a derivative thereof. In some embodiments, the derivative comprises a chemically modified amino acid. In some embodiments, the chemically modified amino acid comprises a phenylthiocarbamoyl group, a hydantoin phenylthiourea group, or an aniline thiazolinone derivative of an amino acid.

[0094] In some embodiments, the second small molecule includes a hapten. In some embodiments, the hapten includes or is conjugated to a nucleic acid molecule. In some embodiments, the nucleic acid molecule includes a nucleic acid barcode molecule. In some embodiments, the nucleic acid barcode molecule identifies a bispecific binder. In some embodiments, the nucleic acid barcode molecule contains time information. In some embodiments, the hapten is selected from the group consisting of digoxigenin, 2,4-DNP, and penicillin.

[0095] In some implementations, the bispecific binder comprises an antibody, antibody fragment, nanobody, or aptamer.

[0096] On the other hand, this document provides a method for processing a polymer analyte, the method comprising: (a) providing a polymer analyte, wherein the polymer analyte comprises a plurality of monomers; (b) coupling one of the monomers to a capture portion; (c) cleaving the monomer from the polymer analyte to produce cleaved monomers; (d) contacting the cleaved monomers with a bispecific binder; wherein the bispecific binder is configured to recognize and bind to the cleaved monomers and small molecules; (e) providing small molecules; and (f) contacting the bispecific binder with the small molecules.

[0097] In some implementations, the polymer analyte includes peptides, and the plurality of monomers are plurality of amino acids.

[0098] In some embodiments, (b) a linker is used for mediation. In some embodiments, the linker comprises a monomeric reactive group and a second reactive group. In some embodiments, the second reactive group comprises a click chemistry portion. In some embodiments, the method further includes contacting the linker with a linker nucleic acid molecule, wherein the linker nucleic acid molecule comprises a third reactive group capable of reacting with the second reactive group.

[0099] In some embodiments, the cleaved monomer comprises an amino acid or a derivative thereof. In some embodiments, the derivative comprises a chemically modified amino acid. In some embodiments, the chemically modified amino acid comprises a phenylthiocarbamoyl group, a hydantoin phenylthiourea group, or an aniline thiazolinone derivative of an amino acid.

[0100] In some embodiments, the second small molecule includes a hapten. In some embodiments, the hapten includes a nucleic acid molecule. In some embodiments, the nucleic acid molecule includes a nucleic acid barcode molecule. In some embodiments, the nucleic acid barcode molecule identifies a bispecific binder. In some embodiments, the nucleic acid barcode molecule contains time information. In some embodiments, the method further includes conjugating the nucleic acid barcode molecule to another nucleic acid molecule. In some embodiments, the other nucleic acid molecule is located near the cleaved monomer on the substrate. In some embodiments, the hapten is selected from the group consisting of digoxigenin, 2,4-DNP, and penicillin.

[0101] In some implementations, the bispecific binder comprises an antibody, antibody fragment, nanobody, or aptamer.

[0102] In some embodiments, the capture portion includes nucleic acid molecules. In some embodiments, the capture portion is coupled to a substrate. In some embodiments, the substrate contains beads.

[0103] Another aspect of this disclosure provides a non-transitory computer-readable medium including machine-executable code that, when executed by one or more computer processors, implements any of the methods described above or elsewhere herein.

[0104] Another aspect of this disclosure provides a system comprising one or more computer processors and computer memory coupled thereto. The computer memory includes machine-executable code that, when executed by the one or more computer processors, implements any of the methods described above or elsewhere herein.

[0105] Further aspects and advantages of this disclosure will become apparent to those skilled in the art from the following detailed description, wherein only exemplary embodiments of the disclosure are shown and described. As will be appreciated, the disclosure is capable of other and different embodiments, and certain details thereof can be modified in various obvious ways without departing from the disclosure. Therefore, the drawings and detailed description are to be considered illustrative in nature and not restrictive.

[0106] Incorporate by reference

[0107] All publications, patents, and patent applications mentioned in this specification are incorporated herein by reference to the extent that each individual publication, patent, or patent application is specifically and individually indicated to be incorporated by reference. In the event that a publication or patent or patent application incorporated by reference contradicts the disclosure contained in this specification, the specification is intended to supersede and / or take precedence over any such contradictory material. Attached Figure Description

[0108] The novel features of the invention are set forth in the appended claims. A better understanding of the features and advantages of the invention will be obtained by referring to the following detailed description of exemplary embodiments utilizing the principles of the invention, along with the accompanying drawings (also referred to herein as “Figure” and “FIG.”):

[0109] Figure 1A An example workflow for processing polymer analyte molecules (e.g., peptides) described herein is illustrated schematically. Figure 1B This schematically illustrates another example workflow for processing polymer analytes in solution. Figure 1C This schematically illustrates another example workflow for processing and detecting polymer analytes. Figure 1D This schematically illustrates another workflow for processing polymer analytes using nucleases for barcoding. Figure 1EThis schematically illustrates another example workflow for processing and detecting polymer analytes using a nanopore sequencing system. Figure 1F This schematically illustrates another example workflow for processing and detecting polymer analytes using a nanopore sequencing system. Figure 1G This schematically illustrates another example workflow for processing and detecting polymer analytes in solution. Figure 1H An example workflow for treating polymer analytes using a dual-specificity binder is illustrated schematically.

[0110] Figure 2 An example connector for attaching polymerizable molecules to polymeric analytes is illustrated schematically.

[0111] Figure 3 This schematically illustrates a method for identifying amino acids using binding modes of multispecific binders.

[0112] Figure 4 An example of a sequencing method for analyzing or characterizing polymer analytes is illustrated schematically. Figure 4 SEQ ID NO: 113 has been disclosed.

[0113] Figure 5 A computer system that is programmed or otherwise configured to implement the methods provided herein is illustrated schematically.

[0114] Figure 6 Data are shown for various binding agents described herein that bind specifically or semi-specifically to amino acid derivatives.

[0115] Figure 7 ELISA data for the anti-leucine binding agent are shown.

[0116] Figure 8 ELISA data for a variety of binding agents are shown, which selectively bind to a single amino acid type or its derivative, rather than selectively binding to all other amino acid types or their derivatives.

[0117] Figure 9 The following ELISA data are shown for binders that recognize amino acid derivatives but not native amino acids when attached to peptides.

[0118] Figure 10 The study presents ELISA data demonstrating that using multiple methods for bioconjugation does not adversely affect the binding profile of the binder.

[0119] Figure 11A Exemplary biofilm layer interferometry (BLI) data for an anti-phenylalanine binder purified from a mouse monoclonal sample are shown. Figure 11BExemplary BLI data for recombinant anti-phenylalanine binders are shown. Figure 11C Exemplary BLI data for an anti-leucine binder purified from mouse monoclonal antibodies are shown. Figure 11D Exemplary BLI data for recombinant antileucine binders are shown. Figure 11E Exemplary BLI data for recombinant antitryptophan binders are shown. Figure 11F Exemplary BLI data for an antivaline binding agent purified from mouse monoclonal antibodies are shown.

[0120] Figure 12A Exemplary imaging data indicating colocalization of anti-Phe binders with their homologous molecules, Phe derivatives, are shown. Figure 12B Exemplary imaging data indicating colocalization of anti-Phe binders with their homologous molecules are shown. Figure 12C Exemplary imaging data indicating the specificity or co-localization efficiency of anti-Phe binders with their homologs and the relative percentage of the target are shown.

[0121] Figure 13 Example data are shown for identifying multiple amino acid types using a unique binding mode with six different binding agents.

[0122] Figure 14 Example data for a directed evolution approach (yeast demonstration) used to improve binder specificity are shown. Detailed Implementation

[0123] While various embodiments of the invention have been shown and described herein, it will be apparent to those skilled in the art that such embodiments are provided by way of example only. Many variations, modifications, and substitutions will occur to those skilled in the art without departing from the invention. It should be understood that various alternative embodiments of the invention described herein may be employed.

[0124] definition

[0125] Whenever the terms "at least," "greater than," or "greater than or equal to" precede the first value in a series of two or more values, the terms "at least," "greater than," or "greater than or equal to" apply to each value in the series. For example, greater than or equal to 1, 2, or 3 is equivalent to greater than or equal to 1, greater than or equal to 2, or greater than or equal to 3.

[0126] Whenever the terms "not exceeding," "less than," or "less than or equal to" precede the first value in a series of two or more values, the terms "not exceeding," "less than," or "less than or equal to" apply to each value in the series. For example, less than or equal to 3, 2, or 1 is equivalent to less than or equal to 3, less than or equal to 2, or less than or equal to 1.

[0127] References to “one implementation,” “implementation,” “example implementation,” “some implementations,” “certain implementations,” “various implementations,” etc., indicate that an implementation of the disclosed technology as described may include a particular feature, structure, or characteristic, but not every implementation must include that particular feature, structure, or characteristic. Furthermore, repeated use of the phrase “in one implementation” does not necessarily refer to the same implementation, but the phrase may refer to the same implementation.

[0128] A range herein may be expressed as “about” or “approximately” or “substantially” a particular value and / or to “about” or “approximately” or “substantially” another particular value. Other exemplary embodiments, when expressing such a range, include a particular value and / or to another particular value. Furthermore, the term “about” means within an acceptable margin of error for a particular value as determined by one of ordinary skill in the art, which will depend in part on how the value is measured or determined, i.e., the limitations of the measurement system. For example, according to practice in the art, “about” may mean within an acceptable standard deviation. Alternatively, “about” may mean a range of up to ±20%, preferably up to ±10%, more preferably up to ±5%, and even more preferably up to ±1% of a given value. Alternatively, particularly with respect to biological systems or biological processes, the term may mean within an order of magnitude of the value, preferably within twice the value. In cases where a particular value is described in this application and claims, the term “about” is implied unless otherwise stated and in that context means within an acceptable margin of error for the particular value.

[0129] "Comprising," "containing," or "including" means that at least the named compound, element, particle, or method step is present in the composition, article, or method, but does not exclude the presence of other compounds, materials, particles, or method steps, even if such other compounds, materials, particles, or method steps have the same function as those named.

[0130] Throughout this specification, various components with specific values ​​or parameters may be identified; however, these items are provided as exemplary embodiments. In fact, the exemplary embodiments do not limit the aspects and concepts of this disclosure, as many comparable parameters, sizes, ranges, and / or values ​​can be implemented. The terms “first,” “second,” etc., “primary,” “secondary,” etc., do not indicate any order, quantity, or importance, but are used to distinguish one element from another.

[0131] As used herein, the term "protein" generally refers to a molecule containing two or more amino acids linked by peptide bonds. Proteins may also be referred to as "polypeptides," "oligopeptides," or simply "peptides." Proteins can be naturally occurring molecules or synthetic molecules (e.g., artificial proteins, peptides, enzymes). Proteins may include one or more non-natural amino acids, modified amino acids, or non-amino acid linkers. Proteins may contain D-amino acid enantiomers, L-amino acid enantiomers, or both. The amino acids of a protein may be modified in a natural or synthetic manner, such as through post-translational modifications or through chemical modifications. In some cases, different proteins can be distinguished from each other based on the different genes they express in an organism, different primary sequence lengths, or different primary sequence compositions. However, proteins expressed by the same gene can be different protein variants, for example, distinguished based on different lengths, different amino acid sequences, or different post-translational modifications. Different proteins can be distinguished based on one or both of the originating gene and the protein variant state.

[0132] As used herein, the term "peptide" can refer to any short single polypeptide chain. The length of a peptide can be no more than about 100, 95, 90, 85, 80, 75, 70, 65, 60, 55, 50, 45, 40, 35, 30, 25, 20, 15, 10, 5, or less than about 5 amino acids. Peptides can have known or unknown biological functions or activities. Peptides can include natural, synthetic, modified, or degraded proteins or peptides, or combinations thereof. Peptides can include proteinaceous, natural, synthetic, or modified amino acids or amino acid residues, or combinations thereof.

[0133] As used herein, the term "single analyte" can refer to an analyte that is operated on alone or distinguished from other analytes. A single analyte can include a biomolecule or a synthetic molecule. A single analyte can include a small molecule. A single analyte can be a single molecule (e.g., a single biomolecule, such as a single protein, nucleic acid molecule, affinity reagent, lipid, carbohydrate, etc.), a single complex of two or more molecules (e.g., a multimeric protein having two or more separable subunits, a single protein attached to a nucleic acid molecule, or a single protein attached to an affinity reagent), a single particle, etc. Unless the context or expressly indicates otherwise, references to "single analyte" in the context of the compositions, systems, or methods herein do not necessarily exclude the application of a composition, system, or method to multiple single analytes that are operated on or distinguished separately.

[0134] As used herein, a "peptide" refers to two or more amino acids linked together by peptide bonds. The term "peptide" includes proteins known in the art that have C-termini and N-termini and may be derived from synthetic or naturally occurring sources. As used herein, "at least a portion of a polypeptide" refers to two or more amino acids of a polypeptide. A polypeptide may comprise one or more peptides. Optionally, a portion of a polypeptide comprises the complete amino acid sequence of the polypeptide or at least 1, 5, 10, 20, 30, or 50 consecutive or nicked amino acids of the complete amino acid sequence of the polypeptide.

[0135] As used herein, the term "sample" refers to a collection of substances or materials that contain or are suspected of containing one or more analytes of interest (e.g., biomolecules, such as peptides). Samples may be modified for purposes such as storage or stability. Samples may be naturally occurring or synthetic. Samples may be treated to separate or remove unwanted fractions or impurities from the analyte of interest. Samples may be enriched or purified. For example, a sample may be part of a separation process (e.g., chromatography, fractionation, electrophoresis, etc.). Alternatively, the sample may not be treated to separate or remove any unwanted fractions or impurities from the analyte of interest. Samples may be obtained from any suitable source or location, including from organisms, cells, tissues, cell preparations, cell-free compositions, and the environment (e.g., air, water, dirt, soil, agriculture, dust, sewage). Samples may be obtained from an organism or a portion of an organism, such as from fluids, tissues, or cells. Samples may include biological and / or non-biological components. As used herein, the terms "biological sample" or "biological source" refer to a sample derived primarily of a biological system or organism, such as one or more viral particles, cells (e.g., individualized cells), organelles (e.g., individualized organelles), tissues, body fluids, bone, cartilage, and exoskeleton. Biological samples may contain prokaryotic cells (e.g., bacteria) or eukaryotic cells (e.g., fungi, protists, algae, plants, animals). Biological samples may contain a majority of biological material on a mass basis (excluding the weight of fluids within the sample). Biological samples may contain one or more proteins, referred to herein as protein samples. Biological samples can be obtained from a variety of sources, such as clinical patient samples (e.g., blood, serum, plasma, cerebrospinal fluid (CSF), saliva, mucosal secretions, sputum, urine, lymph, sweat, vaginal fluid, semen, feces, amniotic fluid, synovial fluid, fine-needle aspirate, tissue biopsy samples, tumor biopsy samples, etc.), bispecific antibodies, monospecific antibodies, single-domain antibodies (sdAbs), and bifunctional antibodies (i.e., bispecific antibodies, such as bispecific T-cell connectors). Biological samples can be processed to purify and retain one or more biomolecules (e.g., proteins, nucleic acids, carbohydrates, lipids, glycoproteins, lipoproteins, metabolites, etc.). Biological samples (e.g., protein samples) can be derived from cultured cells, which can be treated or untreated. Biological samples (e.g., protein samples) can also be generated from tissue samples (such as biopsy samples), which can optionally be processed to release the biomolecules (e.g., proteins) contained therein. Tissue samples can also be derived from in vivo samples, including fresh, frozen, acute, and fixed tissues.

[0136] As used herein, the terms “antibody” and “immunoglobulin” generally refer to proteins that recognize and bind to specific antigens. Antibodies or immunoglobulins can refer to antibody isotypes, full-length antibodies (e.g., containing two identical light chains and two identical heavy chains), antibody fragments (including but not limited to Fab, Fv, scFv, and Fd fragments), chimeric antibodies, humanized antibodies, single-chain antibodies, linear antibodies, and fusion proteins (including the antigen-binding portion of an antibody and non-antibody proteins). Antibodies can be detectably labeled, for example, with fluorophores, radioisotopes, enzymes that produce detectable products (e.g., peroxidases), fluorescent proteins, nucleic acid barcode sequences, etc. Antibodies can be further conjugated to other parts, such as members of specific binding pairs, such as biotin (a member of the biotin-avidin specific binding pair) or analogues (e.g., dethiobiotin, neutral avidin, streptoavidin, etc.), SpyCatcher and SpyTag, SNAP-tag. ® Antibodies can comprise fusion proteins (e.g., antibodies and detectable proteins such as fluorescent proteins), such as green fluorescent protein (GFP), red fluorescent protein (RFP), yellow fluorescent protein (YFP), mVenus, tdTomato, etc. These terms also encompass nanobodies, Fab', Fv, F(ab')2, scFv, bispecific scFv, biantibodies, triantibodies, tetraantibodies, microantibodies, and other antibody fragments that retain specific binding to antigens. Antibodies can exist in a variety of other forms, including, for example, Fv, Fab and (Fab)2, biantibodies, monoantibodies, single-domain antibodies (sdAbs), and bifunctional antibodies (i.e., bispecific antibodies, such as bispecific T-cell connectors) hybrid antibodies (e.g., Lanzavecchia et al., Eur. J. Immunol. 17, 105 (1987)) and in single-chain form (e.g., Huston et al., Proc. Natl. Acad. Sci. USA, 85, 5879-5883 (1988) and Bird et al., Science, 242, 423-426 (1988), which are incorporated herein by reference). (See also Hood et al., Immunology, Benjamin, NY, 2nd edition (1984), and Hunkapiller and Hood, Nature, 323, 15-16 (1986), which are incorporated herein by reference). Naturally occurring immunoglobulin or antibody types include immunoglobulin A, immunoglobulin G, immunoglobulin D, immunoglobulin E, immunoglobulin M, or other immunoreactive components. Antibodies can contain complementarity-determining regions (CDRs) in both the light and heavy chains, while the more conserved portions of the variable domains are called frames (FRs).

[0137] The term "antibody fragment" can generally refer to fragments of an antibody, including but not limited to Fab, Fv, scFv and Fd fragments, nanobodies, Fab', F(ab')2, bispecific scFv, biantibodies, triantibodies, tetraantibodies, microantibodies, chimeric antibodies, humanized antibodies, single-chain antibodies, and fusion proteins that include the antigen-binding portion of the antibody. Fab fragments typically refer to fragments containing a constant domain of the light chain, a first constant domain (CH1) of the heavy chain, and variable domains (V) of both the light and heavy chains. L and V H The Fab' fragment is an antibody fragment. Compared to the Fab fragment, the Fab' fragment typically contains additional residues at the carboxyl terminus of the heavy chain CH1 domain, including one or more cysteine ​​residues from the antibody hinge region. The Fab' fragment is generated by cleaving the disulfide bonds at the hinge cysteine ​​residues of the F(ab')2 pepsin digestion product. Both the Fab and F(ab')2 fragments lack the fragment crystallizable (Fc) region of the intact antibody. Single-chain antibodies (or single-domain antibodies) consist of a single V H or V L Composition of structural domains.

[0138] As used herein, the "Fc" region typically refers to a crystallizable constant region of an antibody fragment. Fc regions generally do not contain antigen-specific binding regions. In IgG, IgA, and IgD antibody isotypes, the Fc region contains two identical protein fragments derived from the second and third constant domains (CH2 and CH3) of the two heavy chains of the antibody. IgM and IgE Fc regions contain three heavy chain constant domains (CH2, CH3, and CH4) in each polypeptide chain. Variant Fc regions contain amino acid sequences that differ from the amino acid sequence of the native Fc region. "Natural Fc region" or "natural Fc sequence" includes amino acid sequences identical to those found in naturally occurring Fc regions. Non-naturally occurring Fc regions or variant sequences can refer to Fc regions with one or more amino acid mutations at one position, or Fc regions or variant sequences that share a percentage homology or identity with the native Fc region.

[0139] As used in this article, the "Fv" fragment typically refers to the smallest fragment of an antibody containing complete target recognition and binding sites. Fv contains a heavy chain variable region (V0). H ) and a light chain variable region (V L Dimers that associate with each other. Each variable domain (V) H or V L The three complementary determinant regions (CDRs) can interact to form V H -V LThe surface of the dimer defines the target binding site. In some cases, the six CDRs confer target binding specificity; however, in others, a single variable domain (e.g., only V) defines the target binding site. H or V L Having only three CDRs, such as single-domain antibodies (sdAbs), allows them to specifically recognize and bind to a target. A single-chain Fv (“scFv”) contains V... H Domain and V L Domain. The V of scFv H Domain and V L Domains can be contained within a single polypeptide chain, typically in V H Domain and V L There are peptide linkers between the domains. Bispecific scFv (bis-scFv) can contain two V... H and V L Domains, and can be homodimers (identical V H and V L ) or heterodimers (different V H and V L ).

[0140] As used herein, the term "biantibody" generally refers to a substance containing two groups of V antibodies linked interchain. H Domain and V L Antibody fragments with structural domains. Biantibodies are typically divalent (having two antigen-binding sites). Biantibodies can be heterodimers of two scFv fragments, where the V of the two scFv fragments... H Domain and V L Domains exist on different polypeptide chains.

[0141] As used herein, the term "variable region" generally refers to a domain in the antibody heavy or light chain that participates in antigen binding. A variable domain typically contains four conserved framework regions and three complementarity-determining regions. In some cases, a single V... H Domain or V L The domain is sufficient to bind to the antigen; in other cases, it needs to contain V. H Domain and V L The Fv domain is used to recognize and bind to antigens.

[0142] As used herein, the term "complementarity-determining region (CDR)" generally refers to the hypervariable region of an antibody's variable domain, which is sequence-highly variable and can form structurally defined loops or contain antigen contact residues. Antibodies typically contain six CDRs, three of which are located in the V... H In the middle, three are in V L In the middle. The CDRs in each chain can be distributed among the four FR regions.

[0143] As used herein, the term "monoclonal antibody" generally refers to an antibody derived from a single clone. Monoclonal antibodies are not limited to antibodies produced through hybridoma technology. Monoclonal antibodies can be derived from a single clone, including any eukaryotic, prokaryotic, archaea, or viral (e.g., bacteriophage) clone. Monoclonal antibodies can be prepared using any useful techniques or combinations thereof, such as hybridoma, recombination, and directed evolution methods, and techniques such as yeast display, yeast surface display, bacteriophage display, bacterial display, mammalian display, mammalian surface display, ribosome display, or mRNA display.

[0144] The term "epitope" or "antigenic determinant" generally refers to a site on an antigen that an antibody specifically binds to, for example, as defined by the specific method used to identify it. An epitope may comprise a portion of a peptide (e.g., several amino acids), which can be formed from sequential or discontinuous amino acids (which can be spatially juxtaposed through the tertiary folding of the peptide). As described herein, an epitope may comprise a single amino acid or a derivative thereof (e.g., a chemically or biologically modified amino acid), which may not be part of or a fraction of a peptide (e.g., a physically isolated amino acid or a derivative thereof).

[0145] As mentioned in this article, a "chimeric antibody" is an antibody in which part of the heavy chain, light chain, or both are derived from a first species or antibody class of immunoglobulin, and the remaining regions are derived from a second species or antibody class different from the first species or antibody class.

[0146] The antibodies disclosed herein can be derived from animals, including but not limited to mice, rats, rabbits, non-human primates, llamas, chickens, or other animals. Alternatively or otherwise, the antibodies can be engineered into non-animal or chimeric antibodies, such as recombinant antibodies, in vitro derived antibodies, such as antibodies grown in useful cell lines, such as monkey kidney CV1 line (COS-7) transformed with SV40; human embryonic kidney line (293 or 293 cells); young hamster kidney cells (BHK); mouse Support cells (TM4 cells); monkey kidney cells (CV1); African green monkey kidney cells (VERO-76); human cervical cancer cells (HELA); canine kidney cells (MDCK); buffalo rat hepatocytes (BRL 3A); human lung cells (W138); human hepatocytes (Hep G2); mouse mammary tumor cells (MMT060562); TRI cells; MRC 5 cells; and FS4 cells. Other useful mammalian host cell lines include Chinese hamster ovary (CHO) cells (including DHFR-CHO cells) and myeloma cells (such as YO, NS0, and Sp2 / 0). Other examples of mammalian cell lines for antibody production can be found in Yazaki and Wu, Methods in Molecular Biology, Vol. 248 (BKC Lo, ed., Humana Press, Totowa, NJ), pp. 255–268 (2003), which is incorporated herein by reference.

[0147] The antibodies disclosed herein may be polyclonal, monoclonal, genetically engineered, or otherwise modified, including but not limited to chimeric antibodies, humanized antibodies, and phenotypic engineered antibodies.

[0148] The antibodies and nucleic acid molecules described herein may be described by peptide or nucleic acid sequences. Unless otherwise specified, peptide sequences are provided in the N-C orientation and nucleic acid sequences are provided in the 5'-3' orientation.

[0149] As described herein, the “percentage (%) amino acid sequence identity” or “homology” of the peptide or antibody sequence identified herein refers to the percentage of amino acid residues in the candidate sequence that are identical to amino acid residues in the compared peptide or antibody after considering any conserved substitutions as part of the sequence identity alignment. Alignments used to determine the percentage of amino acid sequence identity can be performed in various ways within the scope of the art, for example, using publicly available computer software such as BLAST, BLAST-2, ALIGN, or Megalign (DNASTAR) software. Those skilled in the art can determine appropriate parameters for measuring alignments, including any algorithms required to achieve maximum alignment across the full length of the compared sequences. For the purposes of this document, the amino acid sequence identity percentage values ​​are generated using the sequence comparison computer program BLAST.

[0150] As used herein, “binding” or “coupling” generally refers to a covalent or non-covalent interaction between two molecules (referred to herein as “binding couplers,” such as a base and an enzyme or an antibody and an epitope). Binding between binding couplers can be specific or non-specific.

[0151] As used herein, "specific binding" or "specifically binding" generally refers to an interaction between binding partners (e.g., a binding partner and a homologous molecule) such that, under a set of conditions, the binding partners bind to each other but not to other molecules that may be present in the environment (e.g., in biological samples, in tissues, in in vitro assays). Specific binding interactions may require binding partners that bind to homologous molecules. Specific binding interactions may require binding partners to their homologous molecules at significantly or substantially higher levels or with greater affinity compared to binding partners to non-homologous molecules. Specific binding interactions may require a first binding partner that has greater selectivity for binding to homologous molecules compared to non-homologous molecules.

[0152] The binding agents described herein (e.g., antibodies, antibody fragments, nucleic acid binding agents, etc.) may include derivatized binding agents. Derivatized binding agents may include chemical or enzymatic modifications. Modifications may be covalent or non-covalent. Non-limiting examples of modifications include, but are not limited to, acylation, acetylation, alkylation (including methylation), biotinylation, butyrylation, carbamoylation, carbonylation, amidation, deamidation, deiminization, diphthylamide formation, disulfide bridge formation, elimination, flavin attachment, formylation, γ-carboxylation, glutamylation, glycylation, glycosylation, glycosylphosphatidylinositolization, heme C attachment, hydroxylation, hydroxybutyryl lysine formation, iodination, isopreneation, esterification, ester acylation, malonylation, methylation, myristylation, oxidation, transglutamatelation, palmitoylation, polyethylene glycolation, phosphopanthryl ethylamineation, phosphorylation, isopentenylation, propionylation, retinyl Schiff base formation, S-glutathioneization, S-nitrosylation, S-sulfinylation, selenization, succinylation, sulfinization, ubiquitination, threonization, disulfide bond formation, and C-terminal amidation. Modifications can occur at the amino terminus and / or carboxyl terminus of the peptide. Modifications of the terminal amino group include, but are not limited to, deamination, N-lower alkyl, N-dilower alkyl, and N-acyl modifications. Modifications of the terminal carboxyl group include, but are not limited to, amides, lower alkylamides, dialkylamides, and lower alkyl esters (e.g., where the lower alkyl group is C1-C4 alkyl). Modifications can include chemical modifications such as protecting or blocking groups. The binder may contain one or more non-natural amino acids.

[0153] The terms “nucleic acid,” “nucleic acid molecule,” “oligonucleotide,” and “polynucleotide” are used interchangeably herein and generally refer to a polymer of any length of naturally occurring or synthetic nucleotides or their analogues. Nucleic acid molecules may contain one or more deoxyribonucleotides, deoxynucleoside triphosphates, dideoxynucleoside triphosphates, ribonucleotides, hexitol nucleotides, cyclohexane nucleotides, or their analogues or combinations. Nucleic acid molecules may include, for example, DNA, RNA, HNA, and their modified forms. Nucleic acid molecules may contain nucleotides linked by phosphodiester bonds. Nucleic acid molecules may have any two-dimensional or three-dimensional structure and may perform any known or unknown function. Nucleic acid molecules may be single-stranded, double-stranded, or partially double-stranded. Non-restrictive examples of polynucleotides include genes, gene fragments, exons, introns, messenger RNA (mRNA), transfer RNA, ribosomal RNA, non-coding RNA, small interfering RNA, short hairpin RNA, microRNA, scaRNA, ribozymes, riboswitches, viral RNA, complementary DNA (cDNA), granular DNA, mitochondrial DNA, chromosomal or genomic DNA, viral DNA, recombinant polynucleotides, branched polynucleotides, plasmids, vectors, isolated DNA of any sequence, control regions, isolated RNA of any sequence, nucleic acid probes, nucleic acid adaptors, and primers. Nucleic acid molecules can be linear, circular, or any other geometric structure. Examples of polynucleotide analogs include, but are not limited to, xenobiotic nucleic acids (XNA), bridged nucleic acids (BNA), glycol nucleic acids (GNA), hexitol nucleic acids (HNA), cyclohexane nucleic acids (CeNA), 2'-F-arabinose nucleic acids (2'-F-ANA), peptide nucleic acids (PNA), γ-PNA, morpholinopolynucleotides, locked nucleic acids (LNA), threonine nucleic acids (TNA), 2'-O-methyl polynucleotides, 2'-O-alkylribosyl-substituted polynucleotides, thiophosphate polynucleotides, and borosilicate polynucleotides. Polynucleotide analogs may have purine or pyrimidine analogs, including, for example, 7-denitropurine analogs, 8-halopurine analogs, 5-halopyrimidine analogs, or universal base analogs that can pair with any base, including hypoxanthine, nitroazole, isoquinolone analogs, azolecarbamate, and aromatic triazole analogs, or base analogs with additional functions, such as a biotin moiety for affinity binding.

[0154] As used herein, the term "amino acid" generally refers to an organic compound that combines to form a protein or peptide. Amino acids typically contain an amine group, a carboxylic acid group, and a side chain specific to each amino acid, which acts as a monomeric subunit of the peptide. Amino acids can include 20 standard, naturally occurring, or typical amino acids, as well as non-standard amino acids. Standard, naturally occurring, or typical amino acids include alanine (A or Ala), cysteine ​​(C or Cys), aspartic acid (D or Asp), glutamic acid (E or Glu), phenylalanine (F or Phe), glycine (G or Gly), histidine (H or His), isoleucine (or Ile), lysine (K or Lys), leucine (L or Leu), methionine (M or Met), asparagine (N or Asn), proline (P or Pro), glutamine (Q or Gln), arginine (R or Arg), serine (S or Ser), threonine (T or Thr), valine (V or Val), tryptophan (W or Trp), and tyrosine (Y or Tyr). Amino acids can be L-amino acids or D-amino acids. Non-standard amino acids can be naturally occurring or chemically synthesized modified amino acids, amino acid analogs, amino acid mimics, non-standard protein amino acids, or non-protein amino acids. Examples of non-standard amino acids include, but are not limited to, selenocysteine, pyrrolidone and N-formylmethionine, (3-amino acids, high-amino acids, proline and pyruvate derivatives, 3-substituted alanine derivatives, glycine derivatives, cyclic-substituted phenylalanine and tyrosine derivatives, linear core amino acids, and N-methyl amino acids.

[0155] As used herein, the term "amino acid type" generally refers to one of the standard, naturally occurring, or typical amino acids, such as a member of the group consisting of: alanine (A or Ala), cysteine ​​(C or Cys), aspartic acid (D or Asp), glutamic acid (E or Glu), phenylalanine (F or Phe), glycine (G or Gly), histidine (H or His), isoleucine (or Ile), lysine (K or Lys), leucine (L or Leu), methionine (M or Met), asparagine (N or Asn), proline (P or Pro), glutamine (Q or Gln), arginine (R or Arg), serine (S or Ser), threonine (T or Thr), valine (V or Val), tryptophan (W or Trp), tyrosine (Y or Tyr), their derivatives, and any of the aforementioned amino acids in modified forms. The term “amino acid type” may be used in this document to distinguish multiple amino acids containing different side chain groups, rather than the same multiple amino acids (e.g., amino acids at different positions in a single peptide with the same side chain).

[0156] As used herein, the term "post-translational modification" refers to modifications that occur on peptides or amino acids after translation. Post-translational modifications can be covalent or enzymatic. Examples of post-translational modifications include, but are not limited to, acylation, acetylation, alkylation (including methylation), biotinylation, butyrylation, carbamoylation, carbonylation, deamidation, deiminization, diphthylamide formation, disulfide bridge formation, elimination, flavin attachment, formylation, γ-carboxylation, glutamylation, glycylation, glycosylation, glycosylphosphatidylinositolization, heme C attachment, hydroxylation, hydroxybutyryl lysine formation, iodination, isopreneation, esterification, esterification, malonylation, methylation, myristylation, oxidation, transglutamate, palmitoylation, polyethylene glycolation, phosphopanthryl ethylamineation, phosphorylation, isopentenylation, propionylation, retinyl Schiff base formation, S-glutathioneization, S-nitrosylation, S-sulfinylation, S-adenosylation, selenization, succinylation, sulfinization, ubiquitination, threonization, disulfide bond formation, and C-terminal amidation. Post-translational modifications include modifications to the amino-terminal and / or carboxyl-terminal of a peptide. Modifications of the terminal amino group include, but are not limited to, deamination, N-lower alkyl, N-dilower alkyl, and N-acyl modifications. Modifications of the terminal carboxyl group include, but are not limited to, amide, lower alkylamide, dialkylamide, and lower alkyl ester modifications (e.g., where the lower alkyl group is C1-C4 alkyl). Post-translational modifications also include modifications of amino acids falling between the amino and carboxyl terms, such as, but not limited to, those described above. The term post-translational modification may also include peptide modifications containing one or more detectable tags. Post-translational modifications can be naturally occurring or synthetic.

[0157] As used herein, the term "binding agent" refers to a molecule that binds to, associates with, conjugates with, recognizes, or combines with another molecule, such as nucleic acid molecules, peptides, polypeptides, proteins, carbohydrates, synthetic molecules, or small molecules. Binding agents can bind to macromolecules or components or features of macromolecules. Binding agents can form covalent or non-covalent associations with molecules, macromolecules, or components or features of macromolecules. Binding agents can also be chimeric binding agents composed of two or more types of molecules, such as nucleic acid-peptide chimeric binding agents, carbohydrate-peptide chimeric binding agents, or lipid-peptide chimeric binding agents. Binding agents can be naturally occurring, synthetically produced, or recombinantly expressed molecules. Binding agents can bind to a single monomer or subunit of a polymer analyte, such as a macromolecule (e.g., a single amino acid of a peptide) or to multiple linked subunits of a macromolecule (e.g., dipeptides, tripeptides, or higher-order peptides of longer peptides, polypeptides, or protein molecules). Binding agents can bind to linear molecules or molecules having a three-dimensional structure (also known as conformation). For example, antibody binders can bind to linear peptides, polypeptides, or proteins, or to conformational peptides, polypeptides, or proteins. The binder can bind to the N-terminal, C-terminal, or intermediate peptides of a peptide, polypeptide, or protein molecule. The binder can bind to the N-terminal, C-terminal, or intermediate amino acids of a peptide molecule. The binder can preferably bind to chemically modified or labeled amino acids, rather than to unmodified or unlabeled amino acids. For example, the binder can preferably bind to amino acids that have been modified with acetyl, amidoyl, dansyl, PTC, DNP, SNP, etc., rather than to amino acids that do not have such a moiety. The binder can bind to naturally occurring or synthetically derived post-translational modifications of peptide molecules. The binder can exhibit selective binding to components or features of macromolecules (e.g., the binder can selectively bind to one residue from 20 possible natural amino acid residues and bind with the other 19 natural amino acid residues with very low affinity or not bind to them at all). Binders can exhibit less selective binding, meaning they can bind multiple components or features of a macromolecule (e.g., a binder can bind to two or more different amino acid residues with similar affinity). Binders can contain tags that can be coupled to the binder via linkers.

[0158] As used herein, the term "connector" generally refers to a molecule or part that participates in linking two or more molecules. A connector can facilitate covalent or non-covalent interactions between two or more molecules. A connector can be a cross-linking agent. A connector can be monofunctional, bifunctional, trifunctional, tetrafunctional, or multifunctional. A connector can be or include nucleotides, nucleotide analogs, amino acids, peptides, polypeptides, or non-nucleotide chemical parts, such as organic or inorganic compounds. A connector can include polymers such as polyethylene glycol (PEG), poly-L-lysine (PLL), poly(DL-lactic acid) (PLA), poly(DL-lactide-co-glycoside) (PLGA), polyornithine, polyarginine, etc. A connector can include one or more reactive ends, such as amine reactive groups, carboxyl reactive groups, thiol reactive groups, hydroxyl reactive groups, etc. In some examples, connectors can be used to link different molecular types, such as different biomolecule types, such as peptides and nucleic acid molecules, lipids and peptides, carbohydrates and peptides, etc.; non-biomolecule types; or biomolecules to non-biomolecules. For example, adapters can be used to link binders to tags, tags to macromolecules (e.g., peptides, nucleic acid molecules), macromolecules to solid carriers, and tags to solid carriers. Adapters can link two molecules via enzymatic or chemical reactions (e.g., click chemistry). Adapters can also link more than two molecules, for example, via enzymatic or chemical reactions.

[0159] As used herein, the term “combination” generally refers to a covalent or ionic interaction between two entities (e.g., molecules, compounds, or combinations thereof).

[0160] As used herein, the term "tag" generally refers to a molecule or part conjugated to a molecule. Tags can include detectable markers such as fluorophores or fluorescent proteins, radioisotopes, enzymes (e.g., chromophores or fluorescent proteins, proteins that can catalyze chromophore substrates), mass tags, haptens (e.g., biotin, digoxigenin, urushiol, fluorescein), vibrational or FTIR tags (e.g., alkyne groups). Tags can include biomolecules such as nucleic acid molecules, proteins, lipids, carbohydrates, or combinations thereof. Tags can include one or more nucleic acid molecules that may optionally encode information about the tag or a molecule (e.g., a binder, such as an antibody) to which the tag is conjugated. For example, a tag can include a nucleic acid barcode molecule. Tags can include organic or inorganic compounds.

[0161] As used herein, the term "barcode" generally refers to an identifiable feature that can be used to distinguish similar items. Barcodes can contain approximately 2 to approximately 30 bases (e.g., 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36...). 1, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74 1, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100, 105, 110, 115, 120, 125, 130, 135, 140, 145 or A nucleic acid molecule (150 bases) can provide a unique identifier tag or origin information for molecules (e.g., proteins, polypeptides, peptides), binding agents, a group of binding agents from a binding cycle, sample molecules, a group of samples, molecules within a compartment (e.g., droplets, beads, partitions, or separated locations), macromolecules within a group of compartments, fractions of macromolecules, a group of macromolecular fractions, a spatial region or a group of spatial regions, a macromolecular library, or a binding agent library. Barcodes can be artificial sequences or naturally occurring sequences, including peptides, proteins, protein complexes, carbohydrates, and synthetic polymer materials. In some embodiments, each barcode within a group of barcodes is different. In other embodiments, a subset of barcodes within a group of barcodes is different, for example, at least about 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 97%, or 99% of the barcodes in the group of barcodes are different. Barcode clusters can be randomly generated or non-randomly generated. Barcode clusters may include error-corrected barcodes. Barcodes can be used to computationally deconvolve sequence reads derived from individual molecules, samples, libraries, etc. Barcodes can contain multiplexing information, for example, generated from different samples, compartments, individual molecules, etc. Barcodes can also be used to deconvolve assemblies of molecules already distributed into small compartments for enhanced mapping.For example, instead of mapping peptides back to the proteome, it's possible to map peptides back to the protein molecules or protein complexes from which they originate. Barcodes can contain any useful structural parts or motifs, such as hairpins, loop sequences, or spacer regions. Barcodes can contain artificial or modified nucleic acids, such as locked nucleic acids (LNAs), protein nucleic acids (PNAs), hexitol nucleic acids (HNAs), cyclohexane nucleic acids (CeNAs), or combinations thereof. Barcodes can contain proteins (e.g., Tal effectors, Cas proteins (e.g., Cas9), Argonaut, or coiled helices) or be generated using proteins.

[0162] As used in this article, "sample barcode," also known as "sample label," typically refers to a barcode molecule that contains identification information about the sample from which the barcode molecule is derived.

[0163] As used herein, a "spatial barcode" generally refers to a barcode molecule that contains identification information about the region of a 2-D or 3-D sample (e.g., a tissue section) from which the molecule originates or is derived. Spatial barcodes can be used for molecular pathology on tissue sections. Spatial barcodes can allow for multiplexing of multiple samples or libraries from tissue sections.

[0164] As used herein, a “time barcode” generally refers to a barcode molecule that contains time-based information associated with the barcoded molecule. The types of time-based data encoded in a time barcode can include information such as the lifetime of the barcoded molecule, sample collection time, time or duration since the start of the experiment or induction by stimulus, information about the age of the cell or tissue, sequences of interactions between molecules, etc. It is possible to combine different types of barcodes (e.g., spatial, temporal, cell-specific) into a single multiplexed barcode.

[0165] As used herein, the terms “nucleic acid sequence” or “oligonucleotide sequence” generally refer to a continuous string of nucleotide bases and may refer to the specific placement of nucleotide bases relative to each other when they appear in an oligonucleotide. Similarly, the terms “peptide sequence” or “amino acid sequence” refer to a continuous string of amino acids and may refer to the specific placement of amino acids relative to each other when they appear in a polypeptide.

[0166] The "nucleic acid molecule" according to the invention may include any polymer or oligomer of nucleotides such as pyrimidine and purine bases, such as cytosine, thymine, and uracil, and adenine and guanine, respectively, and combinations thereof. The nucleotide sequence may include any deoxyribonucleotide, ribonucleotide, hexitol-nucleotide, cyclohexane-nucleotide, peptide nucleic acid component, and any chemical variant thereof, such as methylated, 7-denitropurine analogs, 8-halopurine analogs, hydroxymethylated or glycosylated forms of these bases, etc. The polymer or oligomer may be heterogeneous or homogeneous in composition and may be isolated from naturally occurring sources or may be produced artificially or synthetically. The nucleic acid molecule may include DNA, RNA, HNA, CeNA, or mixtures thereof and may exist permanently or transitionally in single-stranded or double-stranded form (including homoduplexes, heteroduplexes, and hybrid states).

[0167] The term "complementarity" or "complementarity" refers to a polynucleotide (i.e., a nucleotide sequence) associated with the Watson-Crick base pairing rule. For example, the sequence "5'-AGT-3'" is complementary to the sequence "5'-ACT-3'". Complementarity can be "partial," where only some bases of the nucleic acid match according to the base pairing rule, or it can exist as "complete" or "full" complementarity between nucleic acids. The degree of complementarity between nucleic acid chains can have a significant impact on the hybridization efficiency and strength under certain conditions.

[0168] As used herein, the term "hybridization" refers to the pairing of complementary nucleic acids. Hybridization and hybridization strength (e.g., the strength of association between nucleic acids) are influenced by factors such as the degree of complementarity between the nucleic acids, the stringency of the conditions involved, and the melting temperature of the resulting hybrid. Hybridization methods involve annealing one nucleic acid to another complementary nucleic acid, for example, based on Watson-Crick base pairing.

[0169] As used herein, the term "proteomics" generally refers to the quantitative and / or qualitative analysis of the proteome within a sample, such as a biological sample from cells, tissues, or body fluids. Proteomics can include analyzing the spatial distribution of proteins within a sample (e.g., cells and / or tissues). Proteomics can also include the study of the dynamic state of the proteome, such as how one or more proteins change over time. The proteome can include multiple “groups”, such as the kinase group; the secretion group; the receptor group (e.g., the GPCR group); the immune proteome; the nutritional proteome; subgroups of the proteome defined by post-translational modifications (e.g., phosphorylation, ubiquitination, methylation, acetylation, glycosylation, oxidation, lipidation, and / or nitrosation), such as the phosphorylated proteome (e.g., the phosphotyrosine proteome, tyrosine kinase group, and tyrosine phosphatase group), glycoproteome, etc.; subgroups of the proteome associated with tissues or organs, developmental stages, or physiological or pathological conditions (e.g., cancer or disease); subgroups of the proteome associated with cellular processes (e.g., cell cycle, differentiation (or dedifferentiation), cell death, senescence, cell migration, transformation, or metastasis); or any combination thereof.

[0170] The terminal amino acid with a free amino group at one end of the peptide chain may be referred to herein as an "N-terminal amino acid" (NTAA). The terminal amino acid with a free carboxyl group at the other end of the chain may be referred to herein as a "C-terminal amino acid" (CTAA). The amino acids that make up a peptide may be numbered sequentially, and the length of the peptide is "n" amino acids. As used herein, in some cases, NTAA may be considered as the nth amino acid (also referred to herein as "n NTAA"). In such cases, the next amino acid is the (n-1)th amino acid, then the (n-2)th amino acid, and so on, decreasing the length of the peptide from the N-terminus to the C-terminus. Alternatively, CTAA may be considered as the nth amino acid (also referred to herein as "n CTAA"). In such cases, the next amino acid is the (n-1)th amino acid, then the (n-2)th amino acid, and so on, decreasing the length of the peptide from the C-terminus to the N-terminus. NTAA, CTAA, or both may be chemically modified or labeled.

[0171] As used herein, the terms “determine,” “measure,” “evaluate,” and “determine” are used interchangeably and include both quantitative and qualitative determinations.

[0172] As used in this article, the term “unique molecular identifier” or “UMI” generally refers to a molecular barcode that contains indexing information. UMIs can contain lengths from approximately 3 to approximately 150 bases (3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57...). Nucleic acid molecules with 1, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100, 105, 110, 115, 120, 125, 130, 135, 140, 145, or 150 bases. A UMI can provide a unique identifier tag for each molecule (e.g., peptide, binder, nucleic acid molecule) that contains or is associated with a UMI. A UMI can contain random sequences (e.g., random N-mers).

[0173] As used herein, a “derivative” of a nucleic acid molecule generally refers to a nucleic acid molecule derived from an originating nucleic acid molecule. A derivative may have the same or substantially the same nucleotide sequence as the originating nucleic acid molecule, or it may contain a complementary or partially complementary sequence to the originating nucleic acid molecule. A derivative may be the same type of nucleic acid as the originating nucleic acid molecule (e.g., DNA or RNA), or it may be a different type of nucleic acid (e.g., cDNA derived from an RNA molecule). Nucleic acid molecule derivatives may exhibit sequence identity with the originating nucleic acid molecule. Further treatments derived from the originating nucleic acid molecule may also be applied to the derivative nucleic acid molecule, such as chemical or enzymatic modifications, splicing, ligation, polymerization, fragmentation, tagging (e.g., using transposases), digestion, etc.

[0174] Derivative peptides or peptides can be derived from the originating peptide (or peptide). The derivative may contain the same amino acid sequence as the originating peptide, or the sequence may be different. Derivative peptides can be generated from the originating peptide or subjected to additional treatments derived from the originating peptide, such as chemical or enzymatic modifications. Derivative peptides may contain one or more tags, nucleic acid molecules, barcode molecules, labels (e.g., detectable labels), fluorophores, probes, adapters, post-translational modifications, chemical protecting groups, or other chemical motifs.

[0175] Derivatives of amino acids generally refer to modified amino acids derived from their origin. Derivatives may have the same or substantially the same chemical composition as the origin amino acid, or they may contain different chemical compositions. Derivatives can be chemically modified or enzymatically modified amino acids. Derivatives can be naturally occurring or synthetic.

[0176] As used herein, the term "compartment" or "zone" generally refers to a physical region or volume that separates or isolates a molecular subgroup from the molecular sample. For example, a compartment can separate individual cells from other cells, or separate a subgroup of the proteome of a sample from the rest of the sample's proteome. A compartment can be an aqueous compartment (e.g., a microfluidic droplet), a solid compartment (e.g., a microburette or microtiter well on a plate, tube, vial, gel bead), or a separation region on a surface. A compartment can contain one or more beads to which macromolecules can be immobilized.

[0177] As used herein, the terms “solid support,” “solid surface,” “solid substrate,” or “substrate” refer to any solid material to which molecules can associate directly or indirectly, including porous and non-porous materials. Molecules can associate with the substrate through covalent or non-covalent interactions or combinations thereof. The substrate can be two-dimensional (e.g., a flat surface) or three-dimensional (e.g., a gel matrix or beads). In non-limiting examples, solid supports can include beads, microspheres, arrays, glass surfaces, silicon surfaces, plastic surfaces, filters, membranes, nylon or other polymers, silicon wafers, flow-through chips, and flow-through cells (e.g., custom-designed or commercially available flow-through cells, such as those from Illumina). ®HiSeq, iSeq, MiniSeq, NextSeq, NovaSeq, MiSeq flow cells), microfluidic devices or chips or their surfaces, biochips including signal transduction electronic devices, channels, microtiter wells, ELISA plates, rotating interferometric disks, nitrocellulose membranes, nitrocellulose-based polymer surfaces, polymer matrices, nanoparticles or microspheres. Materials used for solid supports include, but are not limited to, acrylamide, agarose, cellulose, nitrocellulose, glass, gold, quartz, polystyrene, polyvinyl acetate, polypropylene, polymethacrylate, polyethylene, polyethylene oxide, polysilicate, polycarbonate, Teflon, fluorocarbons, nylon, silicone rubber, polyanhydride, polyglycolic acid, polylactic acid, polyorthoester, functionalized silanes, polypropyl fumarate, collagen, glycosaminoglycans, polyamino acids, dextran, or any combination thereof. Solid supports also include films, membranes, bottles, disks, fibers, woven fibers, molded polymers such as tubes, particles, beads, microspheres, microparticles, or any combination thereof. For example, when the solid surface is a bead, the beads can include, but are not limited to, ceramic beads, polystyrene beads, polymer beads, methylstyrene beads, agarose beads, acrylamide beads, solid beads, porous beads, magnetic or paramagnetic beads, glass beads, carboxyl beads, or controlled-pore beads. The beads can be spherical or irregularly shaped. The size of the beads can range from nanometers (e.g., 100 nm) to millimeters (e.g., 1 mm). In some embodiments, the size of the beads ranges from about 0.2 micrometers to about 200 micrometers or from about 0.5 micrometers to about 5 micrometers.In some embodiments, the diameter of the beads can be approximately 1 μm, 1.5 μm, 2 μm, 2.5 μm, 2.8 μm, 3 μm, 3.5 μm, 4 μm, 4.5 μm, 5 μm, 5.5 μm, 6 μm, 6.5 μm, 7 μm, 7.5 μm, 8 μm, 8.5 μm, 9 μm, 9.5 μm, 10 μm, 11 μm, 12 μm, 13 μm, 14 μm, or 15 μm. m, 16μm, 17μm, 18μm, 19μm, 20μm, 21μm, 22μm, 23μm, 24μm, 25μm, 26μm, 27μm, 28μm, 29μm , 30μm, 31μm, 32μm, 33μm, 34μm, 35μm, 36μm, 37μm, 38μm, 39μm, 40μm, 41μm, 42μm, 43μm, 4 4μm, 45μm, 46μm, 47μm, 48μm, 49μm, 50μm, 51μm, 52μm, 53μm, 54μm, 55μm, 56μm, 57μm, 58 μm, 59μm, 60μm, 61μm, 62μm, 63μm, 64μm, 65μm, 66μm, 67μm, 68μm, 69μm, 70μm, 71μm, 72μm The sizes of the solid carriers are 73μm, 74μm, 75μm, 76μm, 77μm, 78μm, 79μm, 80μm, 81μm, 82μm, 83μm, 84μm, 85μm, 86μm, 87μm, 88μm, 89μm, 90μm, 91μm, 92μm, 93μm, 94μm, 95μm, 96μm, 97μm, 98μm, 99μm, or 100μm. In some embodiments, the "bead" solid carrier can refer to a single bead or multiple beads. The solid carrier can exhibit any useful geometry, such as pyramid, cube, cylinder, spiral, sphere, globular, rod, disk, arrowhead, spring shape, teardrop, prism, tetrapod shape, or any other useful geometry.

[0178] As used herein, “sequencing” generally refers to determining the sequence of nucleotides (base sequences) in (A) a nucleic acid sample (e.g., DNA or RNA); or determining the sequence of amino acids in (B) all or part of a polymer (such as a protein, peptide, or other multimeric molecule). Many technologies are available, such as Sanger sequencing or high-throughput sequencing (HTS). Sanger sequencing can involve sequencing via detection through (capillary) electrophoresis, in which up to 384 capillaries can be sequenced in a single run. High-throughput sequencing involves sequencing thousands or millions or more sequences in parallel at once. HTS can be defined as next-generation sequencing (NGS), a technique based on solid-phase pyrosequencing, or as next-generation sequencing based on single-nucleotide real-time sequencing (SMRT). HTS technologies are available, such as those offered by Roche, Illumina, and Applied Biosystems (Life Technologies). Other high-throughput sequencing technologies are described and / or available from Helicos, Pacific Biosciences, Complete Genomics, Ion Torrent Systems, Oxford Nanopore Technologies, Nabsys, ZS Genetics, and GnuBio.

[0179] As used in this article, "next-generation sequencing" refers to high-throughput sequencing methods that allow the parallel sequencing of millions to billions of molecules. Examples of next-generation sequencing methods include sequencing-by-synthesis, sequencing-by-ligation, sequencing-by-hybridization, polymerase cloning sequencing, ion semiconductor sequencing, nanopore sequencing, and pyrosequencing. By attaching primers to a solid substrate and complementary sequences to nucleic acid molecules, the nucleic acid molecules can hybridize with the solid substrate via primers. Multiple copies can then be generated in discrete regions on the solid substrate using polymerase amplification (these groups are sometimes referred to as polymerase communities or polymerase colonies). Therefore, during the sequencing process, nucleotides at specific locations can be sequenced multiple times (e.g., hundreds or thousands of times)—this depth of coverage is called "deep sequencing." Examples of high-throughput nucleic acid sequencing technologies include platforms provided by Illumina, BGI, Qiagen, ThermoFisher, and Roche, including parallel bead arrays, sequencing-by-synthesis, sequencing-by-ligation, capillary electrophoresis, electronic microarrays, “biochips,” microarrays, parallel microarrays, and single-molecule arrays, as well as sequencing based on zero-mode waveguides, some of which are reviewed by Service (Science 311:1544-1546, 2006).

[0180] As used herein, “analysis of macromolecules” means the quantification, characterization, differentiation, or combination thereof of all or part of the components of a molecule (e.g., macromolecules, biomolecules such as proteins, amino acids, nucleic acid molecules, etc.). For example, analyzing peptides, polypeptides, or proteins may include determining all or part of the amino acid sequence (continuous or discontinuous) of the peptide. Analyzing macromolecules may include the partial identification of macromolecular components. For example, partial identification of amino acids in a protein sequence can identify amino acids in a protein as belonging to possible subgroups of amino acids. Analysis can be performed sequentially, for example, starting with the analysis of n NTAAs and then proceeding to the next amino acid of the peptide (i.e., n-1, n-2, n-3, etc.). In such cases, sequencing can be performed by cleaving the n NTAAs, thereby converting the (n-1)th amino acid of the peptide into an N-terminal amino acid (referred to herein as “n-1 NTAA”). Similarly, peptide analysis can begin from the C-terminus to the N-terminus, with each round of cleavage from the C-terminus producing a new CTAA. Cleavage of the n CTAA converts the (n-1)th amino acid of the peptide into a C-terminal amino acid, referred to herein as “n-1 CTAA”. Analyzing peptides may also include determining the presence and frequency of post-translational modifications on the peptide, which may or may not include information about the sequence order of post-translational modifications on the peptide. Analyzing peptides may also include determining the presence and frequency of epitopes within the peptide, which may or may not include information about the sequence order or location of epitopes within the peptide. Analyzing peptides may include combining different types of analyses, such as obtaining epitope information, amino acid sequence information, post-translational modification information, or any combination thereof.

[0181] As used herein, the term "analyte" generally refers to a substance of interest to be further identified, characterized, or measured. In non-limiting examples, an analyte can be an ion, chemical, compound, small molecule, element, particle, metal, biomolecule, macromolecule, metabolite, lipid, carbohydrate, peptide or protein, nucleic acid molecule, organelle, or cell. An analyte can be naturally occurring or synthetic. An analyte can be a solid, semi-solid, liquid, semi-liquid, gas, or plasma. An analyte can be characterized qualitatively or quantitatively. A portion of the analyte can be analyzed. For example, an analyte can be a peptide, and the constituent amino acids can be analyzed. An analyte can include polymers, also referred to herein as "polymer analytes," which generally refers to an analyte of interest comprising one or more monomers. In non-limiting examples, a polymer analyte can be a group of ions, chemicals, compounds, small molecules, elements, particles, metals, or biomolecules, macromolecules, metabolites, lipids, carbohydrates, peptides or proteins, nucleic acid molecules, organelles, or cells.

[0182] As used herein, the term "array" generally refers to a group of molecules attached to one or more solid carriers such that a molecule at one address can be distinguished from molecules at other addresses. An array may include different molecules located at different addresses on a solid carrier. Alternatively, an array may include individual solid carriers, each functioning as an address carrying a different molecule, wherein the different molecules can be identified based on the position of the solid carrier on the surface to which it is attached or based on the position of the solid carrier in a liquid (such as a fluid flow). The molecules in an array may be, for example, nucleic acids such as SNAPs, polypeptides, proteins, peptides, oligopeptides, enzymes, ligands, or receptors such as antibodies, functional fragments of antibodies, or aptamers. The addresses of an array may optionally be optically observable, and in some configurations, adjacent addresses may be optically distinguishable when detected using the methods or apparatus described herein.

[0183] As used herein, the term "functionalized" refers to any material or substance that has been modified to include functional groups. Functionalized materials or substances can be naturally or synthetically functionalized. For example, peptides can be naturally functionalized with phosphate groups, oligosaccharides (e.g., glycosyl, glycosylphosphatidylinositol, or phosphate glycosyl), nitrosyl, methyl, acetyl, lipids (e.g., glycosylphosphatidylinositol, myristoyl, or isoprenyl), ubiquitin, or other naturally occurring post-translational modifications. Functionalized materials or substances can be functionalized for any given purpose, including altering chemical properties (e.g., altering hydrophobicity or surface charge density) or reactivity (e.g., the ability to react with a moiety or reagent to form a covalent bond with that moiety or reagent).

[0184] As used herein, the terms “click reaction,” “click chemistry,” or “bioorthogonal reaction” refer to a single-step, thermodynamically favorable conjugation reaction utilizing biocompatible reagents. Click reactions may not utilize toxic or biologically incompatible reagents (e.g., acids, bases, heavy metals) or produce toxic or biologically incompatible byproducts. Click reactions may utilize aqueous solvents or buffers (e.g., phosphate buffers, Tris buffers, saline buffers, MOPS, etc.). A click reaction may be thermodynamically favorable if it has a negative reaction Gibbs free energy, such as less than about -5 kJ / mol, -10 kJ / mol, -25 kJ / mol, -50 kJ / mol, -100 kJ / mol, -200 kJ / mol, -300 kJ / mol, -400 kJ / mol, or less than -500 kJ / mol. Exemplary bioorthogonal and click reactions are described in detail in WO2019 / 195633A1, the full text of which is incorporated herein by reference. Exemplary click reactions may include metal-catalyzed azide-alkyne cycloaddition, strain-promoted azide-alkyne cycloaddition, strain-promoted azide-nitroketone cycloaddition, strained olefin reactions, thiol-ene reactions, Diels-Alder reactions, electron-demanding Diels-Alder reactions, [3+2] cycloaddition, [4+1] cycloaddition, nucleophilic substitution, dihydroxylation, thiol-alkyne reactions, photoclicking, nitroketone dipolar cycloaddition, norbornene cycloaddition, oxanorbornediene cycloaddition, tetrazine linkage, and tetrazolium photoclicking. Exemplary functional groups or reactive handles for performing click reactions may include alkenes (e.g., straight-chain alkenes or cyclic alkenes, such as trans-cyclooctene (TCO)), alkynes (e.g., straight-chain alkynes or cyclic alkynes (e.g., cyclooctene or derivatives thereof, such as azido-dimethoxycyclooctene (DIMAC), symmetric pyrrolocyclooctene (SYPCO), pyrrolocyclooctene (PYRROC), difluorocyclooctene (DIFO), α,α-bis(trifluoromethyl)pyrrolocyclooctene (TRIPCO), bicyclo[6.1.0]nonyne (BCN), dibenzocyclooctene (DIBO), and difluorinated cyclooctene (DIFO). Difluorobenzocyclooctyne (DIFBO), dibenzozacyclooctyne (DBCO), difluoro-aza-dibenzocyclooctyne (F2-DIBAC), biaryl-azacyclooctyneone (BARAC), difluorodimethoxydibenzocyclooctyne alcohol (FMDIBO), difluorodimethoxydibenzocyclooctyneone (keto-FMDIBO), and 3,3,6,6-tetramethylthioheptyne (TMTH), TMTH-sulfonylimine (TMTHSI), azides, epoxides, amines, thiols, nitrones, isonitriles, isocyanates, aziridines, activated esters, and tetrazides, triazoles, and combinations, variants, or derivatives thereof.The click chemistry portion can be subjected to conditions sufficient to allow the first click chemistry portion to react with the second click chemistry portion, such as the provision of a metal catalyst, a suitable solvent, pH, temperature, ion concentration, or light / energy, and for any useful duration.

[0185] As used herein, the terms “group” and “part” are intended to be synonymous when used to refer to the structure of a molecule. These terms refer to a component or part of a molecule. Unless otherwise specified, these terms do not necessarily indicate the relative size of a component or part compared to the molecule. Unless otherwise specified, these terms do not necessarily indicate the relative size of a component or part compared to any other component or part of the molecule. A group or part may contain one or more atoms.

[0186] As used herein, "primer" generally refers to a nucleic acid molecule that can initiate the synthesis of a nucleic acid molecule (e.g., DNA or RNA). Primers can be single-stranded. Primers may contain one or more recognition sites of proteins (e.g., polymerases, restriction enzymes, lyases, nucleases, etc.) for binding to the primer or to a template strand. Primers may contain DNA, RNA, or other nucleic acid analogs or atypical bases (e.g., spacer regions, uracil, abase-free sites). Primers may optionally contain any number of functional sequences, such as sequencing primer sequences (e.g., P5 or P7 sequences), sequencing primer-binding sequences, read sequences (e.g., R1 or R2 sequences), enzyme recognition sites (e.g., restriction sites or nuclease recognition sites), abase-free sites, cleavage sites, transposition sites, spacer sequences, barcode sequences, unique molecular identifiers (UMIs), etc.

[0187] "Amplification" or "the process of amplifying" typically refers to a polynucleotide amplification reaction, which involves the replication of a group of polynucleotides from one or more starting sequences. Amplification can refer to various amplification reactions, including but not limited to polymerase chain reaction (PCR), linear polymerase reaction, nucleic acid sequence-based amplification, rolling circle amplification, and similar reactions. Amplification reactions can produce amplicones.

[0188] As used herein, "adaptor" generally refers to a short nucleic acid molecule (e.g., about 10 to about 100 base pairs in length). Adaptors may comprise short double-stranded DNA molecules. Adaptors may be attached to the ends of DNA fragments or amplicon, for example, via polymerization or ligation. Adaptors may comprise synthetic oligonucleotides, such as oligonucleotides having nucleotide sequences that are at least partially complementary to each other. Adaptors may have blunt ends, staggered ends (also referred to herein as 3' or 5' "protruding end sequences" or "sticky ends"), or both blunt and staggered ends. Adaptors may be attached (e.g., via ligation) to fragments to provide adaptor-ligated fragments; the fragments ligated by the adaptor may serve as starting points for subsequent operations (e.g., for amplification or sequencing). Adaptors may be functionalized, for example, conjugated to tags, probes, detectable markers, or affinity-capturing reagents (e.g., biotin or streptavidin).

[0189] As used herein, the term "capture moiety" generally refers to a molecule configured to couple to another moiety or molecule. Capture moieties can be biomolecules, such as lipids, carbohydrates, sugars, amino acids, peptides or proteins, nucleotides, nucleic acid molecules, metabolites, or combinations thereof (e.g., glycoproteins, lipoproteins, glycosaminoglycans, etc.). Capture moieties can be small molecules, organic compounds, inorganic compounds, metals, polymers, ions, or other molecules or molecular compounds. Capture moieties can include macromolecules. Capture moieties can include enzymes, antibodies, antibody fragments, nanobodies, aptamers, biotin, streptavidin, avidin, neutral avidin, or analogs or derivatives thereof. Capture moieties can include more than one molecule, such as dimers, trimers, tetramers, pentamers, hexamers, heptamers, octamers, etc. Capture moieties can be a solid substrate or part of a solid substrate, or they can be separated from the substrate, for example, in a fluid medium (e.g., air, in a liquid solution). Capture moieties can be specific to one or more binding partners. The capture portion can bind to one molecule or part (monovalent) or multiple molecules or parts (polyvalent).

[0190] As used herein, the abbreviations for natural 1-enantiomers of amino acids are conventional and may be as follows: alanine (A, Ala); arginine (R, Arg); asparagine (N, Asn); aspartic acid (D, Asp); cysteine ​​(C, Cys); glutamic acid (E, Glu); glutamine (Q, Gln); glycine (G, Gly); histidine (H, His); isoleucine (I, Ile); leucine (L, Leu); lysine (K, Lys); methionine (M, Met); phenylalanine (F, Phe); proline (P, Pro); serine (S, Ser); threonine (T, Thr); tryptophan (W, Trp); tyrosine (Y, Tyr); valine (V, Val). Unless otherwise specified, X may represent any amino acid. In some respects, X may be asparagine (N), glutamine (Q), histidine (H), lysine (K), or arginine (R). These amino acids are also referred to in the form of “[amino acid][residue / residue]” (e.g., lysine residue, lysine residues, leucine residue, leucine residues, etc.).

[0191] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure pertains. Although similar or equivalent methods and materials may be used in the practice or testing of this disclosure, suitable methods and materials are described below.

[0192] amino acid binders

[0193] This document provides methods, compositions, and kits for processing and analyzing polymer analytes (e.g., peptides, polymers, nucleic acid molecules, etc.) in a highly parallel and accurate manner. This disclosure provides methods for sequencing polymer analytes containing monomers using binding agents that specifically or semi-specifically identify and bind to individual monomer types of the polymer analyte. The binding agents can be tagged with detectable markers (e.g., fluorophores or fluorescent proteins, mass tags, radioisotopes, nucleic acid molecules (e.g., nucleic acid barcode molecules)) that can identify or be used to identify the binding agent or its homologs. In some aspects, the systems and methods of this disclosure may include providing a polymer analyte containing multiple monomers, cleaving monomers from the polymer analyte to produce free monomers, coupling monomers to a capture portion (e.g., by local tethering), and providing a monomer-specific binding agent that identifies the cleaved (or free) monomers. Further exemplary methods and systems for identifying monomers of polymer analytes are described elsewhere herein and are also provided in PCT / US2023 / 017954, filed April 7, 2023, the entire contents of which are incorporated herein by reference.

[0194] In some aspects of this disclosure, the binding agents provided herein are used for sequencing or characterizing proteins or peptides containing amino acids. The binding agents of this disclosure can bind with a single type of protein amino acid or its derivative with higher specificity or higher affinity than the binding agents bind with all other (e.g., all 19 other) types of protein amino acids or their derivatives, such as chemically or enzymatically derivatized amino acids, which may include post-translational modified amino acids. Therefore, the binding agents can be used to identify individual amino acids or their derivatives with high accuracy. In some aspects, the binding agents provided herein are configured to bind free amino acids or their derivatives. Free amino acids (or their derivatives) can be generated, for example, by cleaving amino acids from a peptide or protein. Free amino acids can be internal amino acids or terminal amino acids, such as N-terminal or C-terminal amino acids.

[0195] In some embodiments, the binder specifically binds to alanine (A, Ala) or a derivative thereof; in some embodiments, the binder specifically binds to arginine (R, Arg) or a derivative thereof; in some embodiments, the binder specifically binds to asparagine (N, Asn) or a derivative thereof; in some embodiments, the binder specifically binds to aspartic acid (D, Asp) or a derivative thereof; in some embodiments, the binder specifically binds to cysteine ​​(C, Cys) or a derivative thereof; in some embodiments, the binder specifically binds to glutamic acid (E, Glu) or a derivative thereof; in some embodiments, the binder specifically binds to glutamine (Q, Gln) or a derivative thereof; in some embodiments, the binder specifically binds to glycine (G, Gly) or a derivative thereof; in some embodiments, the binder specifically binds to histidine (H, His) or a derivative thereof; in some embodiments, the binder specifically binds to isoleucine (I, Ile) or a derivative thereof. In some embodiments, the binding agent specifically binds to leucine (L, Leu) or a derivative thereof; in some embodiments, the binding agent specifically binds to lysine (K, Lys) or a derivative thereof; in some embodiments, the binding agent specifically binds to methionine (M, Met) or a derivative thereof; in some embodiments, the binding agent specifically binds to phenylalanine (F, Phe) or a derivative thereof; in some embodiments, the binding agent specifically binds to proline (P, Pro) or a derivative thereof; in some embodiments, the binding agent specifically binds to serine (S, Ser) or a derivative thereof; in some embodiments, the binding agent specifically binds to threonine (T, Thr) or a derivative thereof; in some embodiments, the binding agent specifically binds to tryptophan (W, Trp) or a derivative thereof; in some embodiments, the binding agent specifically binds to tyrosine (Y, Tyr) or a derivative thereof; in some embodiments, the binding agent specifically binds to valine (V, Val) or a derivative thereof. In some embodiments, the binding agent specifically binds to a post-translational modified amino acid or a derivative thereof.

[0196] In some embodiments, the binder specifically binds to an amino acid derivative. The amino acid derivative may be a chemically or enzymatically modified amino acid and may or may not contain post-translational modifications. The amino acid derivative may be derived from a cleaved or free amino acid, or the amino acid derivative may be generated by a cleavage process. For example, in the methods and systems described herein, the amino acid of a peptide may contact a linker containing an amino acid reactive group and cleave from the peptide to produce a free amino acid or an amino acid-linker complex; thus, the binder provided herein can specifically bind to and recognize free amino acids, amino acid-linker complexes, or portions thereof. In one such example, the peptide may react with a linker containing a phenyl isothiocyanate (PITC) moiety (amino acid reactive group). The PITC moiety can react with the N-terminal amino acid of the peptide to produce a phenylthiocarbamoyl (PTC)-derived amino acid, which can be recognized by the binder. Alternatively, the binder can recognize other derivatives of PITC that react with amino acids, including amino acids derived from anilinethiazolinone (ATZ) or hydantoin (PTH). In some cases, ATZ derivatives are generated by cleaving amino acids from a peptide. In other cases, the derivatized amino acid (e.g., the ATZ derivative) is further derivatized (e.g., derivatized into PTH or PTC derivatives), and this further derivatized amino acid can be recognized by a binding agent.

[0197] Amino acid derivatives may contain any number of chemical modifications, which may be naturally occurring or synthetic. Non-limiting examples of chemical modifications of amino acids include, for example, alkylation of cysteine ​​residues (e.g., using 4-vinylpyridine, iodoacetamide, iodoacetate, chloroacetate); acetylation, such as reacting serine or threonine residues to form esters (e.g., using acetyl chloride) or using acetic anhydride; oxidation, such as converting cysteine ​​residues to sulfoalanine; reduction (e.g., using reducing agents such as dithiothreitol, β-mercaptoethanol, or TCEP); addition of protecting groups, such as phosphorylated residues may be protected (e.g., using β-elimination of phosphate groups and optional Michael addition of thiol groups, e.g., as described in Knight et al., Nat. Biotechnology. 21, 1047-1054 (2003), the entire text of which is incorporated herein by reference), etc. The amino acid can be modified or derivatized by a protecting group or part thereof, such as methyl, formyl, ethyl, acetyl, tert-butyl, anisyl, benzyl, trifluoroacetyl, N-hydroxysuccinimide, tert-butoxycarbonyl (Boc), benzoyl, 4-methylbenzyl, thioanizyl, thiotolyl, benzyloxymethyl, 4-nitrophenyl, benzyloxycarbonyl, 2-nitrobenzoyl, 2-nitrophenylsulfinyl, 4-toluenesulfonyl, pentafluorophenyl, diphenylmethyl, 2-chlorobenzyloxycarbonyl, 2,4,5-trichlorophenyl, 2-bromobenzyloxycarbonyl, 9-fluorenylmethoxycarbonyl (FMOC), triphenylmethyl, or 2,2,5,7,8-pentamethyl-chromium-6-sulfonyl group. Amino acids can be modified or derivatized using protecting agents such as carboxyethyl methyl thiosulfonate (CEMTS), thiazolidinyl, mercaptophenylacetic acid, cyanobenzothiazole (e.g., for the esterification of N-terminal cysteine), acetamidomethyl, 2-methylsulfonyl ethyl-oxycarbonyl, etc. In some cases, isothiocyanates (e.g., PITC) can be used to modify or derivatize lysine residues (e.g., to react the primary amine of lysine residues) to produce modified lysine residues.

[0198] In some cases, the binder specifically binds to amino acid derivatives but not to native amino acids, such as when the amino acid is present in the peptide. In other cases, the binder specifically binds to or recognizes a free amino acid (e.g., a cleaved amino acid) or a free amino acid derivative (e.g., an amino acid-proximity complex), but not to the amino acid when it is attached to the peptide. For example, a linker containing an amino acid reactant can react with the terminal amino acid of a peptide to produce an amino acid-proximity complex. An amino acid-proximity complex can be cleaved from a peptide to produce a cleaved amino acid-proximity complex. In some cases, the binder can bind to and recognize a cleaved amino acid-proximity complex, but not to the amino acid-proximity complex attached to the peptide. In some cases, the binder can bind to and recognize a cleaved amino acid-proximity complex, but not to the amino acid when it is attached to the peptide.

[0199] Derivatives of amino acids can refer to free or cleaved amino acids, amino acid-proximity complexes, or portions of amino acids, free (e.g., cleaved) amino acids, or amino acid-proximity complexes. Derivatives of amino acids can also include further derivatized amino acids (e.g., amino acids that have undergone more than one type of derivatization). For example, derivatized amino acids produced using a linker containing PITC include PTC, PTH, or ATZ-derived amino acids. Therefore, the binding agents described herein are capable of recognizing and binding to specific amino acid types or derivatized amino acid types, amino acid-proximity complexes, cleaved amino acids, or amino acid-proximity complexes. Further examples of linkers are considered and described elsewhere herein.

[0200] In some cases, amino acid derivatives can be enzymatically modified amino acids. Enzymes used to produce enzymatically modified amino acids may include, for example, modified lyases, such as proteases, including edemases, cruzain, lysins (e.g., ClpS, ClpX), proteinase K, exopeptidases, aminopeptidases, diaminopeptidases, serine proteases, cysteine ​​proteases, threonine proteases, aspartic proteases, aspartic proteases, glutamate proteases, metalloproteinases, asparagine peptidase, pepsin, trypsin, trypsin, Lys-C, Glu-C, Asp-N, chymotrypsin, carboxypeptidases (e.g., carboxypeptidase A, carboxypeptidase B, carboxypeptidase Y), SUMO protease, elastase, papain, endopeptidase, protease, and TrypZean. ®Bromelain, collagenase, hyaluronidase, thermophilic protease, fig protease, keratinase, trypsinoids, fibroblast activator, enterokinase, chymotrypsinogen, chymotrypsin, clostridium protease, calpain, α-cleavage protease, proline-specific endopeptidase, furin, thrombin, subtilisin, gene enzyme, PCSK9, cathepsin, aminoacylproline dipeptidase, methionine aminopeptidase, cathepsin C, 1-cyclohexene-1-yl-boronate pinacol ester, pyroglutamate aminopeptidase, renin, kininogen, kallikrein, DPPIV / CD26, phorate oligopeptidase, prolyl oligopeptidase, leucine aminopeptidase, dipeptidyl peptidase or other enzymes or proteases, or combinations or variants thereof (e.g., engineered mutants or variants). Enzymes may be chemically modified and / or cleave amino acids to produce amino acid derivatives.

[0201] In some embodiments, the binding agent may be bound to a C-terminal amino acid or a derivative thereof. The C-terminal amino acid or its derivative may be linked to the peptide, or the C-terminal amino acid or derivative may be free or cleaved from the peptide. In some cases, C-terminal degradation methods are used, where the C-terminal amino acid comprises a free or cleaved amino acid. C-terminal degradation may include Edman-like degradation methods. C-terminal degradation may involve the use of an activating agent and a derivatizing agent (e.g., a thiocyanate that produces a peptide-thiohydantoin) that reacts with the C-terminal carboxyl group of the peptide. Non-limiting examples of activating agents include acetyl chloride and acetic anhydride. Alternatively or otherwise, a single-step C-terminal derivatization of the peptide can be performed to generate a peptide-thiohydantoin, for example, using the Schlack-Kumpf method, where the peptide reacts with thiocyanate (e.g., in acetone) to generate a peptide-thiohydantoin. The peptide-thiohydantoin may be cleaved (e.g., using alkaline conditions) to yield an amino acid thiohydantoin and the remaining peptide. In this example, the binder can recognize the amino acid thiohydantoin or its derivatives.

[0202] In some embodiments, the binder does not bind to a specific amino acid or its derivative. For example, under a given set of conditions, the binder may bind to a first amino acid type but not to a second amino acid type. In some cases, the binder binds to one or more amino acid types having similar side-chain groups; for example, the binder may bind to one or more nonpolar aliphatic side-chain amino acids (such as alanine, glycine, valine, leucine, methionine, and / or isoleucine). The binder may bind to one or more amino acids containing positively charged side chains (such as lysine, arginine, and / or histidine). The binder may bind to one or more amino acids containing negatively charged side chains (e.g., aspartic acid and glutamic acid). The binder may bind to one or more amino acids containing nonpolar aromatic side chains (such as phenylalanine, tyrosine, and / or tryptophan). The binder may bind to one or more amino acids containing polar, uncharged side chains (such as serine, threonine, cysteine, proline, asparagine, and / or glutamine). In some embodiments, the binder may bind to two or more amino acids with different side chain characteristics (e.g., amino acids with aliphatic side chains and amino acids with charged side chains).

[0203] Binders can bind to one or more amino acid types with varying specificities or affinities. In one example, under a given set of conditions, a binder (e.g., an antibody, antibody fragment) may bind to isoleucine and valine or derivatives thereof; however, when measured using an ELISA, the binder may bind to isoleucine with a greater affinity than the binder has for valine. In another example, under a given set of conditions, a binder may bind to phenylalanine, tyrosine, and tryptophan or derivatives thereof; however, under a given set of conditions, such as when measured using an ELISA with a specific concentration of the binder, the amino acid target, or both, the binder may preferentially bind to phenylalanine over tyrosine and tryptophan. It should be understood that the examples listed above are for illustrative purposes only and cover other binders that can bind differentially to different combinations of amino acids or their derivatives.

[0204] The binding properties of the binder to the target (e.g., a specific amino acid type) can be measured using any applicable technique, including but not limited to immunoassays such as Western blotting, direct or indirect enzyme-linked immunosorbent assay (ELISA), surface plasmon resonance (SPR), biofilm interferometry (BLI), fluorescence-activated cell sorting (FACS), imaging, or radioimmunoprecipitation. Non-limiting examples of measurable useful binding properties include affinity (e.g., equilibrium dissociation constant (K0)). d ) or the binding constant (K) a ), and kinetic rate constant (e.g., k) on or k off), specificity, affinity, and selectivity. In some examples, binding can be measured, for instance, by comparing the binding of the binder to a specific amino acid type or its derivative with the binding of the binder to a control molecule (which may be a structurally similar molecule). In some examples, specific binding can be determined using competitive assays (e.g., competitive ELISA).

[0205] The binding agents provided herein can bind to amino acids or their derivatives with relatively high affinity. The binding agents can bind to amino acids or their derivatives with dissociation constants of about 500 nanomolars (nM), about 100 nm, about 10 nm, about 1 nm, about 500 picomolars (pM), about 100 pM, about 10 pM, about 1 pM, or less, as determined, for example, by BLI or SPR. In some cases, the binding agents bind to amino acids or their derivatives with dissociation constants of up to about 500 nanomolars (nM), up to about 100 nm, up to about 10 nm, up to about 1 nm, up to about 500 picomolars (pM), up to about 100 pM, up to about 10 pM, up to about 1 pM, or less. The dissociation constants can fall within, for example, values ​​from about 100 pM to about 1 nm, from about 500 pM to about 900 pM, etc.

[0206] As described elsewhere herein, the conjugate may include a protein or peptide. In some cases, the conjugate comprises an antibody, an antibody fragment (e.g., Fab), a single-chain variable fragment (scFv), or a nanobody. In some embodiments, the conjugate comprises an antibody, such as immunoglobulin G (IgG). The antibody may be a monoclonal antibody or a polyclonal antibody (e.g., a group of antibodies). Monoclonal antibodies may be obtained from a polyclonal mixture of antibodies. The conjugate may be a purified conjugate (e.g., a purified antibody from animal serum or from a monoclonal hybridoma), or the conjugate may be recombinant. In some cases, the conjugate comprises scFv.

[0207] In some cases, binders are engineered using any suitable method, such as directed evolution methods, including but not limited to phage display, yeast display, yeast surface display, ribosome display, mammalian surface display, bacterial display, mRNA display, droplet-based directed evolution, and continuous evolution methods (e.g., OrthoRep). Display techniques may include or incorporate genetic diversification strategies (e.g., error-prone mutagenesis), as described elsewhere in this document. In some cases, binders are engineered using rational design or computational modeling, such as using RosettaDesign, PyMol, Schrodinger, MOE, and WAM, as described in Sevy and Meiler, 2018 Microbiology Spectrum.

[0208] Alternatively or otherwise, amino acid substitutions (which can be conserved or non-conserved) can be performed to identify key residues in the binder that bind to amino acids or their derivatives, thereby allowing for the rational design to alter the binding properties of the antibody (e.g., binding affinity, specificity, selectivity, dissociation rate, etc.). In one example of conserved amino acid substitution, a native amino acid residue can be substituted with another amino acid residue such that the polarity, charge, or other properties (e.g., hydrophobicity, hydrophilicity, isoelectric point) of the amino acid residue at that position has little effect. In another example, alanine scanning mutagenesis can be performed to identify residues crucial for binding to the binder (e.g., antibody or antibody fragment). Conserved substitutions can encompass non-naturally occurring amino acid residues that can be incorporated through chemical peptide synthesis or synthesis in biological systems (e.g., orthogonal translation, extension of the genetic code).

[0209] The binder may contain or be conjugated to a detectable tag (e.g., a fluorophore, fluorescent protein, radioisotope, mass tag, or recognizable polymerizable molecule (e.g., nucleic acid barcode molecule)). The detectable tag can be conjugated to the binder using any suitable method, such as protein engineering to produce fusion proteins containing a detectable protein (e.g., a fluorescent or chromogenic protein, such as green fluorescent protein (GFP), peroxidase), SNAP-tag, CLIP-tag, cysteine ​​tag, SpyCatcher, SpyTag; chemical conjugation strategies, such as using NHS esters, click chemistry; and affinity-based interactions, such as biotin-avidin or similar molecules, such as dethiobiotin, streptoavidin, neutral avidin, etc. Other conjugation strategies and exemplary connectors are provided elsewhere in this document.

[0210] In some embodiments, the binder comprises atypical amino acids (ncAAs) that can be used to couple molecules (such as detectable tags) to the binder. The binder can be engineered to contain ncAAs using any useful technique, such as site-specific incorporation of atypical amino acids into bacterial or eukaryotic systems using extensions to the biological genetic code, for example, using archaeal tRNA and aminoacyl-tRNA synthases; genome engineering or recoding (e.g., adding or removing codons from the genetic code); engineered ribosomes; molecular evolutionary approaches, such as those described in Arranz-Gibert et al., 2018. Curr Opinion Chem. Bio., 46; 203-211, and reviewed in De la Torre and Chin, 2021. Nature Reviews Genetics, 22, 169-184, each of which is incorporated herein by reference in its entirety. In some cases, atypical amino acids include click chemistry motifs (e.g., azides, DBCO, BCN, TCO, tetrazine, or other click chemistry motifs) that can facilitate the conjugation of detectable tags via complementary click chemistry (e.g., azide-DBCO click reaction).

[0211] Binder Library: Binders may be provided as part of a binder library that can collectively recognize and specifically or semi-specifically bind to any number of protein amino acid types or their derivatives under a given set of conditions. In some embodiments, the binder library comprises multiple binders that can recognize and specifically bind to at least 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or all 20 different types of protein amino acids. The binder library may contain any number of binders to recognize any number of amino acid types or their derivatives. The binder library may contain binders with single specificity or specificity to groups of amino acid types. In one example, the library may contain a first binding agent specific to a single amino acid type, while a second binding agent may be specific to more than one amino acid type (e.g., it may bind to 2, 3, 4, or 5 amino acids). In some cases, fewer than 20 binding agents are needed to identify all 20 protein amino acids or their derivatives. In other cases, more than 20 binding agents may be available to identify all 20 protein amino acids or their derivatives. In one such example, multiple queries can be made to a single amino acid or its derivative using one or more binding agents that identify at least one type of amino acid or its derivative, which can improve the accuracy of amino acid identification.

[0212] A binding library can contain different types of binding agents. For example, the first binding agent may include an antibody, the second binding agent may include an antibody fragment, and the third binding agent may include a tRNA synthetase. Alternatively, all binding agents in the library may be of the same type (e.g., all antibodies, all antibody fragments, all engineered tRNA synthetases, all engineered ClpS or ClpX proteins, etc.).

[0213] Advantageously, using a binding agent library can increase the efficiency and specificity of some binding agents within the library. For example, two or more binding agents within the library can bind differentially to a target amino acid or its derivative, and thus can competitively bind to the target amino acid, thereby reducing non-specific binding. Similarly, binding agents with different levels of specificity can be used in a specific order or sequence, which can help make binding agents with lower specificity more specific simply based on the sequence they provide. For example, a first binding agent may be able to bind specifically to a first amino acid or its derivative, and a second binding agent may be able to bind to both the first amino acid or its derivative and the second amino acid or its derivative. A first binding agent can be provided and brought into contact with the first amino acid or its derivative and the second amino acid or its derivative. Because the first binding agent is specific to the first amino acid or its derivative, it will bind only to the first amino acid or its derivative. Subsequently, a second binding agent can be provided; however, because the first amino acid or its derivative binds to the first binding agent, the first amino acid or its derivative may be inaccessible to the second binding agent (e.g., spatially closed). Therefore, the second binding agent can bind only to the second amino acid or its derivative. Therefore, the identification of the first and second binders (e.g., by detecting a marker / tag or by sequencing polymerizable molecules coupled to the binder and optionally transferred to the substrate) can allow for the identification of the first amino acid or a derivative thereof, as well as the second amino acid or a derivative thereof.

[0214] Binders can be provided at any useful concentration. Binders can be provided at concentrations of about 1 pg / mL, about 10 pg / mL, about 100 pg / mL, about 1 μg / mL, about 10 μg / mL, about 100 μg / mL, about 1 mg / mL, or higher. In some cases, binders can be provided at concentrations of up to about 1 mg / mL, up to about 100 μg / mL, up to about 10 μg / mL, up to about 1 μg / mL, up to about 0.1 μg / mL, or lower. Binders in the binder library can be provided at the same or different concentrations. For example, each binder can be titrated to its homologous molecule (e.g., amino acid type or derivative thereof) to achieve optimal specificity.

[0215] The binder may contain or be coupled with any useful number of functional portions, as described elsewhere herein. For example, the binder may have 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100 or more functional portions attached thereto.

[0216] This document also provides kits and compositions that may contain multiple binding agents. The kits or compositions disclosed herein may contain multiple binding agents that specifically or semi-specifically bind to multiple different types of protein amino acids or their derivatives. In some embodiments, the multiple binding agents may bind to 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or all 20 types of protein amino acids or their derivatives. In some embodiments, the multiple binding agents may bind to at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 11, at least 12, at least 13, at least 14, at least 15, at least 16, at least 17, at least 18, at least 19, or all 20 different types of protein amino acids or their derivatives. In some embodiments, the binding agent comprises an antibody or an antibody fragment.

[0217] This article also provides a heavy chain variable region (V) containing three complementarity-determining regions (CDRs). H ) and light chain variable regions containing three CDRs (V L Antibodies or antibody fragments, wherein V H The CDR #1 sequence has at least 80% homology with the sequences selected from Table 1, where V H The CDR #2 sequence has at least 80% homology with the sequences selected from Table 2, where V H The CDR #3 sequence has at least 80% homology with the sequences selected from Table 3, where V L The CDR #1 sequence has at least 80% homology with the sequences selected from Table 4, where V L The CDR #2 sequence has at least 80% homology with the sequences selected from Table 5, and where V L The CDR #3 sequence is at least 80% homologous to the sequences selected from Table 6.

[0218] Table 1

[0219]

[0220]

[0221] Table 2

[0222]

[0223] Table 3

[0224]

[0225] Table 4

[0226]

[0227] Table 5

[0228]

[0229] Table 6

[0230]

[0231]

[0232] In some implementations, the heavy chain sequence has at least 80% homology with the sequences selected from Table 7.

[0233] Table 7

[0234]

[0235]

[0236]

[0237]

[0238] In some implementations, the light chain sequence has at least 80% homology with the sequences selected from Table 8.

[0239] Table 8

[0240]

[0241]

[0242] In some implementation schemes, V H The sequence has at least 80% homology with the sequences selected from Table 9.

[0243] Table 9

[0244]

[0245]

[0246]

[0247] In some implementation schemes, V L The sequence has at least 80% homology with the sequences selected from Table 10.

[0248] Table 10

[0249]

[0250]

[0251] In some aspects of this disclosure, a binding agent is provided that selectively binds to tryptophan or a derivative thereof, rather than selectively binding to all other types of protein amino acids or their derivatives.

[0252] In some aspects of this disclosure, a binding agent is provided that selectively binds to phenylalanine or a derivative thereof, rather than selectively binding to all other types of protein amino acids or derivatives thereof.

[0253] In some aspects of this disclosure, a binding agent is provided that selectively binds to leucine or a derivative thereof, rather than selectively binding to all other types of protein amino acids or derivatives thereof.

[0254] In some aspects of this disclosure, a binding agent is provided that selectively binds to isoleucine or a derivative thereof, rather than selectively binding to all other types of protein amino acids or derivatives thereof.

[0255] In some aspects of this disclosure, a binding agent is provided that selectively binds to valine or a derivative thereof, rather than selectively binding to all other types of protein amino acids or their derivatives.

[0256] In some aspects of this disclosure, a binding agent is provided that selectively binds to tyrosine or a derivative thereof, rather than selectively binding to all other types of protein amino acids or their derivatives.

[0257] In some aspects of this disclosure, a binding agent is provided that selectively binds to proline or a derivative thereof, rather than selectively binding to all other types of protein amino acids or derivatives thereof.

[0258] Methods and compositions for producing amino acid binders

[0259] In some aspects, this document also provides methods for generating the binders described herein that identify the type of a single monomer (e.g., an amino acid). In one aspect, the binders described herein are generated using an immunoassay method. In some embodiments, the binder comprises an antibody or an antibody fragment. In some embodiments, the binder is generated by a method comprising: immunizing an animal with an immunogenically effective amount of a composition comprising an amino acid or a derivative thereof and a polymeric linker molecule, and obtaining the binder; optionally, obtaining the binder may include further treatment or engineering of the binder. The composition for immunizing animals may also comprise a carrier molecule conjugated to the polymeric linker molecule and the amino acid or a derivative thereof.

[0260] In one example, the binder may be generated using a method comprising: immunizing an animal with an immunogenically effective amount of a composition comprising a carrier molecule conjugated to (i) a polymeric linker molecule and (ii) an amino acid or a derivative thereof, and obtaining the binder from the animal. The binder may include antibodies, such as immunoglobulin G (IgG). In some cases, the binder may be further processed or engineered to produce an amino acid binder capable of recognizing or binding to single or multiple amino acid types. Examples of processing or engineering antibodies or antibody fragments (e.g., single-chain variable fragments (scFv)) include hybridoma techniques for generating monoclonal antibodies, or cloning and directed evolution methods such as phage display and phage panning, yeast display, yeast surface display, mammalian display, bacterial display, ribosome display, mRNA display, droplet-based directed evolution and continuous evolution methods, such as OrthoRep, phage-assisted continuous evolution (PACE), or periplasmic PACE (pPACE). In some cases, directed evolution methods may include attaching one or more sequences of an antibody or antibody fragment (e.g., V... H and V L The sequence is cloned into an expression vector, expressed, purified, and the binding is measured in host cells (e.g., bacteria, yeast, mammalian cells).

[0261] Polymer adapter: Compositions for immunizing animals may contain a carrier molecule coupled to at least one polymer adapter. The polymer adapter may contain any useful polymer. Polymers can be naturally occurring or synthetic, such as acrylic acid, nylon, silicone, viscose, rayon, polyester, polycarboxylic acid, polyvinyl acetate, polyacrylamide, polyacrylate, polyethylene glycol (PEG), polyurethane, polylactic acid, polystyrene, polyacrylonitrile, polybutadiene, polycarbonate, poly(chlorotrifluoroethylene), poly(ethylene oxide), poly(ethylene terephthalate), polyethylene, polyisobutylene, poly(methyl methacrylate), polyoxymethylene, polyoxymethylene, polypropylene, polystyrene, poly(tetrafluoroethylene), poly(vinyl alcohol), polyvinyl chloride, polyvinylidene chloride, polyvinylidene fluoride, polyvinyl fluoride, poly-L-lysine (PLL), poly(DL-lactic acid) (PLA), poly(DL-lactide-co-glycoside) (PLGA), polycaprolactone (PCL), polyvinyl alcohol (PVA), polyethylene terephthalate (PET), polystyrene (PSt), polyornithine, polyarginine, or other useful polymers. Polymers can include biopolymers, such as polypeptides or nucleic acid molecules, such as DNA molecules, RNA molecules, or DNA:RNA hybrids.

[0262] Polymer connectors may contain one or more functional groups that allow the polymer connector to conjugate with amino acids or their derivatives and / or carrier molecules. Examples of useful functional parts include, but are not limited to, click chemistry parts such as azides, alkynes and cycloalkynes (such as BCN and DBCO), nitrones, alkenes (e.g., strained alkenes), tetrazides, methyltetrazides, triazoles, tetrazolium, phosphites, and phosphines; photoreactive parts such as phenyl azides, benzophenone, psoralen, and diazacyclopropenes; carboxyl reactive crosslinking groups; amino reactive crosslinking groups such as NHS esters; EDC groups, maleimides, thiols, cystamines, aldehydes, succinimides, epoxides, acrylates, or other crosslinking or reactive groups. Further examples of connectors and functional parts are provided elsewhere herein. Polymer connectors can contain any useful number of functional groups, such as 1, 2, 3, 4, 5, 6, 7, 8, 9, 10 or more functional groups, which can be the same or different functional groups.

[0263] Polymer linkers can contain any useful number of monomers or have any desired length or molecular weight, for example, to control the molecular distance between the carrier protein and the amino acid or its derivative. Polymer linkers can contain about 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100 or more monomers. Polymer linkers can contain at least about 2, at least about 3, at least about 4, at least about 5, at least about 6, at least about 7, at least about 8, at least about 9, at least about 10, at least about 20, at least about 30, at least about 40, at least about 50, at least about 60, at least about 70, at least about 80, at least about 90, at least about 100 or more monomers. Alternatively, the polymer connector may contain up to about 100, up to about 50, up to about 10, up to about 5 or fewer monomers.

[0264] Under a given set of conditions, polymer joints can contain any useful length. For example, the length of polymer joints in aqueous solutions can be about 0.01 nm, about 0.1 nm, about 0.2 nm, about 0.3 nm, about 0.4 nm, about 0.5 nm, about 0.6 nm, about 0.7 nm, about 0.8 nm, about 0.9 nm, about 1 nm, about 2 nm, about 3 nm, about 4 nm, about 5 nm, about 6 nm, about 7 nm, about 8 nm, about 9 nm, about 10 nm, about 20 nm, about 30 nm, about 40 nm, about 50 nm, about 60 nm, about 70 nm, about 80 nm, about 90 nm, about 100 nm, or greater. The length of the polymer connector can be at least about 0.01 nm, at least about 0.1 nm, at least about 0.2 nm, at least about 0.3 nm, at least about 0.4 nm, at least about 0.5 nm, at least about 0.6 nm, at least about 0.7 nm, at least about 0.8 nm, at least about 0.9 nm, at least about 1 nm, at least about 2 nm, at least about 3 nm, at least about 4 nm, at least about 5 nm, at least about 6 nm, at least about 7 nm, at least about 8 nm, at least about 9 nm, at least about 10 nm, at least about 20 nm, at least about 30 nm, at least about 40 nm, at least about 50 nm, at least about 60 nm, at least about 70 nm, at least about 80 nm, at least about 90 nm, at least about 100 nm or greater. Alternatively, the length of the polymer connector may be up to about 100 nm, up to about 90 nm, up to about 80 nm, up to about 70 nm, up to about 60 nm, up to about 50 nm, up to about 40 nm, up to about 30 nm, up to about 20 nm, up to about 10 nm, up to about 5 nm, up to about 1 nm or less.

[0265] Carrier molecules: Carrier molecules may include any useful molecules that promote an animal’s immune response to the composition. Carrier molecules may include nanoparticles (e.g., gold nanoparticles), liposomes, vesicles, or biomolecules such as proteins. In non-limiting examples, examples of carrier molecules include keyhole hemocyanin (KLH), Chilean abalone hemocyanin (CCH), Qβ virus-like particles, outer membrane protein complex (OMPC), genetically modified cross-reactive substance of diphtheria toxin (CRM), tetanus toxoid (TT), diphtheria toxoid (DT), Haemophilus influenzae protein D (HiD), glycoengineered outer membrane vesicles (geOMV), recombinant Hla, rEPA, and recombinant pneumococcal hemolysin, etc., as reviewed, for example, in Micolia et al., 2018. Molecules. 23(6): 1451, which is incorporated herein by reference. Other examples of carrier proteins include albumins (e.g., bovine serum albumin, cationic bovine serum albumin, ovalbumin, cationic ovalbumin), biotin, streptavidin, hemocyanin carrier proteins, and Blue Carrier. ™ protein.

[0266] As described elsewhere in this document, binders can be configured to recognize and bind to amino acid derivatives, including chemically modified amino acids. In some cases, the amino acid derivatives include amino acids that have already been modified by reacting with a linker. In one example, the linker contains an amino acid reactive group (e.g., an isothiocyanate, such as PITC), which can react with the N-terminal amino acid of the peptide, for example, under basic conditions. The reaction of PITC with the N-terminal amino acid produces a PTC-derived amino acid, which can further undergo cleavage, for example, using acidic and / or thermal conditions, to produce an ATZ-derived amino acid without peptide cleavage. In some cases, the ATZ-derived amino acid can undergo further derivatization conditions, for example, to produce PTH or PTC-derived amino acids.

[0267] Therefore, the methods for generating binders provided herein include providing compositions comprising amino acid derivatives. The compositions may comprise a carrier molecule, such as a carrier protein, coupled to a polymeric linker (e.g., a bifunctional linker comprising a polymeric region) and a derivatized amino acid (e.g., an amino acid-linker complex). In a particular example, the derivatized amino acid may be generated by reacting a first bifunctional linker comprising a PITC moiety and a click chemical moiety (e.g., an azide), as described elsewhere herein, to generate an amino acid-linker complex comprising a PTC-derived amino acid and a click chemical moiety. The click chemical moiety (e.g., an azide) of the amino acid-linker complex may be used to couple the amino acid-linker complex to a carrier molecule (e.g., a carrier protein) (e.g., a carrier molecule comprising a click chemical moiety), or via a polymeric linker comprising a click chemical moiety (e.g., a DBCO). In some examples, the carrier molecule comprises bovine serum albumin (BSA) protein conjugated to a PEG polymeric linker comprising a click chemical moiety (e.g., a DBCO). PEG linkers can be coupled to BSA using, for example, an NHS ester reaction with an amine group of BSA (e.g., an amine group of a lysine residue) (e.g., NHS-PEG-DBCO). The click chemistry of the PEG polymer linker can be linked to the click chemistry of an amino acid-linker complex (e.g., a DBCO-azide click reaction), thereby producing a BSA-PEG-derived amino acid complex. In some examples, the amino acid does not need to be derivatized before being conjugated to the carrier molecule or polymer linker.

[0268] The carrier molecule can be coupled to any useful number of amino acid molecules or their derivatives, for example, to generate an immune response in animals. The carrier molecule can be coupled to about 1, about 2, about 3, about 4, about 5, about 6, about 7, about 8, about 9, about 10, about 11, about 12, about 13, about 14, about 15, about 16, about 17, about 18, about 19, about 20, about 30, about 40, about 50, about 60, about 70, about 80, about 90, about 100 or more amino acid molecules or their derivatives. The carrier molecule may be coupled to at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 15, at least 20, at least 30, at least 40, at least 50, at least 60, at least 70, at least 80, at least 90, at least 100 or more amino acid molecules or their derivatives.

[0269] Similarly, any number of polymeric linkers can be used to couple an amino acid or a derivative thereof to a carrier molecule. The carrier molecule can be coupled with about 1, about 2, about 3, about 4, about 5, about 6, about 7, about 8, about 9, about 10, about 11, about 12, about 13, about 14, about 15, about 16, about 17, about 18, about 19, about 20, about 30, about 40, about 50, about 60, about 70, about 80, about 90, about 100, or more polymeric linkers. The polymeric linkers can be provided at any useful ratio to the amino acid or a derivative thereof. For example, the polymeric linkers can be provided at a ratio of about 1:1, 1:2, 1:3, 1:4, 1:5, 1:6, 1:7, 1:8, 1:9, 1:10, 1:50, 1:100, 1:1000, 1:10,000, or less to the amount of the amino acid or a derivative thereof.

[0270] The compositions described herein cover all 20 protein amino acid types, as well as modified amino acids (e.g., amino acid derivatives, post-translational modified amino acids). Individual compositions for each amino acid type or subgroup of amino acid types can be provided. The methods described herein for generating binders using immunomodulatory approaches may include using one or more animals for each amino acid type in each composition. For example, a first composition comprising a carrier molecule, a polymer linker, and one of the 20 amino acid types can be generated. A second composition comprising a carrier molecule, a polymer linker, and another of the 20 amino acid types different from the amino acid type in the first composition can be generated. The first and second compositions can be injected separately into animals to generate unique binders (e.g., antibodies) against a target antigen (a specific amino acid type).

[0271] The compositions described herein may additionally contain or be provided with adjuvants, including but not limited to aluminum hydroxide, complete Freund's adjuvant (CFA or FCA), incomplete Freund's adjuvant (IFA or FIA), Ribi adjuvant, Titermax, Spector, or others.

[0272] In some embodiments, the method for generating the binder includes obtaining a polyclonal mixture of antibodies from an animal, the polyclonal mixture containing a variety of antibodies that can recognize and preferentially bind to a type of protein amino acid or a derivative thereof (e.g., the type of amino acid the animal is immunized against).

[0273] In some cases, the binding agent is a monoclonal antibody. Monoclonal antibodies can be produced using any useful technique, such as hybridoma technology. In some examples, monoclonal antibodies are produced by immunizing an animal with a composition containing a target antigen (e.g., an amino acid or an amino acid derivative), recovering lymphocytes (e.g., B cells) from the animal, and fusing the recovered cells with an immortalized cell line (e.g., a myeloid-like cell line, such as myeloma cells) to generate a hybridoma. The hybridoma cell line can then be screened and selected to identify individual hybridoma cells that produce antibodies specific to the target antigen. Antibodies can be harvested directly from the cells (e.g., from a culture medium). In some cases, the antibody-producing hybridoma can be sequenced, and recombinant antibodies or antibody fragments (e.g., scFv) can be produced using DNA or RNA sequences (e.g., RNA sequences obtained from hybridoma cells, reverse transcribed, and sequenced).

[0274] Alternatively or otherwise, the binder may contain a recombinant antibody or antibody fragment generated using genetic recombination. For example, monoclonal cells (e.g., B cells harvested from an animal) or hybridoma cells that produce a specific antibody may be genome- or transcriptome-sequencing to obtain one or more DNA or RNA sequences corresponding to the antibody or a portion thereof (e.g., a variable heavy chain region, a variable light chain region, etc.). The DNA or RNA sequence can then be used to generate the recombinant antibody (e.g., through translation of the DNA or RNA sequence, conversion or transduction of the gene in a host cell, or protein synthesis of the corresponding amino acid sequence). In some cases, recombinant antibody fragments (e.g., scFv) are generated by isolating RNA from B cells of an immunized animal and reverse transcribing the RNA to obtain V. H Sequence and V L Sequence, and recombine V in the vector H Sequence and V L Sequence. V H Sequence and V LThe sequence can be introduced into phage vectors, yeast vectors, mammalian cell vectors, plasmids, etc., and transformed into vectors or host cells (e.g., yeast, mammals, bacteria, bacteriophages) to generate recombinant antibody fragments. Bioconjugation of the binder can be performed to attach useful functional parts to the binder. For example, bioconjugation methods can allow the attachment of detectable markers (such as nucleic acid barcodes, fluorophores, proteins (e.g., fluorescent proteins, antibodies, or antibody fragments), mass tags, etc.) to the binder. Non-limiting examples of biological conjugation strategies include chemical strategies such as the use of silanes, such as aminosilanes (e.g., APTES), amino-PEG-silanes; oxidation of sugars and reaction with amines; PLP transamination and reaction with alkoxyamines, aminooxy groups, hydroxylamines, or hydrazines; reactions of aldehydes with organoboroesters; reactions of o-aminophenols with catechols; reactive esters (e.g., NHS), sulfonyl chlorides, isocyanates, isothiocyanates, aldehydes, sulfonyl acrylates, maleimides, haloacetamides, alkynyl carboxylic acid derivatives, 5-methylenepyrrolidone, 5,5'-dithiobis(2-nitrobenzoic acid), cyclooctyne, and diazo. Salts, in-situ generated imines, 4-phenyl-3H-1,2,4-triazol-3,5(4H)-diones, phenthiazides, 9-azabicyclo[3.3.1]non-3-one-N-oxy, N-substituted pyridinium salts, oxazolidinyl propane, high-valent iodine, luminescent flavonoids, phosphorus (V) reagents and DBU, tetrazolium derivatives, aldehydes and isocyanates, H-aziridine-based reagents, dicarbonyl compounds, p-azidophenylglyoxal hydrate, 2-cyclohexenone, thiophosphoalkynyl dichloroester, sulfinates, photocatalytic reactions with C4-alkyl-1,4-dihydropyridine reagents, or other reactive moieties. Examples of enzymatic methods for biological conjugation include: sorting enzyme A, microbial transglutaminase, tyrosinase (e.g., tyrosinase-mediated oxidative coupling), farnesyltransferase, N-myristyltransferase, phosphopanylthioethylaminotransferase, tubulin tyrosine ligase, lipoic acid ligase, biotin ligase, formylglycine synthase, SnoopTag / SnoopCatcher, SnoopLigase, SpyTag / SpyCatcher, etc.

[0275] In some cases, biological conjugation can be achieved using extensions of the genetic code, as described elsewhere herein. Non-limiting examples of atypical amino acids include the incorporation of click chemistry motifs that can mediate the conjugation of additional molecules via click chemistry reactions, such as CuAAC, SPAAC, IEDDA, ​​and other click reactions. Further examples of conjugation methods (e.g., for peptides with substrates) are provided elsewhere herein and are applicable to the biological conjugation of useful molecules with binding agents.

[0276] Any useful number of functional portions may be attached to the binder. The binder may have 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100 or more functional portions attached thereto. The binder may contain at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 20, at least 30, at least 40, at least 50, at least 60, at least 70, at least 80, at least 90, at least 100 or more functional portions attached thereto. If multiple binders are used, each binder may contain the same or different numbers of functional portions attached thereto.

[0277] In some cases, methods for generating binders also include engineering precursor binders to obtain binders that have improved binding properties (e.g., affinity, specificity) for their target antigens (e.g., amino acid types). Examples of engineered antibodies include directed evolution methods, which may include deriving antibodies into antibody fragments (such as scFv), and / or using any number of useful evolutionary methods, such as phage display and phage panning, yeast display, yeast surface display, mammalian display, mammalian surface display, bacterial display, ribosome display, mRNA display, etc. To introduce additional genetic diversity (e.g., additional mutations for generating binder sequences), various techniques can be employed, such as error-prone mutagenesis (e.g., error-prone PCR, error-prone rolling circle amplification (epRCA)), random insertion-deletion strand exchange (RAISE), tandem repeat insertion (TRINS), micromu transposon-based techniques, DNA shuffling, SELEX, staggered amplification process (StEP), random chimerism on transient templates (RACHITT), incremental truncation (ITCHY) and SCRATCHY for generating hybrid enzymes, and sequence homology-independent protein recombination (SHIPR). EC, sequence-independent site-specific chimerism (SISDC), Kunkel mutagenesis, mutagenesis-mediated tissue recombination via homologous in vivo grouping (MORPHING), saturation mutagenesis, site-specific saturation mutagenesis, stepwise loop insertion strategy (StLois), FACS-enabled high-throughput screening, plate-based automated enzyme assays, MS-based methods, QUEST, cofactor regeneration coupling, in vitro compartmentalized self-replication, multiplex automated genome engineering (MAGE), CRISPR-based mutagenesis, phage-assisted continuous evolution (PACE), periplasmic PACE, etc., as reviewed by Vidal et al., 2023. RSC Chem. Biol. 4, 271-291. In some cases, genetic diversity can be increased by using degenerate codons for cloning or amplification, DNA shuffling, CDR recombination, or other methods.

[0278] Polymer analytes were analyzed by localized tethering.

[0279] This disclosure also provides methods for processing and analyzing polymer analytes (e.g., peptides, polymers, nucleic acid molecules, etc.) in a highly parallel and accurate manner. The systems and methods of this disclosure may include cleaving monomers from the polymer analyte, coupling monomers to a capture moiety (e.g., by local tethering), and detecting the monomers. In some embodiments, detection includes using a monomer-specific binder to identify and bind to the cleaved monomers. In some embodiments, the monomer-specific binder is used for direct or indirect detection; for example, the binder may contain a directly detectable marker (e.g., a fluorophore, mass tag, radioisotope, fluorescent protein), or the binder may contain a polymerizable molecule with encoded information that can be transferred by coupling or copying the encoded information to the capture moiety or another polymerizable molecule. In some cases, the additional polymerizable molecule or capture moiety is located in proximity to the polymer analyte. Alternatively or additionally, the binder may be used to sort the cleaved monomers, for example, into separate compartments or chambers for downstream labeling, such as with identification barcode molecules. These operations may be iterated or repeated any number of times to obtain information about all or subgroups of monomers of the polymer analyte and optionally the sequence of the monomers relative to the polymer analyte. Information can be read from polymerizable molecules using methods such as conventional next-generation sequencing or nanopore sequencing. Advantageously, monomers are removed from adjacent monomers by cleaving them from the polymer analyte, and the binder can bind specifically to each monomer without being affected by surrounding monomers. Therefore, the method disclosed herein enables more accurate molecular identification and polymer sequencing, which has applications in disease diagnosis, monitoring protein dynamics or protein interactions, single-cell proteomics, and the development or characterization of therapeutics.

[0280] In some embodiments, the methods provided herein include intramolecular expansion of a polymer analyte (e.g., a peptide) to generate a modified monomer (e.g., a modified amino acid). This intramolecular expansion process may include providing a linker, coupling the linker to a monomer of the polymer analyte to generate a monomer-linker complex, coupling the linker or monomer-linker complex to a capture moiety, and cleaving the monomer from the polymer analyte to generate the modified monomer. In some cases, the method may also include repeating intramolecular expansion on a next monomer of the polymer analyte, or repeating one or more operations of intramolecular expansion to generate another modified monomer. In some cases, the monomer-linker complex or modified monomer generated by a round or cycle of intramolecular expansion may be coupled to the monomer-linker complex or modified monomer from a previous round or cycle of intramolecular expansion to generate a stack of multiple modified monomers, for example, linked by a polymerizable molecular backbone. In some cases, the modified monomer or the stack of multiple modified monomers is detected using a nanopore sequencer to output the identity of each monomer of the polymer analyte.

[0281] One or more methods in this disclosure may employ a connector capable of coupling to (i) a monomer of a polymer analyte and (ii) a trapping portion, which, once the monomer is cleaved, can be used to locally tether the monomer to the vicinity of the polymer analyte or directly attach it to the polymer analyte. In some cases, the methods and systems disclosed herein may additionally include a substrate for local tethering; the substrate may be coupled, for example, to the polymer analyte, the trapping portion, and additional polymerizable molecules. In some cases, the additional polymerizable molecules are encoded with information from polymerizable molecules derived from a binder. In other cases, the binder contains a detectable marker that can indicate a binding event between the binder and the monomer. Alternatively or additionally, one or more methods described herein may be carried out in solution or in the absence of a substrate.

[0282] The method disclosed herein for processing a polymer analyte comprising multiple monomers may include cleaving the monomers of the polymer analyte and coupling the monomers to a capture fraction (e.g., to a substrate, the polymer analyte, or provided in solution form) for subsequent processing or analysis. In one example, the method of this disclosure may include: providing (i) a polymer analyte comprising multiple monomers and (ii) a capture fraction; coupling one of the monomers to the capture fraction to produce a monomer-capture fraction complex; cleaving the monomer; contacting the cleaved monomer-capture fraction complex with a binder; and coupling a first polymerizable molecule to a second polymerizable molecule or to the capture fraction. In some cases, the first polymerizable molecule is coupled to a binder and contains information about the binder, such as the identity of the binder or its homologs, which may be transferred to the second polymerizable molecule or the capture fraction. In some cases, the capture fraction may be a copy or an identical molecule to the second polymerizable molecule. Alternatively or additionally, the binder may be used to sort mixtures of cleaved monomers by identity or type, and after sorting, identification tags or barcodes identifying monomer types may be coupled to the capture fraction. Example methods and systems of this kind are described in U.S. Patent No. 11,499,979, International Publication No. WO / 2023 / 114732, International Patent Application Publication No. WO / 2023 / 196642, and U.S. Patent Application No. 18 / 740,088 (filed June 11, 2024), and L. Zheng et al., 2024. Peptide sequencing via reverse translation of peptides into DNA. BioRxiv., each of which is incorporated herein by reference in its entirety.

[0283] In some implementations, nucleases are used to mediate the coupling of barcode molecules with polymeric analytes, capture fractions, or polymerizable molecules. The nuclease can be an editing endonuclease, such as a CRISPR-associated (Cas) enzyme or a guide editor enzyme. Cas enzymes or guide editor enzymes can include any useful Cas enzyme, such as DNA-targeting Cas enzymes (e.g., Cas9, Cas3, Cas12a, Cas14), or RNA-targeting Cas enzymes (e.g., Cas13 or Cas7-11), or modified variants thereof. Guide editor enzymes can include polymerases, such as polymerases like DNA polymerase, RNA polymerase, reverse transcriptase, or modified variants thereof. In some cases, guide editor enzymes include modified Cas nickases (e.g., modified Cas9 nickases) and modified reverse transcriptases. Advantageously, the use of guide editor enzymes can enable accurate polymerase sequencing because it allows for the efficient barcoding of individual monomers without generating double-strand DNA breaks and also enables efficient stacking of barcode molecules.

[0284] In some embodiments, the polymerizable molecule comprises a tandem array of nuclease target sites (e.g., CRISPR-Cas9 target sites, restriction sites, transposition sites, etc.). In one such example, the tandem array comprises multiple CRISPR-Cas9 target sites, and the barcode comprises a guide editing RNA (pegRNA). The coupling or copying of the barcode molecule to the polymerizable molecule can be mediated using a DNA typewriter approach, for example, as described in J. Choi et al., 2022. Nature. 608, pp. 98–107, the full text of which is incorporated herein by reference. In this approach, the polymerizable molecule may comprise a tandem array of nuclease recognition sites (e.g., CRISPR-Cas target sites, such as CRISPR-Cas9 target sites). The tandem array of nuclease recognition sites may, for example, comprise multiple CRISPR-Cas9 target sites (“target sites”) positioned adjacent to each other. The first target site among the multiple target sites may be active, and the other target sites among the multiple target sites may be inactive. For example, multiple target sites, in addition to the first target site, can be truncated so that the target sites are inactive or not recognized by the required nuclease (e.g., Cas9 or a guide editor enzyme including a modified Cas9). For example, multiple target sites, in addition to the first target site, can be truncated at the 5' end of the target site. The barcode molecule can include a pegRNA molecule containing a key sequence or its complement to complete or activate one of the target sites. For example, the pegRNA molecule can contain a key sequence or its complement to activate a second target site among multiple target sites located near the first target site. The pegRNA molecule can additionally contain a barcode sequence (or its complement). In such an example, Cas or a guide editor enzyme can be used to couple the pegRNA or its complement to the first target site, thereby inserting a barcode sequence and a key sequence near the first target site, thereby inactivating or deactivating the first target site while activating the second target site. Such a method can be mediated using a guide editor enzyme that includes a Cas protein (e.g., a Cas enzyme or a Cas nickase) and a polymerase (e.g., a reverse transcriptase). For example, in guided editing, the Cas protein can cleave the first strand of a polymerizable molecule, and a portion of the pegRNA (e.g., a primer binding site) can hybridize with a portion of the cleaved strand. An extension reaction (e.g., reverse transcription) can then produce a complementary DNA segment that is complementary to at least a portion of the remaining pegRNA. This complementary DNA segment can then be incorporated into the polymerizable molecule by hybridization with the second strand. Additional operations, such as flanking cleavage and ligation, can be performed. In some cases, guided editors include a Cas nickase that can increase the efficiency of cDNA segment incorporation into the polymerizable molecule.Additional barcoding operations (e.g., in an iterative barcoding process or throughout a workflow for processing additional monomers of a polymer analyte) can subsequently inactivate the second target site while activating a third target site adjacent to the second target site, and so on.

[0285] In some cases, barcode molecules include pegRNA molecules used to edit polymerizable molecules, for example, without inserting additional nucleotides. In one such example, the barcode molecule may contain a key sequence and a barcode sequence that replaces the sequence at the first target site, causing inactivation of the first target site and activation of the second target site. In such examples, the polymerizable molecule is edited to contain the barcode sequence, rather than being elongated by inserting additional nucleotides or sequences.

[0286] Alternatively or otherwise, the peCHYRON method can be used to conjugate polymerizable molecules with barcode molecules, for example, as described in T. Loveless et al., 2021. BioRxiv. 11.05.467507, the full text of which is incorporated herein by reference. In this method, the polymerizable molecule may contain a prototype spacer adjacent motif (PAM) and a first nuclease target site, such as a CRISPR-Cas9 target site (“target site”). In this example, the barcode molecule may include a pegRNA molecule containing a barcode sequence (or its complement) and a propagation sequence containing a second target site (or its complement). A Cas or guide editor enzyme can be used to conjugate the pegRNA to the first target site (e.g., between the target site and the PAM), thereby inserting the barcode sequence and the propagation sequence (or its complement) near the first target site. This insertion can induce inactivation of the first target site and activation of the second target site, and elongate the polymerizable molecule. Additional barcoding operations (e.g., in an iterative barcoding process or throughout a workflow for processing additional monomers of polymer analytes) may subsequently add additional target sites to the first and third target sites, such as a third target site, a fourth target site, etc.

[0287] Figure 1A An example workflow for sequencing polymer analytes is illustrated. In workflow 100, a polymer analyte 103 is provided and sequentially broken down into individual monomers through contact with a capture moiety and subsequent cleavage. The individual cleaved monomers, recombined with the capture moiety, are then contacted with a binder containing a polymerizable molecule (e.g., a barcode molecule) that identifies or encodes the binder. The polymerizable molecule (e.g., the barcode molecule) of the binder is coupled to another polymerizable molecule, and the monomer is either cleaved from the capture moiety or blocked to prevent recognition downstream by another binder. Figure 1AIn Figure A, substrate 101 is coupled to polymeric analyte 103 (e.g., a peptide to be sequenced), capture moiety 105 (e.g., a nucleic acid primer), and additional polymerizable molecule 107 (e.g., another nucleic acid primer). In some cases, capture moiety 105 and additional polymerizable molecule 107 are of the same molecular type (e.g., both are nucleic acid molecules containing the same sequence). The polymeric analyte may be contacted with a bifunctional adapter 109 comprising a terminal monomer coupling group (e.g., an amino acid reactive group, such as PITC) and a click chemistry moiety (e.g., an azide). In some cases, bifunctional adapter 109 is coupled to a terminal monomer (e.g., a terminal amino acid, such as an N-terminal amino acid (NTAA)). Figure 1A In Figure B, the linker nucleic acid molecule 111, containing a click chemical motif (e.g., an alkyne), is reacted with and covalently linked to a bifunctional adapter 109. In some cases, the linker nucleic acid molecule 111 and the bifunctional adapter 109 are provided in a pre-coupled manner (see, for example, Figure 2 ).exist Figure 1A In Figure C, the linker nucleic acid molecule 111 is coupled to the capture moiety 105 to generate a monomer-capture moiety complex. Coupling can be mediated by hybridization of the linker nucleic acid molecule 111 and the capture moiety 105 (hybridization not shown) or by using a splinter oligonucleotide 113 containing a sequence complementary to the sequences of both the linker nucleic acid molecule 111 and the capture moiety 105. In some cases (not shown), the linker nucleic acid molecule 111 contains a self-splinting sequence, allowing it to be coupled to the capture moiety 105 in the absence of a separate splinter molecule. A ligase can be used to covalently link the linker nucleic acid molecule 111 to the capture moiety 105. Alternatively, the linker nucleic acid molecule 111 may contain a first reactive moiety (e.g., a click chemistry moiety not shown) that can react with a second reactive moiety (not shown) of the capture moiety 105. Figure 1A In Figure D, the system is subjected to conditions sufficient to cleave terminal monomers (e.g., amino acids) from polymer analyte 103 (e.g., peptide), thereby producing a cleaved monomer-capture moiety complex. These conditions may include performing an Edmund degradation reaction. The monomer cleavage from the polymer analyte produces a cleaved monomer-capture moiety complex comprising the cleaved monomer, a bifunctional linker 109, a linked nucleic acid molecule 111, and a capture moiety 105. Figure 1AFigure E shows a binder 115 (e.g., an antibody) containing another polymerizable molecule 117 (e.g., a nucleic acid barcode molecule). The binder 115 may be specific to a monomer (e.g., to an amino acid type) or to a monomer-connector complex (e.g., a PITC-amino acid complex, which can be derived as a phenylthiocarbamoyl, hydantoin, or anilinethiazolinone moiety). The polymerizable molecule 117 of the binder may contain information about the identity of the binder or the specific monomer (e.g., a single amino acid) bound to it. The polymerizable molecule 117 of the binder may contain additional sequences, such as barcode sequences, UMIs, restriction sites, transposition sites, sequences representing the number of cycles or iterations, or other functional sequences. The polymerizable molecule 117 of the binder may be coupled to another polymerizable molecule 107 coupled to substrate 101. In some cases, an extension reaction (e.g., using a polymerase) may be performed to copy the sequence of the polymerizable molecule 117 of the binder to the other polymerizable molecule 107 coupled to substrate 101. Alternatively, the polymerizable molecule 117 of the binder may be chemically (e.g., via complementary click chemistry) or enzymatically (e.g., using a ligase, ribozyme, or DNase, or a nuclease, such as a guide editor enzyme) linked to another polymerizable molecule 107 (see also...). Figure 1D Optionally, polymerizable molecule 117 can be cleaved from binder 115 (not shown). Figure 1A In Figure F, monomers can be decoupled from the capture portion 105 (e.g., removed or cleaved). For example, all or a portion of the monomer, bifunctional adapter 109, and linker nucleic acid molecule 111 can be cleaved (depicted by an asterisk). Cleavage can be performed chemically, mechanically, or enzymatically. In an example of enzymatic cleavage, the linker nucleic acid molecule 111 may contain restriction sites or other cleavage sites (e.g., uracil), and cleavage is performed by introducing a restriction enzyme or cleavage enzyme (e.g., uracil DNA glycosylase) to cleave the restriction / cleavage sites. Alternatively, the cleaved monomer-capture portion complex can be blocked with a blocking agent (not shown). The workflow 100 can then be iterated or repeated to process and sequence all or a portion of the polymer analyte 103.

[0288] Figure 1BAnother example of processing and characterizing a polymer analyte 103 (e.g., a peptide) is schematically illustrated. The polymer analyte 103 can be tagged (or pre-tagged) with a capture portion 105 (e.g., a polymerizable molecule, such as a nucleic acid molecule). The capture portion 105 can be attached, for example, at the end of the peptide (e.g., as shown at the C-terminus) or at internal residues. In some cases, the capture portion 105 can contain a barcode sequence. In some cases, the capture portion 105 may not be attached to the peptide, but may associate with it (e.g., through indirect interaction). In other examples (not shown), the peptide is coupled (directly or indirectly) to a non-nucleic acid molecule (e.g., a polymerizable molecule, such as another peptide, or other detectable markers, such as mass tags, fluorophores, radioisotopes, etc.). A connector 109 containing a linker to a nucleic acid molecule 111 is provided. The linker to the nucleic acid molecule 111 can contain any useful sequence, such as a primer sequence, a barcode sequence, a UMI, a restriction site, etc. In some cases, the linker nucleic acid molecule 111 contains a cleavable moiety, such as a restriction site, abase-free site, uracil, transposition site, etc. In process 110, the linker 109 is coupled to a monomer (e.g., a terminal amino acid) of the polymer analyte to produce a monomer-linker complex. Before, during, or after the linker-monomer coupling, the linker nucleic acid molecule 111 may be coupled to the capture moiety 105, for example via hybridization (not shown), linking, or splinting (not shown), to produce a monomer-capture moiety complex. In process 113, the monomer is cleaved (e.g., chemically or enzymatically) from the remainder of the polymer analyte 103 to provide a cleaved monomer-capture moiety complex containing the cleaved monomer coupled to the linker 109, the linker nucleic acid molecule 111, and the capture moiety 105 or their complementary sequence. In some cases, the cleaved monomer-capture moiety complex remains coupled to the polymer analyte 103 via the linker nucleic acid molecule 111 and an additional polymerizable molecule 107. The cleaved monomer-capture moiety complex may be further processed for downstream analysis.

[0289] Downstream processing and analysis may include sorting, barcoding, detection, or combinations thereof. In some cases, multiple binding agents 115 are provided; binding agents 115 can recognize and bind to different monomer types (e.g., different amino acid types). Binding agents 115 may be contacted with and bind to their respective targets by multiple cleaved monomer-capture moiety complexes containing different (cleaved) monomer types (e.g., different amino acids). In some cases (not shown), binding agents 115 contain polymerizable molecules that identify the binding agent or its homologs, such as nucleic acid molecules (e.g., nucleic acid barcode molecules); polymerizable molecules may be coupled to or transferred to the capture moiety 105 (not shown) or linked nucleic acid molecules 111 (not shown), for example, via nucleic acid extension, ligation, transposition, primer editing, etc. Alternatively or additionally, in process 119, binding agents 115 may be separated or sorted into individual compartments (not shown), for example, using nucleic acid sequences complementary to the nucleic acid molecules of the binding agent. Alternatively or otherwise, the binder may contain sorting tags that enable sorting of different binder types, such as reporter molecules, mass tags, fluorophores, or fluorescent proteins. In one such example, a first binder for a first monomer type (e.g., amino acid residues) may contain a GFP tag, and a second binder for a second monomer type may contain an RFP tag that can be sorted by fluorescence (e.g., using FACS) or affinity sorting (e.g., using beads with anti-GFP and anti-RFP antibodies).

[0290] Optionally, after process 119, the adapter 109 or the linked nucleic acid molecule 111 or a portion thereof may be removed from the monomer-adaptor complex or the cleaved monomer-adaptor complex, for example, by restriction digestion or cleavage of uracil linking the nucleic acid molecule 111 (e.g., using UDG or USER enzymes).

[0291] In some cases, the cleaved monomer-capture portion complex can be barcoded. For example, as described above, binder 115 may contain a polymerizable barcoded molecule, such as a nucleic acid barcoded molecule, that identifies the binder or its homologs; the polymerizable barcoded molecule can be coupled to or transferred to capture portion 105 (not shown) to barcode the capture portion. Alternatively or additionally, after sorting in process 119, the sorted cleaved monomer-capture portion complex can be barcoded. For example, since each compartment contains known monomers (e.g., amino acid types) based on the binding spectrum of the binder (e.g., specificity for a particular monomer type), the cleaved monomer-capture portion complex can be labeled with a polymerizable molecule 117 (e.g., a nucleic acid barcoded molecule) that identifies the specific monomer type of the cleaved monomer-capture portion complex, thereby producing a barcoded capture portion.

[0292] After barcoding, the contents of individual compartments can then be merged, and the process can be repeated to iteratively cleave and attach identification barcodes for each monomer in the polymer analyte. Subsequent barcoded molecules can be attached to the barcoded capture portion to generate additional multi-barcoded polymerizable molecules (e.g., linked or stacked barcoded nucleic acid molecules). Alternatively or otherwise, barcodes can be added to individual polymerizable molecules, such as the amplicon (not shown) of the capture portion 105 or the linked nucleic acid molecule 111. In some cases, polymerizable molecule 117 or linked nucleic acid molecule 111 may include time information (e.g., round or cycle number). After any useful number of rounds or iterations, additional barcoded polymerizable molecules (or multiple barcoded polymerizable molecules) can be removed, for example, using NGS methods and sequenced to output the identity of each monomer type that has been processed and their order or position in the polymer analyte based on the time information.

[0293] Figure 1C Another example workflow for processing or sequencing polymer analytes is illustrated schematically. Multiple polymer analytes (e.g., peptides) can be coupled to substrate 101 together with multiple capture moieties. For illustrative purposes, further sequencing workflow operations for a single polymer analyte 103 are shown; however, it should be understood that... Figures 1A to 1C The workflow operation can be performed in parallel on all or subgroups of the polymer analyte. In process 110, a connector 109 is provided. Connector 109 is capable of coupling with (i) a monomer of the polymer unit (e.g., a terminal amino acid, such as an N-terminal amino acid) and (ii) a capture portion 105, which can be used to locally tether the monomer in the vicinity of the polymer analyte. In one example, connector 109 may contain an amino acid-reactive group, such as phenyl isothiocyanate (PITC), which enables the connector to couple with an N-terminal amino acid. The connector may additionally contain a substrate-tethering portion that can be coupled with the capture portion. For example, the substrate-tethering portion may include a click chemistry portion that can be coupled with the click chemistry capture portion, for example, via an azide-alkyne or azide-cycloalkyne reaction. In other examples, the substrate-tethering portion may include a nucleic acid molecule that can be coupled with the nucleic acid capture portion, for example, via hybridization, ligation, or both (e.g., as shown in the image). Figure 1A(As shown). In process 112, the base-tethered portion of the linker can be coupled to the capture portion. In process 113, the monomer can be cleaved from the polymer analyte. In some embodiments, monomer cleavage can be mediated by stimulation, such as a chemical reaction or pH change (such as the addition of acid). After cleavage, a binding agent 115 can be provided. The binding agent (e.g., an antibody or antibody fragment) can be specific to a particular monomer among a plurality of monomers, such as specific to a particular amino acid type or derivative thereof. The binding agent may contain a detectable marker (e.g., a fluorophore, a radioisotope, a mass tag, etc.). The detectable marker can be detected (e.g., using microscopy or imaging). The identity of the cleaved terminal monomer (e.g., an amino acid) is determined. Subsequently, the linker can be removed or cleaved, and the process can be repeated or iterative to sequence the remaining monomers of the polymer analyte.

[0294] Figure 1D Another exemplary workflow for processing polymeric analytes for sequencing using a DNA typewriter method is shown, in which a nuclease is used to facilitate the barcoding of polymerizable molecules associated with the polymeric analyte. In this example, polymeric analyte 103 is provided, which is sequentially broken down into individual monomers by contact with a capture moiety and subsequent cleavage, and the individual cleaved monomers compounded with the capture moiety are contacted with a binding agent containing a barcoded molecule that identifies or encodes the binding agent. The barcoded molecule of the binding agent is coupled to the polymerizable molecule using a nuclease, or used as an editing agent to edit the polymerizable molecule to contain a barcode. The polymerizable molecule may contain one or more nuclease target sites, such as CRISPR-Cas target sites. Figure 1D In Figure A, substrate 101 is coupled to polymer analyte 103 (e.g., a peptide to be sequenced), capture portion 105 (e.g., a nucleic acid primer), and polymerizable molecule 107 (e.g., another nucleic acid primer). Polymerizable molecule 107 may comprise a tandem array of CRISPR-Cas9 target sites (“target sites”) as described above. In some cases, the tandem array comprises multiple target sites, including, for example, a first target site 107a and a second target site 107b. In some cases, the first target site 107a is active, while the second target site 107b and all other target sites in the tandem array are inactive (e.g., truncated at the 5' region or end of each target site). Polymerizable molecule 107 may associate with polymer analyte 103; for example, the polymerizable molecule may comprise a barcode sequence for identifying the polymer analyte (e.g., a barcode for identifying a peptide), spatial location, or other useful information. Alternatively or otherwise, polymerizable molecule 107 may be coupled (directly or indirectly) to polymeric analytes, for example, using a joint (e.g., a trifunctional joint), non-covalent interaction, attachment to a substrate, etc.

[0295] exist Figure 1D In Figure B, the polymer analyte can be contacted with a bifunctional connector 109, which comprises a terminal monomer coupling group (e.g., an amino acid reactive group such as PITC) and a click chemistry motif (e.g., an azide). In some cases, the bifunctional connector 109 is coupled to a terminal monomer (e.g., a terminal amino acid such as N-terminal amino acid (NTAA)). Figure 1D In Figure B, the linker nucleic acid molecule 111, containing a click chemical motif (e.g., an alkyne), is reacted with and covalently linked to a bifunctional adapter 109. In some cases, the linker nucleic acid molecule 111 and the bifunctional adapter 109 are provided in a pre-coupled manner (see, for example, Figure 2 ).exist Figure 1D In Figure C, the linker nucleic acid molecule 111 is coupled to the capture moiety 105, thereby generating a monomer-capture moiety complex. Coupling can be mediated by hybridization of the linker nucleic acid molecule 111 and the capture moiety 105 (hybridization not shown) or by using a splinting oligonucleotide 113 containing a sequence complementary to the sequences of both the linker nucleic acid molecule 111 and the capture moiety 105. In some cases (not shown), the linker nucleic acid molecule 111 contains a self-splinting sequence, allowing it to be coupled to the capture moiety 105 in the absence of a separate splinting molecule. A ligase can be used to covalently link the linker nucleic acid molecule 111 to the capture moiety 105. Alternatively, the linker nucleic acid molecule 111 may contain a first reactive moiety (e.g., a click chemistry moiety not shown) that can react with a second reactive moiety (not shown) of the capture moiety 105. Figure 1D In Figure D, the system is subjected to conditions sufficient to cleave terminal monomers (e.g., amino acids) from polymer analyte 103 (e.g., peptide), thereby producing a cleaved monomer-capture moiety complex. These conditions may include the use of chemical stimuli (e.g., performing an Edmund degradation reaction) or biological stimuli (e.g., enzymes). The monomer cleavage from the polymer analyte produces a cleaved monomer-capture moiety complex comprising the cleaved monomer, a bifunctional linker 109, a linker nucleic acid molecule 111, and a capture moiety 105. Figure 1DFigure E shows a binder 115 (e.g., an antibody) containing a barcode molecule 117 (e.g., a nucleic acid barcode molecule). The binder 115 may be specific to a monomer (e.g., to an amino acid type) or to a monomer-linker complex (e.g., a PITC-amino acid complex, which may be derived as a phenylthiocarbamoyl, hydantoin, or anilinethiazolinone moiety, an amino acid-guanidin complex, an amino acid-dithioester, or an amino acid-thiobenzoyl complex). The barcode molecule 117 of the binder may contain information about the identity of the binder or the specific monomer (e.g., a single amino acid) bound to it. The barcode molecule 117 of the binder may contain additional sequences, such as barcode sequences, UMIs, restriction sites, transposition sites, sequences representing cycle or iteration numbers, or other functional sequences. In some cases, the barcode molecule includes a pegRNA containing a barcode sequence. In some cases, pegRNAs also contain a key sequence configured to activate the next (adjacent) inactive target site immediately adjacent to the active target site. For example, if the first target site 107a is active and the second target site 107b and all other target sites are inactive, the key sequence can be configured to be inserted between the first target site 107a and the second target site 107b. In this case, the insertion causes inactivation of the first target site 107a and activation of the second target site 107b. In some cases, the inactive sites in the tandem array are truncated to provide or couple the key sequence to form a complete active target site.

[0296] exist Figure 1D In Figure E, the barcode molecule 117 of the binder can be coupled to or inserted into the polymerizable molecule 107. Coupling can occur using a nuclease, such as a Cas protein or a guide editor containing a modified Cas protein. For example, a guide editor enzyme can be used to insert pegRNA or its complementary sequence near an active target site (e.g., a first target site 107a), thereby inactivating the active target site and activating an adjacent target site (e.g., a second target site 107b). In one such example, the guide editor may include a Cas protein and a reverse transcriptase; the Cas protein can be used to cleave the first strand of the polymerizable molecule, and a portion of the pegRNA (e.g., a primer binding site) can hybridize with a portion of the cleaved strand. An extension reaction (e.g., reverse transcription) can then produce a complementary DNA segment complementary to at least a portion of the remaining pegRNA. The resulting complementary DNA segment can then be incorporated into the polymerizable molecule by hybridization with the second strand of the polymerizable molecule, and optionally, flanking cleavage and ligation. Additional nucleic acid reactions, such as extension, amplification, ligation, restriction digestion, etc., can be performed. Optionally, barcode molecule 117 can be cleaved from binder 115 (not shown). Figure 1AIn Figure F, monomers can be decoupled from the capture portion 105 (e.g., removed or cleaved). For example, all or a portion of the monomer, bifunctional adapter 109, and linker nucleic acid molecule 111 can be cleaved (depicted by an asterisk). Cleavage can be performed chemically, mechanically, or enzymatically. In an example of enzymatic cleavage, the linker nucleic acid molecule 111 may contain restriction sites or other cleavage sites (e.g., uracil), and cleavage is performed by introducing a restriction enzyme or cleavage enzyme (e.g., uracil DNA glycosylase) to cleave the restriction / cleavage sites. Alternatively, the cleaved monomer-capture portion complex can be blocked with a blocking agent (not shown). The workflow can then be iterated or repeated to process and sequence all or a portion of the polymer analyte 103, producing polymerizable molecules 107 containing multiple barcode sequences generated by the binding of a binding agent to its target for all or a subgroup of the polymer analyte monomer.

[0297] Polymerizable molecules can contain any useful number of target sites. Polymerizable molecules can contain 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 30, 40, 50, 60, 70, 80, 90, 100, or more target sites. In some cases, the number of target sites is predetermined, for example, based on the number of iterations or cycles of the workflow to be performed. For example, in some cases, only about 5-6 rounds of sequencing are sufficient to identify the polymer analyte (e.g., peptide). Therefore, polymerizable molecules can contain 5 or 6 target sites.

[0298] In some cases, polymerizable molecules contain a single target site, and additional target sites can be added via pegRNA, for example, using the peCHYRON method. In this method, the polymerizable molecule may contain a prototypical spacer adjacent motif (PAM) and a first target site. The pegRNA molecule of the barcode molecule coupled with the binding agent may contain a barcode sequence and a propagation sequence containing a second target site. Therefore, the coupling of the barcode molecule with the polymerizable molecule results in the insertion of the propagation sequence adjacent to the PAM and the first target site. This insertion can induce inactivation of the first target site and activation of the second target site.

[0299] The peCHYRON and DNA typewriter methods can be similarly used to barcode captured portions, for example, as... Figure 1B As shown. In one such example, and again refer to Figure 1BThe capture portion 105 may contain one or more target sites. For example, the peCHYRON method can be used, where the capture portion 105 contains a first target site and a PAM sequence. The barcode molecule 117 may contain a pegRNA molecule, which may contain a barcode sequence (e.g., encoding the identity of a binder or monomer), a propagation sequence, and a second target site. The coupling of the barcode molecule 117 to the first target site can be performed using a guide editor enzyme as described above, thereby inactivating the first target site and inserting it, as well as activating the second site. Alternatively, the DNA typewriter method can be used, where the capture portion 105 contains a tandem array of target sites, all of which are inactive except for the first site. The barcode molecule 117 may contain a pegRNA molecule, which may contain a barcode sequence (e.g., encoding the identity of a binder or monomer) and a key sequence. The insertion of the barcode molecule 117 or its complementary sequence adjacent to the first target site can be performed using a guide editor enzyme, thereby inactivating the first target site, inserting the barcode sequence (or its complementary sequence), and activating the second site. In either case, the workflow can be iterated to sequentially decompose and barcode the individual monomers of the polymer analyte.

[0300] Further processing can be performed at any useful or convenient step (e.g., after barcoding). For example, nucleic acid reactions such as nucleic acid ligation, nucleic acid extension, amplification, and DNA repair can be performed.

[0301] Figure 1EAnother workflow for sequencing a polymer analyte is illustrated. A polymer analyte 103 and a capture portion 105 are provided, optionally coupled to a substrate 101. The capture portion 105 may include a first nucleic acid molecule (e.g., DNA). In process 106, an adapter 207 and a polymerizable molecule (e.g., a linker nucleic acid molecule 111) are provided. In some cases, the adapter 109 is pre-tethered to the polymerizable molecule (depicted as linker nucleic acid molecule 111); alternatively, the adapter 109 and the polymerizable molecule may be provided separately. In process 106, the adapter 109 may be coupled to a monomer (e.g., an amino acid (e.g., NTAA) of the polymer analyte 103 (e.g., a peptide) to generate a monomer-adaptor complex. In process 112, the monomer-adaptor complex may be coupled to the capture portion 105 to generate a monomer-capture portion complex. The coupling of the monomer-adaptor complex to the capture portion 105 may be mediated by the polymerizable molecule (e.g., linker nucleic acid molecule 111). Optionally, the monomer-linker complex and the capture mole 105 may be covalently linked together using chemical (e.g., click chemistry) or enzymatic (e.g., ligase) methods (e.g., linker nucleic acid molecule 111 may be covalently linked to the capture mole 105). Alternatively or otherwise, the polymerizable molecule may contain a first sequence (not shown) complementary to and capable of hybridizing with a second sequence of the capture mole 105, or the polymerizable molecule may be linked to the capture mole 105 via a splint or bridging molecule that may contain a sequence (not shown) complementary to both the first sequence of the polymerizable molecule and the second sequence of the capture mole 105. In process 113, the monomer is cleaved from polymer analyte 103 to provide a cleaved monomer-capture mole complex containing the cleaved monomer, polymerizable molecule (shown as linker nucleic acid molecule 111), and capture mole 105 coupled to linker 109. In process 114, a binding agent 115 (e.g., antibody, binding protein, etc.) may be contacted with the monomer-capture mole complex. The binder can be configured to recognize all or part of the monomer-capture moiety complex. For example, the binder can recognize a monomer, a monomer-linker complex, or the entire monomer-capture moiety complex. In one example, the linker can contain a PITC moiety, and the binder can recognize a PITC-amino acid or a derivative thereof, such as phenylthiocarbamoyl, thiazolidinone, or hydantoin. In some cases, the binder 115 can contain a detectable moiety (not shown) or can be contacted with another binder (e.g., a secondary antibody) that may optionally contain a detectable moiety (not shown).

[0302] Using additional connectors 109 and polymerizable molecules (optionally containing cycle / round information), and by tethering additional polymerizable molecules together (e.g., tethering additional polymerizable molecules to polymerizable molecules of monomer-captured partial complexes), any of the processes (e.g., 106, 112, 113, or 114) can be iterated and repeated any number of times (“rounds”). Multiple rounds can be performed until all or subgroups of monomers in polymer analyte 103 are cleaved and tethered together. In some cases, processes 106, 112, and 113 can be iterated to produce a stack of polymerizable molecules 123 containing a set of cleaved monomers, e.g., a set of linked monomer-connector-polymerizable complexes. The stack of polymerizable molecules can then be contacted with a binding agent library that can bind to its corresponding monomer target (e.g., amino acid type).

[0303] Further downstream analysis can be performed, for example, using nanopore or nanogap systems. In one such example, a nanopore sequencing system can be used to prepare stacked polymerizable molecules 123 that can optionally be coupled with a binding agent and translocated through the nanopore sequencing system, which can output the identity of the polymerizable molecule (e.g., nucleic acid sequence), monomer type, and (if general) a separate binding agent.

[0304] Alternatively, in some cases, binding agents may not be necessary for sequencing polymerizable molecules. Figure 1F This schematically illustrates an example workflow for sequencing polymer analytes. In this example workflow, the following steps are performed: Figure 1E The same operations are performed in the workflow shown, except for process 114. The workflow can be iterated multiple times to produce stacked polymerizable molecules 123, which can be further processed, for example, analyzed using a nanopore or nanogap sequencing system, to output the polymerizable molecules and the identities of the individual monomers.

[0305] Similarly, Figure 1G This schematically illustrates another example workflow for sequencing a polymer analyte in the absence of a binding agent, and the polymer analyte can be prepared with or without a substrate. In this example, a polymer analyte 103 and a capture portion 105 are provided. The capture portion 105 may include a first nucleic acid molecule (e.g., a DNA molecule) and may include identification information of the polymer analyte 103, such as an identification barcode sequence. The capture portion 105 may additionally include a releasable or cleavable portion. In some cases, the polymer analyte and the capture portion 105 may be coupled to a substrate. Figure 1G(Illustration). In such an example, the substrate (e.g., beads, the surface of a flow cell) may contain an anchoring molecule to which the capture portion 105 can be coupled. For example, the capture portion 105 may contain a nucleic acid sequence complementary to the anchoring sequence on the substrate (e.g., beads, a flat surface). In process 106, a linker 109 and a polymerizable molecule, such as a linker nucleic acid molecule 111, are provided. In some cases, the linker 109 is pre-tethered to the polymerizable molecule (linker nucleic acid molecule 111); alternatively, the linker 109 and the polymerizable molecule (linker nucleic acid molecule 111) may be provided separately. The polymerizable molecule may contain identification time information, such as providing the cycle or round of the molecule. In process 106, the linker 109 may be coupled to a monomer (e.g., an amino acid (e.g., NTAA) of polymer analyte 103) to produce a monomer-linker complex. In process 112, the monomer-linker complex may be coupled to the capture portion 105. The coupling of the monomer-linker complex to the capture moiety 105 can be mediated by a polymerizable molecule and optionally an additional polymerizable molecule 116. Optionally, the monomer-linker complex and the capture moiety can be covalently linked together (e.g., via a linker). Alternatively or additionally, the polymerizable molecule may contain a first sequence (not shown) complementary to and capable of hybridizing with a second sequence of the capture moiety 105, or the polymerizable molecule may be linked to the capture moiety 105 via a splint or bridging molecule, which may contain a sequence (not shown) complementary to both the first sequence of the polymerizable molecule and the second sequence of the capture moiety 105. In process 113, the monomer can be cleaved from the polymer analyte 103 to produce a monomer-capture moiety complex comprising the cleaved monomer, linker 109, polymerizable molecule (e.g., linker nucleic acid molecule 111), and capture moiety 105. Using additional adapters 109 and polymerizable molecules (e.g., linking nucleic acid molecules 111, which may contain the same or different sequences), and tethering additional polymerizable molecules together (e.g., tethering additional polymerizable molecules to a monomer-capture portion complex), processes 106, 112, and 113 can be iterated and repeated any number of times (“rounds”). Multiple rounds can continue until all or a subgroup of the monomers of polymer analyte 103 are tethered together. For example, the process can be iterated to produce a stack of polymerizable molecules 123 containing a set of cleaved monomers, e.g., a set of linked monomer-adaptor-polymerizable complexes. After any useful number of rounds, the stack of polymerizable molecules 123 can be cleaved, for example, from or at the capture portion 105 using a cleavable portion. The cleaved products can then be sequenced using a nanopore or nanogap sequencing instrument.

[0306] In some cases, polymerizable molecules (e.g., linker nucleic acid molecules 111) contain time information about the cycles that provided the molecule; therefore, this time information can be used for quality control. For example, if a missing cycle number is missing, it can be inferred that an amino acid is missing or not present in the peptide, that amino acid cleavage did not occur, or other errors.

[0307] Iteration: In some cases, one or more of the operations described herein can be iterated or repeated. Iteration of operations allows for the sequential processing, analysis, or identification of individual monomers of a polymer analyte, which can allow for the reconstruction of the entire polymer analyte. For example, refer to... Figure 1A The operation of workflow 100 can be performed to encode the identity of a terminal amino acid (e.g., NTAA) (e.g., via polymerizable molecule 117) onto another polymerizable molecule 107. Workflow 100 can then be repeated to encode the identity of the (n-1)th terminal amino acid, the (n-2)th terminal amino acid, the (n-3)th terminal amino acid, etc., until all or part of the peptide has been treated. Encoding can occur on the same (another) polymerizable molecule 107, for example, to produce a polymerizable molecule comprising a stack of multiple polymerizable molecules from multiple binders, or encoding can occur on another polymerizable molecule (not shown) present on a substrate. In the former case, in some cases, a second (or third, fourth, fifth, ..., nth) cycle polymerizable molecule can be configured to couple only to the first (or second, third, fourth, ..., nth) polymerizable molecule. For example, a first cycle binder polymerizable molecule can contain a unique binding sequence not present on another polymerizable (or captured) molecule on the substrate, and a second cycle binder polymerizable molecule can bind to that unique binding sequence. Therefore, the second-cycle binder polymerizable molecule can bind only to the first-cycle binder polymerizable molecule without binding to any other polymerizable (or trapped) molecules on the substrate. In the absence of a binding event (“vacancy”), a bridging polymerizable molecule encoding a vacancy binding event but containing a unique binding sequence can be provided, allowing subsequent cycles to continue even if the binder does not bind the fragmented monomer.

[0308] Similarly, Figures 1B to 1G Any operation described herein can be iterated to sequentially analyze all or subgroups of monomers of the polymer analyte. For example, refer to Figure 1BThe workflow can be operated to encode the identity of the terminal monomer (e.g., NTAA) (e.g., via polymerizable molecule 117) onto another polymerizable molecule (not shown) or onto a barcode capture portion. These operations can be repeated to encode the identity of the (n-1)th terminal amino acid, the (n-2)th terminal amino acid, the (n-3)th terminal amino acid, etc., until all or part of the peptide has been processed. Encoding can occur on the same capture portion 105, for example, to produce a stack of polymerizable molecules containing multiple barcode polymerizable molecules 117, or encoding can occur on another polymerizable molecule (not shown). In the former case, in some cases, the second (or third, fourth, fifth, ..., nth) cycle of polymerizable molecules can be configured to be coupled only to the first (or second, third, fourth, ..., nth) polymerizable molecule, as described above.

[0309] When using one or more additional polymerizable molecules, polymerizable molecule 117 (e.g., coupled with a binding agent or provided separately after sorting) can additionally encode time information, such as cycle or iteration number, so that the order of the individual monomers can be determined. For example, for a given peptide, the terminal amino acid can be coupled to the capture moiety and cleaved, and then contacted with a binding agent containing a barcode molecule that contains a barcode sequence that identifies (i) the identity of the amino acid (e.g., any one of the twenty protein amino acids, a post-translational modified amino acid, etc.) and (ii) the cycle number (e.g., cycle 1) (not shown). The information encoded by the barcode sequence can be coupled to an adjacent (additional) polymerizable molecule (not shown) or to capture moiety 105 or a barcode-coded capture moiety to produce stacked barcode-coded capture moies. After cleaving the monomers from the capture moiety (e.g., as Figure 1B As shown in process 121, this workflow can be repeated for the (n-1)th terminal amino acid, which can again be coupled to the capture portion, cleaved, and contacted with a binding agent that may contain an additional barcode sequence identifying (i) the identity of the amino acid (e.g., any one of the twenty protein amino acids) and (ii) the cycle number (e.g., cycle 2). Information encoded by the additional barcode sequence can be transferred to the same capture portion 105 or a barcoded capture portion, or to another polymerizable molecule (not shown). In the former case, the polymerizable molecule can then contain information about: (i) the identity of the terminal amino acid, (ii) the cycle number of the terminal amino acid (cycle 1), (iii) the identity of the (n-1)th terminal amino acid, and (iv) the cycle number of the (n-1)th terminal amino acid (cycle 2), etc. Alternatively or additionally, temporal information can be provided on another molecule, such as capture portion 105, linker nucleic acid molecule 111, another polymerizable molecule 116, etc.

[0310] Alternatively or otherwise, the binder may contain sorting tags and may provide a barcode sequence after sorting. For example, the binder may be used to sort different cleaved amino acid-linker complexes (as shown in process 119) and may provide a polymerizable molecule 117 containing an identity of the amino acid type for each sorted cleaved amino acid-linker complex. The polymerizable molecule 117 may also contain time information, such as providing the molecule's cycle or rotation.

[0311] In some cases, time information can be provided separately. For example, when combining polymerizable molecule 117 containing barcode information with another polymerizable molecule 107 (… Figure 1A ) or capture portion 105 or linker nucleic acid molecule 111 ( Figure 1B Before, during, or after coupling, it can provide polymerizable molecules 117 ( Figures 1A to 1B ), and another polymerizable molecule 107 ( Figure 1A ), capture portion 105 or linker nucleic acid molecule 111 ( Figure 1A , Figure 1B , Figures 1D to 1G A time barcode, or a combination thereof, may be coupled to a time barcode. The time barcode may include any useful agent, including nucleic acid molecules, peptides, lipids, carbohydrates, enzymes (e.g., chromophores or luciferases) or ribozymes or DNases, fluorophores, dyes, intercalating agents, dideoxynucleotides, fluorescent nucleic acid molecules or nucleotides, radioisotopes, quality tags, or other detectable markers that may indicate the time or number of cycles (or iterations) providing the time barcode. In some cases, the time barcode contains a cycle-specific nucleic acid barcode molecule that may be coupled to a polymerizable molecule 117 (containing the identity of the monomer) or to a terminal polymerizable molecule of a stack of polymerizable molecules containing polymerizable molecules from multiple rounds or iterations. The time barcode may contain any additional useful functional sequences, such as primer sites, sequencing sites, restriction sites, base-free or cleavable sites, etc. In some cases, the time barcode may contain an amplification site that allows bridging amplification of the time barcode and optionally coupled polymerizable molecules with other captured or polymerizable molecules.

[0312] In some cases, barcoding occurs only on a single polymerizable molecule. For example, see reference... Figure 1D Barcoding using the peCHYRON method or the DNA typewriter method can improve the stacking efficiency of barcoded molecules, for example, through specific binding events between the binding agent containing the barcoded molecule and its monomer target. The stacked barcoded molecules can then be efficiently read out using NGS methods as described elsewhere in this document for the identification and sequencing of polymer analytes.

[0313] Polymer analytes: Polymer analytes can be biomolecules, macromolecules, or synthetic molecules. Polymer analytes can be biomolecules or other biomolecules comprising one or more monomers. Non-limiting examples of polymeric biomolecules include nucleic acid molecules (e.g., DNA molecules, RNA molecules, DNA:RNA hybrids, aptamers), peptides and proteins, polysaccharides, and lipid polymers (e.g., diglycerides, triglycerides, and other fatty acids). Polymer analytes can be synthetic molecules, such as peptide-like substances or synthetic polymers, or peptide-like substances (e.g., peptide-like substances, β-peptides, D-peptide-like peptides). Non-limiting examples of synthetic polymers include acrylics, nylon, silicones, viscose, rayon, polyesters, polycarboxylic acids, polyvinyl acetate, polyacrylamide, polyacrylates, polyethylene glycol, polyurethane, polylactic acid, silica, polystyrene, polyacrylonitrile, polybutadiene, polycarbonate, polyethylene terephthalate, poly(chlorotrifluoroethylene), poly(ethylene oxide), polyethylene terephthalate, polyethylene, polyisobutylene, poly(methyl methacrylate), poly(formaldehyde), polyoxymethylene, polypropylene, polystyrene, poly(tetrafluoroethylene), poly(vinyl acetate), poly(vinyl alcohol), poly(vinyl chloride), poly(vinylidene chloride), poly(vinylidene fluoride), poly(vinyl fluoride), or combinations thereof. Polymer analytes may contain a single polymer type (e.g., homopolymer) or more than one polymer type (e.g., copolymer), and may contain random or arranged monomers. The polymer analytes can be block polymers, alternating copolymers, periodic copolymers, statistical copolymers, stereoblock copolymers, gradient copolymers, branched copolymers, graft copolymers, etc. In some embodiments, the polymer analytes described herein are proteins or peptides containing monomeric amino acids.

[0314] Polymer analytes can be of any size or within a certain size range. The size of polymer analytes can be approximately 1 nanometer (nm), approximately 5 nm, approximately 10 nm, approximately 20 nm, approximately 30 nm, approximately 40 nm, approximately 50 nm, approximately 60 nm, approximately 70 nm, approximately 80 nm, approximately 90 nm, approximately 100 nm, approximately 200 nm, approximately 300 nm, approximately 400 nm, approximately 500 nm, approximately 600 nm, approximately 700 nm, approximately 800 nm, approximately 900 nm, approximately 1 micrometer (µm), approximately 10 µm, approximately 100 µm, approximately 1 millimeter (mm), or larger. Multiple polymer analytes can comprise polymer analytes with similar sizes or within a certain size range (e.g., between approximately 10 nm and approximately 100 nm, or between approximately 50 nm and approximately 1 µm). Similarly, polymer analytes can have any molecular weight or a molecular weight range. Polymer analytes can be approximately 10 Daltons (Da), 100 Da, 500 Da, 1 kilodaltons (kDa), 10 kDa, 100 kDa, 1,000 kDa, 10,000 kDa, 100,000 kDa, or larger. Polymer analytes can include polymer analytes with similar molecular weights or within a certain molecular weight range.

[0315] The monomers of the polymer analyte can have any size or size range smaller than the monomers of the entire polymer analyte. The size of the monomers can be about 0.1 nanometers (nm), about 0.5 nm, about 1 nm, about 5 nm, about 10 nm, about 20 nm, about 30 nm, about 40 nm, about 50 nm, about 60 nm, about 70 nm, about 80 nm, about 90 nm, about 100 nm, about 200 nm, about 300 nm, about 400 nm, about 500 nm, about 600 nm, about 700 nm, about 800 nm, about 900 nm, about 1 micrometer (µm), about 10 µm, about 100 µm, about 1 millimeter (mm) or larger. The monomers can have any molecular weight or molecular weight range. Monomers can be approximately 1 Dalton (Da), 10 Da, 100 Da, 500 Da, 1 kilodalton (kDa), 10 kDa, 100 kDa, 1,000 kDa, 10,000 kDa, 100,000 kDa, or larger. Monomer or polymer analytes can be in the range of molecular weights; for example, polymer analytes can include peptides containing amino acid monomers with molecular weights ranging from 75 Da (glycine) to 204 Da (tryptophan).

[0316] Polymer analytes can contain any number of monomers. A polymer analyte can contain approximately 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1,000, 2,000, 3,000, 4,000, 5,000, 6,000, 7,000, 8,000, 9,000, 10,000, 50,000, 100,000, or more monomers. The polymer analyte may contain at least about 2, at least about 3, at least about 4, at least about 5, at least about 6, at least about 7, at least about 8, at least about 9, at least about 10, at least about 20, at least about 30, at least about 40, at least about 50, at least about 60, at least about 70, at least about 80, at least about 90, at least about 100, at least about 500, at least about 1,000, at least about 5,000, at least about 10,000, at least about 50,000, at least about 100,000 or more monomers. Alternatively, the polymer analyte may contain up to about 100,000, up to about 50,000, up to about 10,000, up to about 5,000, up to about 1,000, up to about 500, up to about 100, up to about 50, up to about 10, up to about 50, up to about 10, up to about 500 monomers. The polymer analyte may contain a range of monomers; for example, one polymer analyte may contain about 5 monomers, while another polymer analyte may contain about 500 monomers.

[0317] In some cases, polymeric analytes include peptides containing amino acid monomer units. Peptides can be naturally occurring or synthetic. Peptides can contain any number of amino acids. The amino acids can be one of the 20 protein amino acids and can contain any number of post-translational modifications. Peptides or any of their constituent amino acids can be treated, such as by contacting a protecting group, alkylation, β-elimination of phosphate groups, etc., as described elsewhere herein. In some cases, peptides are derived from larger peptides or proteins and are fragmented.

[0318] Substrate: One or more operations described herein can be performed using a substrate. For example, one or more molecules described herein (e.g., polymeric analytes, trapping portions, polymerizable molecules) can be coupled to a substrate. In some cases, polymeric analytes, trapping portions, and one or more polymerizable molecules (e.g., a first polymerizable molecule or a second polymerizable molecule) or combinations thereof can be provided coupled to one or more substrates. In one example, a polymeric analyte, a trapping portion, and a second polymerizable molecule are coupled to a substrate. In some cases, more than one substrate can be used. In such cases, the substrates can contain the same material or different materials.

[0319] The substrate can be made of any suitable material (e.g., glass, silicon, gel, polymer, etc.), as described elsewhere in this document. In some cases, the substrate may be beads or gel beads (e.g., polyacrylamide, agarose, or TentaGel). ® (Beads). The substrate can be functionalized. One or more molecules, such as the capture fraction and polymer analyte (e.g., peptide), can be coupled to the substrate via covalent or non-covalent interactions. Any suitable chemical can be used to couple the capture fraction and polymer analyte (e.g., peptide) to the substrate, such as click chemistry fractions (e.g., alkyne-azide coupling), photoreactive groups (e.g., benzophenone), 1-ethyl-3-(3-dimethylaminopropyl)carbodiimide hydrochloride (EDC) (e.g., to couple amino-oligonucleotides or peptides), N-hydroxysulfosuccinimide (NHS), sulfonyl-NHS or NHS-esters (e.g. Examples of coupling agents include thiol oligonucleotides, maleimides, hydrazines, hydroxylamines, thiols, biotin-streptavitin interactions, cystamine, glutaraldehyde, formaldehyde, 4-(N-maleimidemethyl)cyclohexane-1-carboxylic acid succinimide ester (SMCC), sulfonyl-SMCC, 4-(4,6-dimethoxy-1,3,5-triazin-2-yl)-4-methylmorpholinium chloride (DMTMM), silanes (e.g., aminosilanes), and combinations thereof. In some cases, the substrate may be functionalized to contain coupling chemicals, thereby coupling the polymer analyte or the captured portion. In a non-limiting example, the substrate (e.g., beads or surface) may contain alkynes, such as dibenzocyclooctyne (DBCO), which may be configured to react with amines (e.g., DBCO-alcohol, DBCO-Boc, DBCO-NHS), carboxyl or carbonyl groups (e.g., DBCO, DBCO-silane), thiols, etc. Azide-functionalized nucleic acids or proteins can react with DBCO to attach the nucleic acid or protein to the DBCO substrate. In other examples, adapters such as bifunctional adapters can be used to attach molecules to the substrate; such bifunctional adapters may contain the same reactive portion at both ends or different portions at each end (e.g., heterobifunctional adapters).

[0320] In some cases, enzymatic methods can be used to couple molecules (e.g., polymeric analytes, capture moieties, polymerizable molecules) to a substrate, for example, as described elsewhere herein. For instance, enzymes can be used to attach chemical connectors or moieties (such as click chemi) to polymeric analytes (e.g., peptides). The chemical connectors or moieties may be able to react with another chemical connector or moiety (e.g., click chemi) of the substrate, capture moieties, or polymerizable molecules.

[0321] The substrate can be coupled with any useful number of molecules (e.g., polymeric analytes, trapping fractions, polymerizable molecules). In some cases, the substrate may contain multiple polymeric analytes, multiple trapping fractions, and / or multiple polymerizable molecules, which can be provided at any useful ratio or density. For example, the ratio of polymeric analytes to trapping fractions or polymerizable molecules may be about 1:1, 1:5, 1:10, 1:20, 1:100, 1:1000, 1:10,000, 1:100,000, 1:100,000, 1:1,000,000, or lower. In some cases, the ratio of polymeric analytes to trapping fractions or polymerizable molecules may be up to about 1:1, up to about 1:5, up to about 1:10, up to about 1:20, up to about 1:100, up to about 1:1000, up to about 1:10,000, up to about 1:100,000, up to about 1:1,000,000, or lower.

[0322] Similarly, molecules (e.g., polymeric analytes, captured fractions, or polymerizable molecules) can be present at any useful density, such as about 1 molecule per square micrometer (µm). 2 Approximately 10 molecules / µm 2 Approximately 100 molecules / µm 2 Approximately 1,000 molecules / µm 2 Approximately 10,000 molecules / µm 2 Approximately 100,000 molecules / µm 2 Approximately 1,000,000 molecules / µm 2 Approximately 10,000,000 molecules / µm 2 Approximately 100,000,000 molecules / µm 2 Approximately 1,000,000,000 molecules / µm 2 Approximately 10,000,000,000 molecules / µm 2 Approximately 100,000,000,000 molecules / µm 2 Or even larger, coupled to the substrate. Polymer analytes, trapped fractions, and polymerizable molecules can be coupled at densities ranging from approximately 100 to approximately 10,000 molecules / µm. 2or approximately 10 to approximately 1,000 molecules / µm 2 The polymer analyte, the trapped portion, and the polymerizable molecule can be the same or different. For example, the density of the polymerizable molecule can be 1, 1 / 2, 1 / 3, 1 / 4, 1 / 5, 1 / 6, 1 / 7, 1 / 8, 1 / 9, 1 / 10, 1 / 100, 1 / 1000, 1 / 10,000, 1 / 100,000, 1 / 1,000,000, or less than the density of the polymer analyte.

[0323] In some cases, molecules coupled to the substrate can be spaced apart at a specified or controlled distance. For example, the average spacing or distance between polymerizable molecules coupled to the substrate can be about 1 nanometer (nm), about 2 nm, about 3 nm, about 4 nm, about 5 nm, about 6 nm, about 8 nm, about 9 nm, about 10 nm, about 20 nm, about 30 nm, about 40 nm, about 50 nm, about 60 nm, about 70 nm, about 80 nm, about 90 nm, about 100 nm, about 500 nm, about 1 µm, or greater. In some cases, the spacing between polymerizable molecules coupled to the substrate can be up to about 1 µm, up to about 500 nm, up to about 100 nm, up to about 90 nm, up to about 80 nm, up to about 70 nm, up to about 60 nm, up to about 50 nm, up to about 40 nm, up to about 30 nm, up to about 20 nm, up to about 10 nm, up to about 5 nm, or less. Similarly, the spacing or distance between the polymeric analyte and the polymerizable molecule or trapping portion can be about 1 nanometer (nm), about 2 nm, about 3 nm, about 4 nm, about 5 nm, about 6 nm, about 8 nm, about 9 nm, about 10 nm, about 20 nm, about 30 nm, about 40 nm, about 50 nm, about 60 nm, about 70 nm, about 80 nm, about 90 nm, about 100 nm, about 500 nm, about 1 µm or greater. In some cases, the average spacing between the polymerizable molecule and the polymeric analyte coupled to the substrate can be up to about 1 µm, up to about 500 nm, up to about 100 nm, up to about 90 nm, up to about 80 nm, up to about 70 nm, up to about 60 nm, up to about 50 nm, up to about 40 nm, up to about 30 nm, up to about 20 nm, up to about 10 nm, up to about 5 nm or less. The average distance range between polymerizable molecules or between polymeric analytes can be used, for example, about 1 nm to about 40 nm, about 2 nm to about 10 nm, etc.

[0324] The concentration or density of molecules attached to a substrate can be adjusted using one or more suitable methods, including patterning or random deposition methods. Examples of methods for controlling the concentration or density of molecules attached to a substrate include limiting dilution, adding a liquid release agent (e.g., guanidine, formamide, urea), using organometallic compounds, etc. Molecules can be attached to the substrate in a patterned manner (e.g., using self-assembled monolayers, photopatterning, photolithography, etching, or combinations thereof), or molecules can be arranged randomly.

[0325] The substrate can have any useful size or dimension (e.g., length, width, height, diameter, radius), surface area, volume, or ratio, or a combination thereof. The substrate may include beads or particles that can have dimensions of about 1 nanometer (nm), about 2 nm, about 3 nm, about 4 nm, about 5 nm, about 6 nm, about 8 nm, about 9 nm, about 10 nm, about 20 nm, about 30 nm, about 40 nm, about 50 nm, about 60 nm, about 70 nm, about 80 nm, about 90 nm, about 100 nm, about 500 nm, about 1 µm, about 2 µm, about 3 µm, about 4 µm, or about [missing value]. 5µm, approximately 6µm, approximately 7µm, approximately 8µm, approximately 9µm, approximately 10µm, approximately 20µm, approximately 30µm, approximately 40µm, approximately 50µm, approximately 60µm, approximately 70µm, approximately 80µm, approximately 90µm, approximately 100µm, approximately 200µm, approximately 300µm, approximately 400µm, approximately 500µm, approximately 600µm, approximately 700µm, approximately 800µm, approximately 900µm, approximately 1 millimeter (mm) or larger in diameter. The substrate may include a surface area of ​​about 1 square nanometer (nm2), about 10 nm2, about 100 nm2, about 1,000 nm2, about 10,000 nm2, about 100,000 nm2, about 1 μm2, about 10 μm2, about 100 μm2, about 1,000 μm2, about 10,000 μm2, about 100,000 μm2, about 1 mm2, about 10 mm2, about 100 mm2, about 1,000 mm2, about 10,000 mm2, about 100,000 mm2, about 1,000,000 mm2, or larger.

[0326] Molecules can be coupled to the substrate in an ordered or random arrangement. In an ordered arrangement, molecules can be patterned using any conventional method, such as photolithography (e.g., soft photolithography, photolithography), etching (e.g., ion etching, photolithography), or other patterning methods. In some cases, connectors (e.g., bifunctional connectors) can be used to facilitate the coupling of molecules (e.g., polymeric analytes, polymerizable molecules, trapping moieties) to the substrate; such connectors can be patterned using any useful technique (e.g., self-assembled monolayers, photopatterning, photolithography, etching). In some cases, molecules can be coupled to the substrate in a random arrangement. For example, molecules can be provided at stoichiometric ratios or controlled concentrations to couple molecules at any useful ratio or density.

[0327] Polymerizable molecules: Polymerizable molecules (including barcode molecules as described herein) can include any useful type of polymerizable molecule. Polymerizable molecules can be naturally occurring, such as biopolymers (e.g., nucleic acid molecules, peptides, polysaccharides, fatty acids) or other naturally occurring polymers, such as rubber, cellulose, starch, polyhydroxyalkanoates, deacetylated chitosan, dextran, structural proteins (e.g., collagen, hyaluronic acid, glycosaminoglycans), agarose, carrageenan, psyllium husk gum, gum arabic, agar, gelatin, shellac, xanthan gum, guar gum, alginate, etc. Polymerizable molecules can be synthetic, such as acrylics, nylon, silicones, viscose, rayon, polyesters, polycarboxylic acids, polyvinyl acetate, polyacrylamide, polyacrylates, polyethylene glycol, polyurethane, polylactic acid, silica, polystyrene, polyacrylonitrile, polybutadiene, polycarbonate, polyethylene terephthalate, poly(chlorotrifluoroethylene), poly(ethylene oxide), polyethylene terephthalate, polyethylene, polyisobutylene, poly(methyl methacrylate), poly(formaldehyde), polyoxymethylene, polypropylene, polystyrene, poly(tetrafluoroethylene), poly(vinyl acetate), poly(vinyl alcohol), poly(vinyl chloride), poly(vinylidene chloride), poly(vinylidene fluoride), poly(vinyl fluoride), and combinations thereof. Polymerizable molecules may contain one or more reactive moieties (e.g., free radical groups) to initiate polymerization, or they may be polymerized via a contact initiator (e.g., ammonium persulfate, peroxide, or other free radicalizing agent). Polymerizable molecules can be polymerized via contact enzymes (e.g., polymerizing enzymes, such as polymerases), ribozymes, or DNases. Alternatively or otherwise, polymerizable molecules can be polymerized via self-assembly. Polymerizable molecules may include a single polymer type (e.g., homopolymer) or more than one polymer type (e.g., copolymer), and may contain randomly or arranged monomers. Polymerizable molecules can be block polymers, alternating copolymers, periodic copolymers, statistical copolymers, stereoblock copolymers, gradient copolymers, branched copolymers, graft copolymers, etc.

[0328] Polymerizable molecules of the same or different types can be used in the methods described herein. For example, the first polymerizable molecule contained in or coupled to the binder can be a nucleic acid molecule, and the second polymerizable molecule can be a peptide. In another example, both the first and second polymerizable molecules are nucleic acid molecules. In this example, the first polymerizable molecule can be coupled to the second polymerizable molecule via ligation or hybridization. For example, the first polymerizable molecule can contain a first nucleic acid sequence, and the second polymerizable molecule can contain a second nucleic acid sequence. The first nucleic acid sequence can be complementary or partially complementary to the second nucleic acid sequence, and coupling can include hybridizing the first nucleic acid sequence or a portion thereof with the second nucleic acid sequence or a portion thereof. Alternatively, the first nucleic acid sequence and the nucleic acid sequence can be complementary to two sequences of a splice oligonucleotide or a bridging oligonucleotide, and coupling can be mediated via hybridization with the splice oligonucleotide. The first nucleic acid sequence can be chemically (e.g., via click chemistry, in which the first and second polymerizable molecules contain one member of a click chemical pair) or enzymatically (e.g., using a ligase) linked to the second nucleic acid sequence.

[0329] Polymerizable molecules may contain functional parts. For example, polymerizable molecules may include nucleic acid molecules that contain functional sequences such as primer sequences (e.g., universal initiation sites), sequencing sequences, read sequences, unique molecular identifiers (UMIs), barcode sequences, cleavage sequences (e.g., restriction sites, Cas binding sequences), transposon sequences (e.g., chimeric end sequences), or combinations thereof.

[0330] Polymerizable molecules can be of any useful size. The size, length, or other dimensions of a polymerizable molecule can be approximately 1 angstrom, approximately 2 angstroms, approximately 3 angstroms, approximately 4 angstroms, approximately 5 angstroms, approximately 6 angstroms, approximately 7 angstroms, approximately 8 angstroms, approximately 9 angstroms, approximately 10 angstroms, approximately 20 angstroms, approximately 30 angstroms, approximately 40 angstroms, approximately 50 angstroms, approximately 60 angstroms, approximately 70 angstroms, approximately 80 angstroms, approximately 90 angstroms, approximately 100 angstroms, approximately 200 angstroms, approximately 300 angstroms, approximately 400 angstroms, approximately 500 angstroms, approximately 600 angstroms, approximately 700 angstroms, approximately 800 angstroms, approximately 900 angstroms, approximately 1000 angstroms, approximately 10,000 angstroms, approximately 100,000 angstroms, or larger. In some cases, polymerizable molecules (e.g., a first polymerizable molecule or a second polymerizable molecule) comprise nucleic acid molecules containing one or more nucleotide bases. Polymerizable molecules can contain any useful number of nucleotide bases, such as about 1 base, about 2 bases, about 3 bases, about 4 bases, about 5 bases, about 6 bases, about 7 bases, about 8 bases, about 9 bases, about 10 bases, about 20 bases, about 30 bases, about 40 bases, about 50 bases, about 60 bases, about 70 bases, about 80 bases, about 90 bases, about 100 bases, about 200 bases, about 300 bases, about 400 bases, about 500 bases, about 600 bases, about 700 bases, about 800 bases, about 900 bases, about 1000 bases, or more.

[0331] Polymerizable molecules can include nucleic acid molecules. Nucleic acid molecules can be single-stranded, double-stranded, or partially double-stranded. Nucleic acid molecules can contain modified nucleotides or atypical bases. For example, polymerizable molecules can include pseudo-complementary bases, bridged nucleic acids (BNA), heteronucleic acids (XNA), locked nucleic acids (LNA), peptide nucleic acids (PNA), γ-PNA molecules, morpholinonucleotides, or combinations thereof. In some cases, polymerizable molecules can include hexitol nucleic acids (HNA) or cyclohexyl nucleic acids (CeNA), which can be used to make the polymerizable molecule more resistant to acid degradation (e.g., as used in conventional Edman degradation). Alternatively or otherwise, polymerizable molecules can contain naturally occurring bases that are more resistant to acid degradation, such as those consisting primarily of thymine or cytosine. For example, nucleic acid molecules can contain at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, or 100% thymine or cytosine, which can make the nucleic acid molecule more resistant to acid compared to nucleic acid molecules containing adenine or guanine.

[0332] Linkers: Linkers can be used to mediate one or more operations of the method. In some cases, linkers are used to mediate the coupling of monomers (e.g., amino acids) with a capture moiety to produce a monomer-capture moiety complex. The coupling of the linker to the monomer or capture moiety can be covalent or non-covalent. In one example, the linker may contain a first reactive group capable of coupling with a monomer (e.g., an amino acid of a peptide) of the polymer analyte and optionally cleaving the amino acid from the peptide. For example, the first reactive group can be an amino acid reactive group, such as isothiocyanates (ITCs) like phenyl isothiocyanate (PITC), 3-pyridyl isothiocyanate (PYITC), 2-piperidinylethyl isothiocyanate (PEITC), 3-(4-morpholino)propyl isothiocyanate (MPITC), 3-(diethylamino)propyl isothiocyanate (DEPTIC), or naphthyl isothiocyanate (NITC), fluorescein isothiocyanate (FITC), ammonium thiocyanate, potassium thiocyanate, trimethylsilyl isothiocyanate (TMS-ITC), phenyl phosphate isothiocyanate, acetyl isothiocyanate (AITC), or an aldehyde group, such as o-phthalaldehyde (OPA), 2,3-naphthalenedicarboxyl (NDA), 2-pyridinaldehyde, which can react with N-terminal amino acids (NTAAs). The linker may additionally contain a second reactive group capable of direct or indirect coupling with the capturing moiety. In one example of direct coupling, the capturing portion may include a click chemistry portion (e.g., an alkyne), and the second reactive group of the linker may include an additional click chemistry portion (e.g., an azide) that can react with the click chemistry portion of the capturing portion. Alternatively, the linker may be indirectly coupled to the capturing portion, for example, via a non-covalent interaction or via an intermediate linker molecule. In some cases, the intermediate linker molecule may include a third polymerizable molecule (e.g., a polymer or nucleic acid molecule) that can couple the linker to the capturing portion. In one such example, the third polymerizable molecule may comprise (i) a third reactive group capable of coupling with the second reactive group of the linker (e.g., via alkyne-azide click chemistry) and (ii) a portion capable of coupling with the capturing portion (e.g., another orthogonal click chemistry reaction, avidin-biotin interaction, nucleic acid coupling, or hybridization). In some cases, the third polymerizable molecule includes a nucleic acid molecule comprising (i) a click chemical moiety (e.g., an alkyne) that can conjugate with a first reactive group (e.g., an azide) of the adapter and (ii) a nucleic acid sequence that can be coupled to a capture moiety, for example via linking, splinting, or hybridization. In some cases, the adapter comprises a linker nucleic acid molecule containing a self-splastic moiety.

[0333] Where applicable, the click chemistry of the linker and the trapping or intermediate linking molecule can include any suitable bioorthogonal motif as described elsewhere herein, such as alkenes, alkynes, azides, epoxides, amines, thiols, nitrones, isonitriles, isocyanates, aziridines, activated esters, and tetrazides, as well as combinations, variants, or derivatives thereof. The linker can withstand conditions sufficient to allow the first click chemistry to react with the second click chemistry, such as the provision of a metal catalyst, a suitable solvent, pH, temperature, ion concentration, or light / energy, for any useful duration.

[0334] The first reactive group of the linker can be an amino acid reactive moiety. The amino acid reactive moiety of the linker can be any useful portion capable of reacting with the reactive moiety with an amino acid conjugate and optionally cleaving the amino acid. In some examples, the first reactive moiety can react with a terminal amino acid (e.g., NTAA or CTAA). In such examples, the first reactive portion may include any primary amine or carboxyl group reactive group, including but not limited to isocyanates, acyl azides, NHS esters, sulfonyl chlorides, aldehydes, glyoxal, epoxides, ethylene oxide, carbonates, aryl halides, imine esters, carbodiimides, acid anhydrides, phenyl esters, isothiocyanates (e.g., phenyl isothiocyanate, sodium isothiocyanate, ammonium isothiocyanate (e.g., tetrabutylammonium isothiocyanate, tetrabutylammonium isothiocyanate), diphenylphosphoisothiocyanate), acetyl chloride, cyanogen bromide, carboxypeptidase, azides, alkynes, DBCO, maleimide, succinimide, thiol-thiol disulfide bond, tetraazine, TCO, vinyl, methylcyclopropene, acryloyl, allyl, etc. Further examples of amino acid reactive groups are provided in U.S. Patent No. 11,499,979, U.S. Patent Publication No. 2020 / 0217853, International Patent Publication No. WO / 2024 / 107755 and International Patent Application No. PCT / US2024 / 013211, each of which is incorporated herein by reference in its entirety.

[0335] The linker can contain any additional useful portion. For example, the linker can contain a releasable or cleavable portion that facilitates the removal of monomers from the polymer analyte or a portion thereof, or from the substrate. Such a releasable or cleavable portion can include, for example, a disulfide bond that can be released by contact with a reducing agent (e.g., DTT, TCEP). In some examples, the linker can be coupled to a third polymerizable molecule via the releasable or cleavable portion, alternatively or otherwise via a click chemistry portion. Thus, the coupling between the polymerizable molecule and the linker can be reversible. The linker can additionally or alternatively contain any number of spacer portions, such as polymers (e.g., PEG, PVA, polyacrylamide), aminocaproic acid, nucleic acids, alkyl chains, etc. Such spacer portions can increase the distance between any other portions of the linker (e.g., amino acid reactive groups and polymerizable molecule reactive groups). The linker can contain or be coupled to a detectable portion (e.g., a fluorophore, a radioisotope, a mass tag, a nucleic acid molecule (which can also act as a releasable or cleavable portion) or other detectable portion). In some examples, the connector contains a fluorophore, which can be visualized using single-molecule imaging or fluorescence lifetime imaging. In another example, the monomer can be labeled with a first fluorophore, and the connector can contain a second fluorophore to enable the visualization of the connector and monomer (e.g., using dual-channel imaging or FRET).

[0336] Using a linker containing two reactive groups allows the linker to be coupled to (i) a monomer of the polymer analyte and (ii) an intermediate linker molecule (e.g., a third polymerizable molecule) or (iii) a capture moiety. In some cases, when using an intermediate linker molecule, the linker can be pre-coupled to the intermediate linker molecule. For example, a precursor linker may contain a monomer-binding group (e.g., PITC) and a click chemistry moiety (e.g., an azide), which can react with a polymerizable molecule (e.g., an oligonucleotide) containing a complementary click chemistry moiety (e.g., an alkyne) to produce a linker capable of coupling to both the monomer and the capture moiety (e.g., another oligonucleotide). In some cases, a linker pre-coupled to an intermediate linker molecule can be provided.

[0337] Figure 2 An example adapter that can be used to sequence polymer analytes such as peptides is illustrated schematically. Figure 2Figure A illustrates a bifunctional linker 203 (e.g., 1-(but-3-yn-1-yl)-4-isothiocyanobenzene) comprising an amino acid reactive moiety (e.g., PITC) and an alkyne click chemistry moiety, which can react with a polymerizable molecule 201 (e.g., a linked nucleic acid molecule) comprising a complementary azide click chemistry moiety. The bifunctional linker may also include a spacer region, such as an alkyl chain of any length (e.g., an ethyl group is depicted), a polymer of any length (e.g., PEG), etc. The spacer region may be located between the amino acid reactive moiety and the click chemistry moiety. Figure 2 Figure B illustrates the product of a click cycloaddition reaction between an azide and an alkyne group to produce a connector molecule comprising a polymerizable molecule and a reactive amino acid moiety. The conjugation of the polymerizable molecule 201 with the bifunctional connector 203 can occur at any useful or convenient step. In an alternative example (not shown), the bifunctional connector 203 may comprise an azide group, such as 1-(2-azidoethyl)-4-isothiocyanobenzene, which can react with the polymerizable molecule 201 comprising an alkyne moiety (e.g., a straight-chain alkyne or a cycloalkyne, such as DBCO).

[0338] In some cases, such as when the polymer analyte contains a peptide with an amino acid monomer, the coupling of the linker to the amino acid (e.g., NTAA or CTAA) modifies the chemical structure of the amino acid. For example, if a linker containing an isothiocyanate moiety is used, the amino acid can be derivatized into a thiocarbamoyl group during or after contact with the isothiocyanate moiety (e.g., under mild alkaline conditions). One or more further derivatizations can be performed. For example, the amino acid or amino acid derivative (e.g., a thiocarbamoyl-derived amino acid) can be further derivatized into a thiazolyl group (e.g., which can cleave the amino acid from the peptide under acidic conditions), a thiohydantoin group, or other chemical moieties. Similarly, the thiazolyl group or the thiohydantoin group can be further derivatized into a thiocarbamoyl group.

[0339] Capture Membrane: The capture membrane can be coupled to the monomer of the polymer analyte via any suitable mechanism. Coupling between the monomer and the capture membrane can include covalent or non-covalent interactions. Coupling can occur through interactions of binding pairs such as biotin and avidin (or streptoavidin), cyclodextrin and small hydrophobic molecules (e.g., alkanes, benzenes, polycyclic aromatic hydrocarbons), cucurbituril and adamantane or trimethylammonium methylferrocene, cyclopentadiene (e.g., calixarenes, cryptanes, columnar aromatics, tetralactams), etc.

[0340] In some cases, the capture portion includes additional polymerizable molecules (e.g., nucleic acid molecules). In such cases, the monomer can first be coupled to a complementary polymerizable molecule (e.g., to generate a peptide-oligonucleotide conjugate) and tethered to the capture portion, for example, directly via complementary base pairing or via a splint molecule. Alternatively, as described above, the monomer can be coupled to the capture portion via a linker. For example, the linker can comprise a monomer-coupling group (e.g., PITC, which can couple to or react with an amino acid of the peptide) and a nucleic acid molecule. The capture portion may include additional nucleic acid molecules, which can be coupled to the nucleic acid molecule of the linker via hybridization, linkage, or both.

[0341] The capture portion may include a nucleic acid molecule, which may contain any naturally occurring, non-natural, or engineered nucleotide bases. For example, the nucleic acid molecule may include pseudo-complementary bases, bridging nucleic acids, heteronucleic acids, locked nucleic acids, peptide nucleic acids (PNA), γ-PNA, morpholino, etc., as described elsewhere herein.

[0342] The capture portion may include one or more functional sequences, including but not limited to priming sequences, sequencing sequences, sequencing read sequences, chimeric end sequences, transposase recognition sequences, cleavage sites (e.g., restriction sites), UMIs, blocking groups, spacer sequences, barcode sequences, or other functional sequences. In some cases, the capture portion includes cleavable or releasable portions (e.g., restriction enzyme recognition sites, base-free sites, or sites that can be used with USER). ® Or uracil cleaved by uracil DNA glycosylation enzymes, or disulfide bonds that can be released upon the addition of a reducing agent.

[0343] In some cases, a capture portion and a polymerizable molecule coupled to the substrate are provided. In one example, the substrate contains a first nucleic acid molecule, a second nucleic acid molecule, and a capture portion coupled thereto, which may be a third nucleic acid molecule. In some cases, the substrate may contain the same nucleic acid molecules spanning the substrate; these same nucleic acid molecules may act as both the capture portion and the polymerizable molecule, with additional polymerizable molecules (e.g., coupled to a binder) coupled to them.

[0344] The capture moiety may include any useful portion or functional group. The capture moiety may have, for example, monomer-capture groups, substrate-tethering groups, or linkers for coupling or tethering to other molecules or for detection, or any additional functional group or portion. In some examples, the capture moiety comprises a nucleic acid molecule containing substrate-tethering groups (e.g., biotin, click chemistry motifs such as azides) that can couple to a substrate (e.g., containing streptavidin or complementary click chemistry). The capture moiety may additionally include a binding sequence to which another nucleic acid molecule (e.g., a linker nucleic acid molecule, linker-monomer complex, or binding nucleic acid barcode molecule) can be coupled, for example, via hybridization, ligation, or both. In some cases, the capture moiety comprises a single-stranded oligonucleotide or a single-stranded region in which a complementary oligonucleotide can hybridize. The complementary oligonucleotide may contain a detectable label (e.g., a fluorophore) that allows detection of the capture moiety.

[0345] Cleavage: The cleavage of monomers from polymer analytes can be achieved using any suitable mechanism, such as by applying a stimulus. The stimulus can be, for example, a chemical stimulus, a biological stimulus, a thermal stimulus (e.g., the application of heat), a light stimulus, a physical or mechanical stimulus, or other types of stimuli or combinations thereof. In some cases, the stimulus can be a chemical stimulus, such as the addition of a pH change, a solubilizer, an initiator, a free radical generator, a reducing agent, etc. In some cases, the stimulus can be a biological stimulus, such as an enzyme that can cleave or catalyze the cleavage of monomers from polymer analytes (e.g., Edman's enzyme, protease, endonuclease, artificial protease such as artificial peptidases).

[0346] In some examples, the polymer analyte comprises a peptide, and the monomer comprises an amino acid (e.g., NTAA, CTAA, or an internal amino acid). The method may include using a linker containing an amino acid reactive group (e.g., PITC) to cleave the amino acid from the peptide by coupling the linker's amino acid reactive group to the amino acid and using a stimulus (e.g., a change in pH, temperature). In one example, PITC may be coupled to NTAA under mildly alkaline conditions to produce a phenylthiocarbamoyl (PTC) derivative of NTAA, and cleavage of NTAA from the peptide may be achieved using an Edmann degradation reaction (e.g., applying an acid such as trifluoroacetic acid under heating) to produce a thiazolinone (ATZ) derivative or a hydantoin (PTH) derivative. As described elsewhere herein, the linker may contain or be coupled to a capture moiety (e.g., a nucleic acid molecule or a polymerizable molecule) such that, after cleavage, the cleaved amino acid can be coupled to the capture moiety.

[0347] In some cases, more than one monomer can be cleaved from the polymer analyte at a single cleavage. Cleavage may include cleaving 2, 3, 4, 5, 6, 7, 8, 9, 10, or more monomers. For example, the polymer analyte may include peptides containing multiple amino acid monomers, and single amino acids, dipeptides, tripeptides, tetrapeptides, or larger peptides may be cleaved in the methods described herein. In some cases, in a given cleavage event, up to about 10, 9, 8, 7, 6, 5, 4, 3, or fewer monomers may be cleaved. In some cases, enzymes capable of recognizing or cleaving more than one single amino acid (e.g., edemases, proteases) may be used to mediate the cleavage of more than one monomer (e.g., amino acids).

[0348] The cleavage of monomers (or multiple monomers) can be performed using biostimuli such as enzymes. Enzymes can be any useful lysin, such as proteases, including edemases, cruzain, X proteins (e.g., ClpS, ClpX), proteinase K, exopeptidases, aminopeptidases, diaminopeptidases, serine proteases, cysteine ​​proteases, threonine proteases, aspartic proteases, aspartic proteases, glutamate proteases, metalloproteinases, asparagine peptidase, pepsin, trypsin, trypsin, Lys-C, Glu-C, Asp-N, chymotrypsin, carboxypeptidases (e.g., carboxypeptidase A, carboxypeptidase B, carboxypeptidase Y), SUMO protease, elastase, papain, endopeptidase, protease, and TrypZean. ®Bromelain, collagenase, hyaluronidase, thermophilic protease, fig protease, keratinase, trypsin, fibroblast activation, enterokinase, chymotrypsinogen, chymotrypsin, clostridium protease, calpain, α-lysin, proline-specific endopeptidase, furin, thrombin, subtilisin, gene enzyme, PCSK9, cathepsin, aminoacylproline dipeptidase, methionine aminopeptidase, cathepsin C, 1-cyclohexene-1-yl-boronate pinacol ester, pyroglutamate aminopeptidase, renin, kininogen, kallikrein, DPPIV / CD26, phorate oligopeptidase, prolyl oligopeptidase, leucine aminopeptidase, dipeptidyl peptidase or other enzymes or proteases, or combinations or variants thereof (e.g., engineered mutants or variants). In some cases, lyases, ribozymes, or DNases can be constructed or engineered to cleave terminal monomers or multiple monomers; alternatively, lyases, ribozymes, or DNases can be constructed or engineered to perform exosite cleavage at non-terminal sites of the polymer analyte, such as at internal monomers within the polymer analyte, at sites n-1, n-2, n-3, n-4, n-5, n-6, n-7, n-8, n-9, n-10, etc. (where n is the number of monomers in the polymer analyte).

[0349] In the case of enzymatic cleavage, additional reagents can be provided to catalyze or induce cleavage. For example, metalloproteinases, aminopeptidases, or exopeptidases can promote the cleavage of one or more amino acids in the presence of a catalyst (e.g., a metal or metal ion (e.g., cobalt)). Therefore, catalysts can be provided to facilitate the binding of an enzyme to an amino acid or the subsequent cleavage of amino acids from a peptide. In some examples, cleavage can be mediated by decoenzyme removal, which is inactive in the absence of a cofactor and a metal catalyst, and cleavage can be controlled by the addition of a metal or metal ion.

[0350] Other examples of cleavage stimuli may include: light stimulation (e.g., applying UV, X-rays, gamma rays, or other wavelengths of light), mechanical stimulation (e.g., sonication, high pressure), thermal stimulation (e.g., applying heat), or chemical stimulation. In some cases, polymeric analytes may contain or be modified to contain cleavable or unstable bonds that can be cleaved upon the application of a suitable stimulus, such as disulfide bonds (e.g., cleavable upon the application of a chemical stimulus such as a reducing agent), ester bonds (e.g., cleavable with pH changes), vicinal diol bonds (e.g., cleavable with sodium periodate), Diels-Alder bonds (e.g., cleavable upon the application of heat), sulfone bonds (e.g., cleavable via a base), silyl ether bonds (e.g., cleavable via an acid), glycosidic bonds (e.g., cleavable via amylase), peptide bonds (e.g., cleavable via a protease), or phosphodiester bonds (e.g., cleavable via a nuclease (e.g., DNase)).

[0351] Monomer Modification: In some cases, one or more monomers of the polymer analyte may be modified. Modifications may be naturally occurring (e.g., post-translational modifications) or non-natural, such as by labeling or tagging with amino acid or amine reactants (e.g., isothiocyanates (e.g., PITC, NITC), 1-fluoro-2,4-dinitrobenzene (DNFB), dansyl chloride, 4-sulfonyl-2-nitrobenzene (SNFB), acetylation agents, acylation agents, alkylation agents, guanidinization agents, thioacetylation agents, thioacylation agents, thiobenzoylation agents, or derivatives or combinations thereof). Alternatively or in addition, one or more monomers may be modified to include any useful portion, such as adducts (e.g., polymers (e.g., PEG), polymerizable molecules (e.g., nucleic acid molecules), nanoparticles or nanotubes, peptides or proteins), lipids, carbohydrates, metabolites, fluorophores, haptens, quenchers, labels (e.g., fluorescent labels, magnetic labels, radioactive labels), barcodes, or other portions. In some cases, monomers of polymer analytes can be modified to facilitate enzyme recruitment, thereby recognizing or cleaving terminal monomers (e.g., NTAA or CTAA of a peptide, 5' or 3' nucleotides of a nucleic acid molecule, or the first or last monomer of a polymer) or groups of monomers. For example, the terminal amino acids of a peptide analyte can be modified with sugars to recruit lectins or lectin-binding proteases. In another example, one or more monomers of a polymer analyte may contain or be coupled to a nucleic acid molecule having a first sequence complementary to a second sequence contained in an oligonucleotide-binding protease. Hybridization of the first and second sequences can facilitate local recruitment of the protease to the monomer to be cleaved. In yet another example, a peptide analyte can be modified with PITC, which can allow recruitment and cleavage by an Edman enzyme. In some examples, modifications to monomers of a polymer analyte may include epitope tags that can facilitate binding to a binder (e.g., after the monomer has cleaved from the polymer analyte). Examples of such epitope tags include fluorophores, nucleic acid molecules, peptides, haptens, polymers, chemical moieties, or other adduct molecules. Further examples of modifications to polymer analytes are described elsewhere herein.

[0352] Polymer analytes may contain one or more modified monomers. Monomer modifications can be naturally occurring or synthetic. Synthetic modifications can be performed before, during, or after cleavage of the monomer from the polymer analyte and can be beneficial for preserving the monomer's identity. For example, during a standard Edmann degradation reaction from the cleavage of terminal amino acids (monomers) of a peptide, some amino acid residues may be altered or become undetectable under the reaction conditions. In one example, the conditions of Edmann degradation may lead to oxidation of cysteine ​​residues, dehydration or destruction of serine or threonine in the form of hydantoin (PTH), reaction with and modification of lysine residues, or rendering some post-translational modifications undetectable. Therefore, modifying the peptide prior to analysis, such as protecting some amino acid residues or performing post-translational modifications, can be used to more accurately identify each amino acid residue. In one example of modifications that can be performed prior to cleavage, the peptide or a portion thereof may be alkylated, for example by alkylating cysteine ​​residues (e.g., using 4-vinylpyridine, iodoacetamide, iodoacetate, chloroacetate, or by aminoethylation, for example, using 2-bromoethylamine); acetylated, for example by reacting serine or threonine residues to form esters (e.g., using acetyl chloride) or by using acetic anhydride; oxidized, for example by converting cysteine ​​residues to sulfoalanine; reduced (e.g., using reducing agents such as dithiothreitol, [β-mercaptoethanol, or TCEP]); subjected to native chemical linkage; contacted with a protecting group, for example, phosphorylated residues may be protected (e.g., by β-elimination of phosphate groups and optional Michael addition of thiol groups, for example, as described in Knight et al., Nat. Biotechnology. 21, 1047-1054 (2003), the full text of which is incorporated herein by reference), etc. The polymer analyte or monomer can be modified with a protecting group or part thereof, such as methyl, formyl, ethyl, acetyl, tert-butyl, anisyl, benzyl, trifluoroacetyl, N-hydroxysuccinimide, tert-butoxycarbonyl (Boc), benzoyl, 4-methylbenzyl, thioanizyl, thiotoluyl, benzyloxymethyl, 4-nitrophenyl, benzyloxycarbonyl, 2-nitrobenzoyl, 2-nitrophenylsulfinyl, 4-toluenesulfonyl, pentafluorophenyl, diphenylmethyl, 2-chlorobenzyloxycarbonyl, 2,4,5-trichlorophenyl, 2-bromobenzyloxycarbonyl, 9-fluorenylmethoxycarbonyl (FMOC), triphenylmethyl, or 2,2,5,7,8-pentamethyl-chromium-6-sulfonyl group. Polymer analytes or monomers can be treated with protective agents, such as carboxyethyl methyl thiosulfonate (CEMTS), thiazolyl, mercaptophenylacetic acid, cyanobenzothiazole (e.g., for the esterification of N-terminal cysteine), acetamidomethyl, 2-methylsulfonyl ethyl-oxycarbonyl, etc.In some cases, isothiocyanates (e.g., PITC) can be used to block lysine residues (e.g., to enable the primary amine reaction of lysine residues), and optionally a single round of Edman degradation can be performed to generate new N-terminal exposed ends.

[0353] In some cases, monomers of polymer analytes can be modified to promote the cleavage of monomers from the polymer analyte. For example, amino acid monomers in peptide polymer analysis can be modified to make them recognizable by enzymes, such as by acetylation of amino acids, which can promote the cleavage of acetylated amino acids by acylpeptides. Additional or alternative modifications to monomers, such as those described herein, can also promote recognition by engineered lyases or interaction with engineered lyases.

[0354] In some cases, monomers containing naturally occurring modifications can be treated to remove or alter these modifications, making the polymer analyte or monomer more suitable for the processing operations disclosed herein. For example, acetylation, formylation, methylation, and post-translational modifications of pyrrolidone carboxylic acid (PCA) can be removed prior to sequencing. Acetylation modifications can be removed using acylpeptidases or acid treatment (e.g., using 1N HCl). Methylation can be removed using aminopeptidases. Formylation modifications can be removed, for example, using acid treatment (e.g., 0.6M HCl). Pyrrolidone carboxylic acid (PCA) can be removed using pyroglutamic acid aminopeptidase. Exemplary C-terminal modifications may include amidation and methylation, both of which can be removed using carboxypeptidases.

[0355] Binder: A binder that allows contact between the binder and the monomer (e.g., after cleavage and monomer-capture moiety coupling). The binder can be any useful molecule that can be coupled to a monomer or monomer-capture moiety complex. For example, as described herein, the binder can be or include proteins or peptides (e.g., antibodies, antibody fragments, single-stranded variant fragments (scFv), nanobodies, anti-carrier proteins (anticalin), tRNA synthetases or tRNA-acyl synthetases, fibronectin domains), peptide mimics, peptide pseudomorphs (e.g., peptide-like substances, β-peptides, D-peptide pseudomorphs), polysaccharides, nucleic acid molecules (e.g., aptamers), somamers, polymers, inorganic compounds, organic compounds, small molecules or derivatives (e.g., engineered variants), or combinations thereof. In cases where the polymer analyte includes a peptide, the binder can be able to bind to a modified or derivatized amino acid (e.g., an amino acid coupled to a linker) or a portion thereof. The binder may contain a recognition site that specifically recognizes an amino acid, a modified amino acid (e.g., an amino acid that binds to a linker containing a PITC moiety), or a derivatized (and optionally modified) amino acid. For example, the binder may be configured to recognize or bind to a portion of a modified amino acid, such as a specific amino acid residue, a residue-linker complex, or a derivatized amino acid (e.g., a thiocarbamoyl-derived residue, a thiazolyl-derived residue, a thiohydantoin-derived residue, etc.) or a portion of a modified amino acid. In some cases, the binder may be derived from or engineered from a naturally occurring enzyme or protein, such as aminopeptidase, exopeptidase, metalloproteinase, antibody, anticarrier protein, N-recognition protein, Clp protease, endopeptidase (e.g., trypsin), or tRNA synthetase. In some examples, the binder may be a lyase (e.g., trypsin, endopeptidase) that has been modified to remove peptidase activity. The binder can also recognize terminal amino acids attached to the substrate; for example, after all monomers except the final monomer of the polymer analyte have been coupled to one or more capture moieties and cleaved, the final monomer can remain coupled to the substrate. Therefore, the binder can recognize and bind to the substrate-coupled monomer.

[0356] The conjugate can be monospecific, bispecific, trispecific, or specific to a variety of different molecules. In some cases, the conjugate may contain multiple binding domains or portions that can recognize and bind to different molecules. Alternatively or otherwise, the conjugate may have multiple binding domains or portions that bind to the same target, such as bivalent, trivalent, or multivalent conjugates. In some embodiments, the conjugate is a bispecific antibody.

[0357] The binder can contact and specifically bind to cleaved monomers, monomer-joint complexes, monomer-joint-capture moieties, or monomer-capture moieties (collectively referred to herein as “monomer analytes”). For example, monomer analytes can fall into any size or size range smaller than any size or size range of the entire polymer analyte. The size of monomer analyte complexes can be about 0.1 nanometers (nm), about 0.5 nm, about 1 nm, about 5 nm, about 10 nm, about 20 nm, about 30 nm, about 40 nm, about 50 nm, about 60 nm, about 70 nm, about 80 nm, about 90 nm, about 100 nm, about 100 nm, about 1 mm, or larger. Monomer analytes can have any molecular weight or molecular weight range. Monomer analytes can be approximately 1 Dalton (Da), 10 Da, 100 Da, 500 Da, 1 kilodalton (kDa), 10 kDa, 100 kDa, 1,000 kDa, 10,000 kDa, 100,000 kDa, or greater. The molecular weight or length of monomer analytes can vary, for example, based on amino acid residues.

[0358] A conjugate may contain or be directly or indirectly coupled to a polymerizable molecule. The polymerizable molecule may be the same type of molecule as the conjugate (e.g., two peptides, two nucleic acid molecules, etc.), or they may be different. In some cases, the conjugate contains a peptide (e.g., an antibody or antibody fragment), and the polymerizable molecule comprises a nucleic acid molecule. In some cases, both the conjugate and the polymerizable molecule comprise proteins (e.g., an antibody and a fluorescent protein, two conjugated antibodies, such as a bispecific antibody).

[0359] Polymerizable molecules can be conjugated with binders via chemical conjugation methods, such as using connectors such as SMCC, (Ne-maleimide hexanoyloxy)succinimide ester (EMCS), succinimide-4-(p-maleimide phenyl)butyrate ester (SMPB), succinimide-(N-maleimide propionamido-ethylene glycol) ester (SMPEG), succinimide (NHS) ester, succinimide-4-formylbenzamide (S-4FB), succinimide-6-hydrazinoamide (S-HyNic), 4-phenyl-3H-1,2,4-triazolin-3,5(4H)dione (PTAD) or other diazo, 1-ethyl-3-3-dimethylaminopropyl carbodiimide hydrochloride (EDC). The synthesis of peptide-nucleic acid conjugates can also be carried out using solid-phase synthesis, fragment conjugation (e.g., using heterobifunctional crosslinking agents, such as those containing an aliphatic chain and a maleimide group at one end and NHS at the other end), click chemistry (e.g., strain-promoted azide alkyne cycloaddition, reverse electron-demanding Diels-Alder reaction), or a combination of methods or chemical actions. In some cases, enzymatic methods can be used to conjugate polymerizable molecules with binding agents. For example, truncated nucleases (e.g., Cas proteins, such as Cas9), relaxants (e.g., VirD2) or other enzymes, ribozymes, or DNases can be used to generate DNA-protein conjugates. In some cases, SpyTag and SpyCatcher interactions (e.g., generating binding agents containing a fusion protein containing a SpyCatcher and conjugated to a polymerizable molecule with a SpyTag), biotin-antibiotin interactions, SNAP-tags, or other interactions can be used to conjugate polymerizable molecules with binding agents. Optional purification can be performed, for example, using ion-exchange chromatography, affinity chromatography, HPLC, or other purification techniques.

[0360] The binder can be coupled to the polymerizable molecule via non-covalent interactions. For example, the binder can contain an avidin or streptavidin tag to which a biotin-conjugated polymerizable molecule can bind. Alternatively, the binder can contain a biotin tag to which an avidin or streptavidin-conjugated polymerizable molecule can bind.

[0361] Polymerizable molecules can contain identification information for the binder. For example, polymerizable molecules can include nucleic acid barcode molecules containing a barcode sequence. The barcode sequence can encode the identity of the binder or binding partner. For example, monomers (e.g., amino acids) of polymeric analytes (e.g., peptides containing multiple amino acids) can be cleaved and conjugated to a capture moiety (e.g., on a substrate) and can be contacted with a binder (e.g., antibody, antibody fragment, nanobody). The binder can specifically recognize amino acid residues or their derivatives (e.g., PTH, PTC, ATZ derivatized forms) rather than other amino acid residues or their derivatives. Nucleic acid barcode molecules can contain information identifying the binder, which can also identify specific amino acid residues (or derivatives) due to the specificity of the binder for its target.

[0362] In some cases, the binder contains or is conjugated with a guide RNA (gRNA) (such as guide editing guide RNA (pegRNA)). For example, the binder may include a barcode molecule containing pegRNA. The pegRNA may contain a barcode sequence that identifies the binder or a homologous molecule of the binder (e.g., monomer identity or type). The gRNA may contain useful functional sequences, such as key sequences (e.g., key sequences configured to insert into a tandem array of CRISPR-Cas9 target sites), UMIs, barcode sequences, primer sites, restriction sites, cleavable portions, base-free sites, transposition sites, spacer portions, etc. In some cases, the gRNA contains a propagator sequence containing a CRISPR-Cas9 target site. The propagator sequence may be configured to insert into a polymerizable molecule containing an additional target site and a motif adjacent to the prototype spacer region. In some cases, the propagator sequence is configured to insert between a prototype spacer region motif and an additional target site in a polymerizable molecule.

[0363] The polymerizable molecule of the binder may contain additional multiplexing information. For example, the polymerizable molecule (e.g., a nucleic acid molecule) may contain sequences encoding cycling or other temporal or spatial information. In one such example, an array of peptides and capture moieties may be provided on a substrate. The array may include multiple individual addressable units, each of which (or a subset thereof) contains the peptide and capture moieties to be analyzed. The binder and the polymerizable molecule contained therein or coupled thereto may contain spatial information (e.g., spatial barcode sequences) that uniquely identify the individual addressable units, thereby identifying the location of the array. The polymerizable molecule may additionally contain temporal information (e.g., a cycling barcode indicating the round or iteration of the binder or polymerizable molecule provided). Subsequent sequencing of the polymerizable molecule may be used to reveal spatial information (e.g., the origin location in the peptide or amino acid array). In some cases, the polymerizable molecule may contain a unique molecular identifier (UMI) that can be used to determine the amount of a given binder or monomer (e.g., an amino acid) in a given peptide, substrate, array, or sample. In some cases, polymerizable molecules can contain one or more detectable markers, such as fluorophores, radioisotopes, mass tags, etc. For example, polymerizable molecules can contain nucleic acid molecules or peptides containing multiple fluorophores. The unique combination of fluorophores can be used as a barcode or to encode additional information into the polymerizable molecule.

[0364] Alternatively, the binder recognizing the monomer-capture moiety complex may not contain or be coupled to a polymerizable molecule. In such cases, after the binder binds to the monomer-capture moiety complex, an additional molecule (e.g., a second binder) containing a detectable marker (e.g., a fluorophore, a radioisotope, a mass tag, or identifying a polymerizable molecule (e.g., a nucleic acid barcode molecule)) may contact and bind to the binder bound to the monomer-capture moiety complex. In some examples, the additional molecule includes identifying a polymerizable molecule, and the identifying polymerizable molecule may be coupled to or transferred to the additional polymerizable molecule. In a non-limiting example, the binder contains a primary antibody or antibody fragment recognizing the monomer-capture moiety complex (e.g., a terminal amino acid-linker-capture moiety complex) or a portion thereof (e.g., a terminal amino acid or a terminal amino acid-linker complex); after the primary antibody or antibody fragment binds to the monomer-capture moiety complex or a portion thereof, a secondary antibody or antibody fragment containing a polymerizable molecule (e.g., a nucleic acid barcode molecule) or coupled thereto is coupled to the primary antibody. The polymerizable molecule of the secondary antibody or antibody fragment may contain information or other information about the secondary antibody or antibody fragment, the primary antibody or antibody fragment. The transfer of the polymerizable molecule of the secondary antibody or antibody fragment to another polymerizable molecule or the conjugation with another polymerizable molecule can be mediated by any suitable technique, such as hybridization of nucleic acid molecules optionally mediated by splice molecules, click chemistry, or association of high-affinity molecules (e.g., streptavidin and biotin).

[0365] In some cases, the method may include contacting the monomer-capture portion complex with a binding agent library. The binding agent library may contain a variety of binding agents specific to different analytes. For example, the binding agent library may contain a variety of binding agents that recognize different amino acids or their derivatives (e.g., derivatized amino acids, such as PTH, PTC, or ATZ forms), amino acid clusters (e.g., dipeptides, tripeptides, etc.), or combinations of amino acids (e.g., amino acids with similar side chain groups). In one such example, a given binding agent may optionally recognize and bind to more than one amino acid with different affinities or binding kinetics. A given binding agent may recognize and bind to a single amino acid, two different amino acids, three different amino acids, four different amino acids, etc. For example, a given binder may bind to amino acids having similar residues, such as amino acids with positively charged side chains (e.g., arginine, histidine, lysine), amino acids with negatively charged side chains (aspartic acid, glutamic acid), amino acids with polar, uncharged side chains (e.g., serine, threonine, asparagine, glutamine), amino acids with hydrophobic side chains (e.g., alanine, valine, isoleucine, leucine, methionine, phenylalanine, tyrosine, tryptophan), or combinations thereof. In summary, the binding agent library can specifically recognize or bind to any number of different amino acids; for example, the binding agent library can be configured to specifically bind to at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 11, at least 12, at least 13, at least 14, at least 15, at least 16, at least 17, at least 18, at least 19, or at least 20 different protein amino acids or their derivatives.

[0366] A binding agent library can contain any useful number of binding agents, each of which can have different binding specificities. For example, a first binding agent may recognize one amino acid, a second binding agent may recognize two amino acids, and a third binding agent may recognize three amino acids. In another example, a first binding agent may recognize one amino acid, a second binding agent may recognize different amino acids, and a third binding agent may recognize multiple amino acids. It should be understood that any number of binding agents can be used, and each binding agent may be specific to one or more amino acids. In summary, a binding agent library can bind to all 20 protein amino acids or their derivatives, or subgroups of amino acids or their derivatives (e.g., 10 or more, 15 or more). Similarly, multiple binding agents can bind to the same amino acid or its derivative. For example, one or more binding agents that recognize at least one type of amino acid or its derivative can be used to query a single amino acid or its derivative multiple times, which can be used to improve the accuracy of amino acid identification. A binding agent library can contain binding agents having combinations of the amino acid sequences described herein (e.g., in Tables 1 to 8).

[0367] The binding agent library can be provided sequentially, in parallel, or in combination. For example, for sequential binding, for a given cyclic iteration, a first binding agent targeting a specific monomer analyte (e.g., a specific amino acid type or derivative thereof) can be provided to query the monomer analyte. The interaction between the first binding agent and its homologous molecules can be recorded. Optionally, the first binding agent can be washed away or removed. Subsequently, a second binding agent for another monomer analyte (e.g., another specific amino acid type or derivative thereof) can be provided. In another example, for parallel binding, the first and second binding agents can be provided simultaneously, for example, in the form of a mixture. In some cases, several sets of binding agents can be provided sequentially; for example, a first set of binding agents containing two or more binding agents that bind to different monomer analytes (e.g., different amino acid types or derivatives thereof) can be provided first, and the first set of binding agents can be bound to their respective targets; optional washing and / or purification can be performed, and then a second set of binding agents containing two or more binding agents that bind to different monomer analytes (e.g., different amino acid types or derivatives thereof) can be provided. This process can be iterated or repeated any useful number of times. In some cases, a group of binders may contain binders that bind to the same target; for example, a first group of binders may contain a first binder that binds to a first amino acid type and a second binder that binds to both the first amino acid type and the second amino acid type.

[0368] Combinatorial decoding using multispecific binders: Identification of multiple amino acid types may require different numbers of binders to identify the number of amino acid types. In some cases, a combination of binders specific to multiple amino acid types can be used to identify all 20 amino acid types. In one example, for N different types of amino acids in a sample or peptide, where N is an integer, M different binders can be provided, where M is also an integer. Each of the M different binders can identify one or more amino acid types, such that a combination of all M different binders can identify N amino acid types. In some cases, a peptide containing multiple amino acids, including N different types of amino acids, can be provided. The peptide or cleaved amino acids (e.g., locally tethered cleaved amino acids) can be contacted with a library of M different binders (e.g., antibodies or antibody fragments), where at least one binder in the library of M different binders binds to more than one amino acid type. The interaction between at least one of the M different binders and the amino acid can be recorded (e.g., using the polymerizable molecules described above or detecting detectable signals, such as fluorescence from a fluorophore conjugated with the binder). In some cases, multiple interactions can be recorded within a given cycle, for example, by coupling a polymerizable molecule of the binder to a polymerizable molecule located near the peptide, or by detecting a detectable signal (e.g., fluorescence, mass tagging, radioisotope) to record individual binding events of the same or different binders. In other cases, multiple binding events of a given cycle can be recorded by washing and re-probing (e.g., by providing a library of the same or different binders). By recording the interactions between M different binders and amino acids, the binding patterns of these M different binders can be generated, which can then be used to identify N different amino acid types.

[0369] Figure 3This diagram schematically illustrat...

Claims

1. A binding agent that binds to one type of protein amino acid or a derivative thereof in the group with higher specificity or higher affinity than with all other types of protein amino acids or derivatives thereof in the group of two or more protein amino acids or derivatives thereof, wherein the one type of amino acid is not tryptophan.

2. The binder according to claim 1, wherein the group of two or more protein amino acids or their derivatives comprises at least 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19 or 20 different protein amino acids or their derivatives.

3. The binder according to claim 1 or 2, wherein the group consisting of two or more protein amino acids or derivatives thereof comprises naturally occurring protein amino acids.

4. The binder according to any one of claims 1 to 3, wherein the group of two or more protein amino acids or their derivatives comprises 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19 or 20 standard protein amino acids or their derivatives.

5. The binder according to any one of claims 1 to 3, wherein the group of two or more protein amino acids or their derivatives comprises 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19 or 20 standard protein amino acids or their derivatives.

6. The binding agent according to any one of claims 1 to 5, wherein the binding agent is part of a binding agent library, wherein the binding agent library comprises a plurality of binding agents that collectively recognize and specifically bind to at least 5 different types of protein amino acids.

7. The binding agent of claim 6, wherein the binding agent is part of a binding agent library, wherein the binding agent library comprises a variety of binding agents that collectively recognize and specifically bind to all 20 types of standard protein amino acids.

8. The binding agent according to claim 7, wherein the binding agent of the binding agent library is capable of recognizing at least one post-translational modified amino acid and specifically binding to said at least one post-translational modified amino acid.

9. The binding agent according to any one of claims 6 to 8, wherein the binding agent library comprises a plurality of polyclonal antibodies.

10. The binding agent according to any one of claims 6 to 9, wherein the binding agent library comprises a plurality of monoclonal antibodies.

11. The binder according to any one of claims 1 to 10, wherein the binder comprises an antibody, an antibody fragment, a single-chain variable fragment (scFv), or a nanobody.

12. The binder of claim 11, wherein the binder comprises an antibody.

13. The binder according to claim 12, wherein the antibody is immunoglobulin G (IgG).

14. The binder according to claim 12 or 13, wherein the antibody is generated by immunizing an animal with an immunogenically effective amount of a composition comprising a carrier protein coupled to an amino acid or an amino acid derivative.

15. The binder of claim 14, wherein the composition comprises the carrier protein, the carrier protein comprising a polymeric linker coupled to the amino acid or an amino acid derivative.

16. The binder of claim 15, wherein the polymeric joint comprises polyethylene glycol (PEG).

17. The binding agent according to any one of claims 12 to 16, wherein the antibody is a monoclonal antibody.

18. The binder according to claims 1 to 17, wherein the binder comprises engineered scFv.

19. The binder of claim 18, wherein the engineered scFv is derived from IgG.

20. The binder according to claim 18 or 19, wherein the engineered scFv is engineered using directed evolution.

21. The binder of claim 20, wherein the engineered scFv undergoes affinity maturation using yeast surface display, bacteriophage display, ribosome display, or sequential evolution.

22. The binder according to claim 20 or 21, wherein the engineered scFv exhibits improved specificity for the type of protein amino acid or its derivatives by using yeast surface exposure.

23. The binder according to any one of claims 1 to 22, wherein the derivative or the derivative thereof comprises chemically modified amino acids.

24. The binder according to claim 23, wherein the chemically modified amino acid comprises hydantoin, phenylthiocarbamoyl, or aniline thiazolinone derivatives of the amino acid.

25. The binder according to claim 23 or 24, wherein the chemically modified amino acid comprises an amino acid coupled to a linker.

26. The binder of claim 25, wherein the linker is coupled to the amino acid at the N-terminus.

27. The binder of claim 26, wherein the linker coupled to the chemically modified amino acid comprises a hydantoin, phenylthiocarbamoyl, or anilinethiazolinone moiety.

28. The binder of claim 25, wherein the linker is coupled to the amino acid at the C-terminus.

29. The binder according to any one of claims 25 to 28, wherein the binder comprises a nucleic acid molecule.

30. The binder according to any one of claims 25 to 28, wherein the linker comprises an amino acid reactive group and a linker nucleic acid molecule.

31. The binder according to any one of claims 1 to 30, wherein the binding of the binder to a specific type of protein amino acid or a derivative thereof is measured using an indirect enzyme-linked immunosorbent assay (ELISA), rather than to all other types of protein amino acids or derivatives thereof in the group.

32. The binder according to any one of claims 1 to 31, wherein the surface plasmon resonance is used to measure the specific binding of the binder to the one type of protein amino acid or its derivative thereof, rather than to all other types of protein amino acids or their derivatives in the group.

33. The binding agent according to any one of claims 1 to 32, wherein the binding agent is specifically bound to the one type of protein amino acid or its derivative thereof, rather than to all other types of protein amino acids or their derivatives in the group, as measured by biomembrane interferometry.

34. The binder according to any one of claims 1 to 33, wherein the binder comprises a detectable marker.

35. The binder of claim 34, wherein the detectable label comprises a nucleic acid molecule, a fluorophore, a mass tag, or a protein.

36. The binder of claim 35, wherein the detectable marker comprises a protein, wherein the protein comprises an additional antibody or an additional antibody fragment.

37. The binder of claim 35, wherein the detectable marker comprises a nucleic acid molecule, wherein the nucleic acid molecule comprises a barcode sequence.

38. The binder according to any one of claims 34 to 37, wherein the detectable marker is coupled to the binder using a portion selected from the group consisting of: biotin, desulfobiotin, avidin, streptavidin, neutral avidin, SpyCatcher, SpyTag, SNAP tag, click chemistry portion, and cysteine ​​tag.

39. The binder according to any one of claims 34 to 38, wherein the binder comprises atypical amino acids, wherein the detectable label is coupled to the binder via the atypical amino acids.

40. The binder of claim 39, wherein the atypical amino acid comprises a click chemistry moiety.

41. The binding agent according to any one of claims 1 to 40, wherein the binding agent specifically binds to phenylalanine or a derivative thereof, rather than specifically binding to all other types of protein amino acids or derivatives thereof in the group.

42. The binding agent according to any one of claims 1 to 40, wherein the binding agent specifically binds to leucine or a derivative thereof, rather than specifically binding to all other types of protein amino acids or derivatives thereof in the group.

43. The binding agent according to any one of claims 1 to 40, wherein the binding agent specifically binds to valine or a derivative thereof, rather than specifically binds to all other types of protein amino acids or derivatives thereof in the group.

44. The binding agent according to any one of claims 1 to 40, wherein the binding agent specifically binds to tyrosine or a derivative thereof, rather than specifically binding to all other types of protein amino acids or derivatives thereof in the group.

45. The binding agent according to any one of claims 1 to 40, wherein the binding agent specifically binds to proline or a derivative thereof, rather than specifically binds to all other types of protein amino acids or derivatives thereof in the group.

46. ​​The binder according to any one of claims 1 to 45, wherein the binder is derived from mice.

47. An antibody fragment that binds to one type of protein amino acid or a derivative thereof in the group with higher specificity or higher affinity than with all other types of protein amino acids or derivatives thereof in the group of two or more protein amino acids or derivatives thereof.

48. The antibody fragment of claim 47, wherein the group of two or more protein amino acids or their derivatives comprises at least 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19 or 20 different protein amino acids or their derivatives.

49. The antibody fragment according to claim 47 or 48, wherein the group consisting of two or more protein amino acids or derivatives thereof comprises naturally occurring protein amino acids.

50. The antibody fragment according to any one of claims 47 to 49, wherein the group of two or more protein amino acids or their derivatives comprises 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19 or 20 standard protein amino acids or their derivatives.

Citation Information

Patent Citations

  • Single-molecule protein and peptide sequencing

    US11499979B2

  • Single-Molecule Protein and Peptide Sequencing

    US20200217853A1

  • Methods of generating nanoarrays and microarrays

    WO2019195633A1

  • Single-molecule peptide sequencing through molecular barcoding and ex-SITU analysis

    WO2023114732A2

  • Methods and systems for processing polymeric analytes

    WO2023196642A1