Methods and compositions for protein sequencing

Recombinant amino acid-binding proteins and peptidases facilitate accurate polypeptide sequencing by detecting and degrading proteins to determine amino acid sequences through real-time signal pulses, addressing the complexity of protein sequencing.

JP2026065001APending Publication Date: 2026-04-14QUANTUM SI INC
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
QUANTUM SI INC
Filing Date
2025-12-17
Publication Date
2026-04-14

AI Technical Summary

Technical Problem

The complexity of protein structure, composition, and modifications poses challenges in determining large-scale protein sequencing information for biological samples.

Method used

The use of recombinant amino acid-binding proteins and peptidases, along with amino acid recognition molecules, in a polypeptide sequencing reaction mixture to determine amino acid sequence information by detecting interactions and degrading polypeptides while monitoring binding and cleavage events.

Benefits of technology

Enables accurate and efficient sequencing of polypeptides, including those with modifications, by detecting sequential amino acids through real-time signal pulses, facilitating detailed protein analysis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026065001000001_ABST
    Figure 2026065001000001_ABST
Patent Text Reader

Abstract

The present invention provides methods for identifying and sequencing proteins, polypeptides, and amino acids, as well as compositions useful for them. [Solution] The present invention provides amino acid recognition molecules, such as amino acid-binding proteins and their fusion polypeptides, which differentially associate with different types of amino acids to generate detectable characteristic signatures that serve as indicators of the amino acid sequence of polypeptides. The present invention also provides amino acid recognition molecules that include a shielding element to enhance photostability during polypeptide sequencing reactions.
Need to check novelty before this filing date? Find Prior Art

Description

[Background technology]

[0001] Proteomics has emerged as a crucial and necessary complement to genomics and transcriptomics in the study of biological systems. Proteomic analysis of individual organisms can provide insights into cellular processes and response patterns, leading to improved diagnostic and therapeutic strategies. The complexity surrounding protein structure, composition, and modifications presents challenges in determining large-scale protein sequencing information for biological samples. [Overview of the Initiative]

[0002] In some embodiments, the present application provides methods and compositions for determining amino acid sequence information from polypeptides (for example, for sequencing one or more polypeptides). In some embodiments, amino acid sequence information can be determined for a single polypeptide molecule. In some embodiments, for example, the relative positions of two or more amino acids of a polypeptide are determined for a single polypeptide molecule. In some embodiments, one or more amino acids of a polypeptide are labeled (for example, directly or indirectly) to determine the relative positions of the labeled amino acids of the polypeptide. In some embodiments, amino acid sequence information can be determined by detecting interactions between a polypeptide and one or more amino acid recognition molecules (for example, one or more amino acid-binding proteins).

[0003] In some embodiments, the present application provides amino acid-binding proteins that can be used in methods for determining amino acid sequence information from polypeptides. In some embodiments, the present application provides recombinant amino acid-binding proteins having an amino acid sequence that is at least 80% identical to a sequence selected from Table 1 or Table 2 and comprising one or more labels. In some embodiments, one or more labels include luminescence labels or conductivity labels. In some embodiments, one or more labels include a tag sequence. In some embodiments, the tag sequence includes one or more of a purified tag, a cleavage site, and a biotinylated sequence (e.g., at least one biotin ligase recognition sequence). In some embodiments, the biotinylated sequence includes two biotin ligase recognition sequences oriented in series. In some embodiments, one or more labels include a biotin moiety having at least one biotin molecule (e.g., a bisbiotin moiety). In some embodiments, the label includes at least one biotin ligase recognition sequence to which at least one biotin molecule is attached. In some embodiments, one or more labels include one or more polyol moieties (e.g., polyethylene glycol). In some embodiments, the recombinant amino acid-binding protein comprises one or more non-natural amino acids to which one or more labels are attached. In some embodiments, the present application provides compositions comprising the recombinant amino acid-binding protein described herein.

[0004] In some embodiments, the Application provides a polypeptide sequencing reaction composition comprising two or more amino acid recognition molecules, wherein at least one of the two or more amino acid recognition molecules is a recombinant amino acid-binding protein as described herein. In some embodiments, the two or more amino acid recognition molecules comprise amino acid recognition molecules of different types from each other. For example, in some embodiments, one type of amino acid recognition molecule interacts with the target polypeptide in a different manner (e.g., distinctly different) from the other types of amino acid recognition molecules in the polypeptide sequencing reaction composition. In some embodiments, the polypeptide sequencing reaction composition comprises at least one type of cleavage reagent. In some embodiments, the Application provides a method of polypeptide sequencing comprising contacting a polypeptide with a polypeptide sequencing reaction composition as described herein. In some embodiments, the method further comprises sequencing the polypeptide by detecting a series of interactions between the polypeptide and at least one amino acid recognition molecule while the polypeptide is being degraded.

[0005] In some embodiments, the present application provides polypeptide sequencing reaction mixtures comprising an amino acid-binding protein and a peptidase. In some embodiments, the molar ratio of the labeled amino acid-binding protein to the peptidase is about 1:1,000 to about 1:1 or about 1:1 to about 100:1. In some embodiments, the amino acid-binding protein includes one or more labels. In some embodiments, the amino acid-binding protein is a ClpS protein. In some embodiments, the amino acid-binding protein is a protein having at least 80%, 80-90%, 90-95%, or at least 95% identical amino acid sequences to sequences selected from Table 1 or Table 2. In some embodiments, the peptidase is an exopeptidase. In some embodiments, the peptidase is an enzyme having at least 80%, 80-90%, 90-95%, or at least 95% identical amino acid sequences to sequences selected from Table 4 or Table 5. In some embodiments, the reaction mixture comprises one or more amino acid-binding proteins and / or one or more peptidases. In some embodiments, the reaction mixture includes peptide molecules immobilized on a surface.

[0006] In some embodiments, the present application provides polypeptide sequencing reaction mixtures comprising a single peptide molecule, at least one peptidase molecule, and at least three amino acid recognition molecules. In some embodiments, the reaction mixture comprises at least one and up to ten peptidase molecules (e.g., at least one and up to five peptidase molecules, at least one and up to three peptidase molecules). In some embodiments, the reaction composition comprises two or more peptidase molecules, each peptidase molecule being of a different type from the others. For example, in some embodiments, one type of peptidase molecule has a different cleavage priority than the other types of peptidase molecules in the reaction mixture. In some embodiments, the reaction mixture comprises at least three and up to 30 amino acid recognition molecules (e.g., up to 20, up to 10, or up to five amino acid recognition molecules). In some embodiments, the at least three amino acid recognition molecules comprise amino acid recognition molecules of different types from the others. For example, in some embodiments, one type of amino acid recognition molecule interacts with the target polypeptide in a way that is different (e.g., detectably different) from other types of amino acid recognition molecules in the reaction mixture.

[0007] In some embodiments, the present application provides a substrate comprising an array of sample wells, wherein at least one sample well of the array contains a polypeptide sequencing reaction mixture described herein. In some embodiments, at least one sample well comprises a bottom surface. In some embodiments, a single polypeptide molecule is immobilized on the bottom surface.

[0008] In some embodiments, the present application provides an amino acid recognition molecule comprising a polypeptide having at least a first amino acid-binding protein and a second amino acid-binding protein linked to each other at their termini, wherein the first amino acid-binding protein and the second amino acid-binding protein are separated by a linker comprising at least two amino acids. In some embodiments, the first amino acid-binding protein and the second amino acid-binding protein are the same. In some embodiments, the first amino acid-binding protein and the second amino acid-binding protein are different from each other.

[0009] In some embodiments, the present application provides a compound of formula (I): (Z 1 -X 1 ) n -Z 2 (I) (wherein Z 1 and Z 2 are independent amino acid-binding proteins, X 1 is a linker comprising at least two amino acids, the amino acid-binding proteins are linked to each other at their termini by the linker, and n is an integer from 1 to 5 (including both end values)). In some embodiments, Z 1 and Z 2 comprise the same type of amino acid-binding protein. In some embodiments, Z 1 and Z 2 comprise different types of amino acid-binding proteins. In some embodiments, Z 1 and Z 2 are independently optionally associated with a label component comprising at least one detectable label. In some embodiments, the polypeptide further comprises a tag sequence.

[0010] In some embodiments, the present application provides a method for polypeptide sequencing. In some embodiments, the polypeptide sequencing method involves contacting a single polypeptide molecule in a reaction mixture with a composition comprising binding and cleaving means. In some embodiments, the binding and cleaving means are configured to achieve at least 10 association events between the binding means and terminal amino acids on the polypeptide before the terminal amino acids are removed from the peptide by the cleaving means. In some embodiments, the binding and cleaving means are configured to achieve at least 10 and up to 1,000 association events before the removal of the terminal amino acids. In some embodiments, the terminal amino acids are exposed at the end of the polypeptide by a cleaving event prior to at least 10 association events. In some embodiments, at least 10 association events occur after the cleaving event.

[0011] In some embodiments, the binding and cleaving means are configured to achieve a time interval of at least 1 minute between cleavage events (e.g., about 1 minute to about 20 minutes, about 5 minutes to about 15 minutes, or about 1 minute to about 10 minutes). In some embodiments, the binding means comprises one or more amino acid recognition molecules, and the cleaving means comprises one or more peptidase molecules. In some embodiments, the molar ratio of amino acid recognition molecules to peptidase molecules is configured to achieve at least 10 association events before the removal of terminal amino acids. In some embodiments, the molar ratio of amino acid recognition molecules to peptidase molecules is about 1:1,000 to about 1:1 or about 1:1 to about 100:1. In some embodiments, the molar ratio of amino acid recognition molecules to peptidase molecules is about 1:100 to about 1:1 or about 1:1 to about 10:1.

[0012] In some embodiments, the present application provides a substrate comprising an array of sample wells, wherein at least one sample well of the array comprises a single polypeptide molecule, a cleaving means, and a binding means. In some embodiments, the binding means and the cleaving means are configured to achieve at least 10 association events between the binding means and a terminal amino acid on the polypeptide before the terminal amino acid is removed from the peptide by the cleaving means. In some embodiments, the binding means and the cleaving means are configured to achieve at least 10 and up to 1,000 association events before the removal of the terminal amino acid. In some embodiments, the terminal amino acid is terminally exposed to the polypeptide by a cleaving event prior to at least 10 association events. In some embodiments, at least 10 association events occur after the cleaving event.

[0013] In some embodiments, the present application provides an amino acid recognition molecule comprising a shielding element for, for example, enhancing photostability during polypeptide sequencing reactions. In some embodiments, the present application provides an amino acid recognition molecule comprising a polypeptide having a terminally linked amino acid-binding protein and a labeled protein. In some embodiments, the amino acid-binding protein and the labeled protein are separated by a linker containing at least two amino acids (e.g., at least two and up to 100 amino acids, about 5 to about 50 amino acids). In some embodiments, the labeled protein has a molecular weight of at least 10 kDa (e.g., about 10 kDa to about 150 kDa, about 15 kDa to about 100 kDa). In some embodiments, the labeled protein contains at least 50 amino acids (e.g., about 50 to about 1,000 amino acids, about 100 to about 750 amino acids). In some embodiments, the labeled protein includes a luminescent label. In some embodiments, the luminescent label contains at least one fluorophore dye molecule. In some embodiments, the amino acid-binding protein is a Gid protein, a UBR box protein or a UBR box domain-containing fragment thereof, a p62 protein or a ZZ domain-containing fragment thereof, or a ClpS protein. In some embodiments, the amino acid-binding protein has at least 80% amino acid sequence identity with an amino acid sequence selected from Table 1 or Table 2.

[0014] In some embodiments, the present application is formula (II): A-(Y) n -D (II) (In the formula, A is an amino acid binding component containing at least one amino acid recognition molecule; each of Y is a polymer that forms a covalent or non-covalent linking group; n is an integer from 1 to 10 (including the values ​​at both ends); and D is a labeling component containing at least one detectable label.) The present invention provides an amino acid recognition molecule. In some embodiments, A comprises at least one amino acid-binding protein having an amino acid sequence that is at least 80% identical to an amino acid sequence selected from Table 1 or Table 2. In some embodiments, the amino acid recognition molecule comprises A and Y linked at their terminal ends. 1 A polypeptide comprising A and Y 1 It is separated by a linker containing at least two amino acids. In some embodiments, Y 1 This is a protein having a molecular weight of at least 10 kDa (for example, about 10 kDa to about 150 kDa). In some embodiments, Y 1 A protein is a protein that contains at least 50 amino acids (for example, approximately 50 to 1,000 amino acids).

[0015] In some embodiments, D has a diameter of less than 20 nm (200 Å). In some embodiments, -(Y) n - is at least 2 nm in length (for example, at least 5 nm, at least 10 nm, at least 20 nm, at least 30 nm, at least 50 nm, or longer). In some embodiments, -(Y) n - is approximately 2 nm to approximately 200 nm in length (for example, approximately 2 nm to approximately 100 nm, approximately 5 nm to approximately 50 nm, or approximately 10 nm to approximately 100 nm). In some embodiments, each of Y is independently a biomolecule or dendritic polymer (e.g., polyol, dendrimer). In some embodiments, A comprises a polypeptide having at least a first amino acid-binding protein and a second amino acid-binding protein linked at their ends (e.g., a fusion polypeptide). In some embodiments, the present application provides a composition comprising an amino acid-recognizing molecule of formula (II). In some embodiments, the amino acid-recognizing molecule is soluble in the composition.

[0016] In some embodiments, the present application is formula (III): AY 1 -D (III) (wherein A is an amino acid binding component containing at least one amino acid recognition molecule; Y 1 (where D is a nucleic acid or polypeptide; D is a labeled component containing at least one detectable label.) The present invention provides an amino acid recognition molecule. In some embodiments, A comprises at least one amino acid-binding protein having an amino acid sequence that is at least 80% identical to an amino acid sequence selected from Table 1 or Table 2. In some embodiments, Y 1 When is a nucleic acid, the nucleic acid forms covalent or non-covalent linking groups. In some embodiments, Y 1 If it is a polypeptide, then the polypeptide is 50 × 10 -9 Dissociation constant less than M (K D ) forms a non-covalent linking group characterized by. In some embodiments, K D is 1 × 10 -9 Less than M, 1 x 10 -10 Less than M, 1 x 10 -11 Less than M, or 1 × 10 -12 It is less than M.

[0017] In some embodiments, the present application provides an amino acid recognition molecule comprising a nucleic acid, at least one amino acid recognition molecule attached to a first attachment site on the nucleic acid, and at least one detectable label attached to a second attachment site on the nucleic acid, wherein the nucleic acid forms a covalent or non-covalent linkage between the at least one amino acid recognition molecule and the at least one detectable label. In some embodiments, the nucleic acid comprises a first oligonucleotide chain. In some embodiments, the nucleic acid further comprises a second oligonucleotide chain hybridized to the first oligonucleotide chain. In some embodiments, the at least one amino acid recognition molecule comprises a polypeptide (e.g., a fusion polypeptide) having at least a first amino acid-binding protein and a second amino acid-binding protein linked at their terminals. In some embodiments, the first amino acid-binding protein and the second amino acid-binding protein are separated by a linker containing at least two amino acids.

[0018] In some embodiments, the present application provides an amino acid recognition molecule comprising a polyvalent protein having at least two ligand-binding sites, at least one amino acid recognition molecule attached to the protein via a first ligand moiety bound to a first ligand-binding site on the protein, and at least one detectable label attached to the protein via a second ligand moiety bound to a second ligand-binding site on the protein. In some embodiments, the polyvalent protein is an avidin protein. In some embodiments, the at least one amino acid recognition molecule comprises a polypeptide (e.g., a fusion polypeptide) having at least one first amino acid-binding protein and a second amino acid-binding protein linked at their terminal ends. In some embodiments, the first and second amino acid-binding proteins are separated by a linker containing at least two amino acids.

[0019] In some embodiments, the shield amino acid recognition molecule may be used in the polypeptide sequencing method according to the Application or in any method known in the Art. Therefore, in some embodiments, the Application provides a polypeptide sequencing method comprising contacting a polypeptide molecule with one or more shield amino acid recognition molecules of the Application (for example, by an Edman-type degradation reaction, a dynamic sequencing reaction, or in any other method known in the Art). For example, in some embodiments, the method comprises contacting a polypeptide molecule with at least one amino acid recognition molecule comprising a shield or shielding element according to the Application, and detecting the association of the at least one amino acid recognition molecule with the polypeptide molecule.

[0020] In some embodiments, the present application provides a method for obtaining data during the degradation process of a polypeptide. In some embodiments, the method further includes analyzing the data to determine portions of the data corresponding to amino acids sequentially exposed at the ends of the polypeptide during the degradation process. In some embodiments, the method further includes outputting the amino acid sequence corresponding to the polypeptide. In some embodiments, the data serves as an indicator of the amino acid identity at the end of the polypeptide during the degradation process. In some embodiments, the data serves as an indicator of a signal generated by one or more amino acid recognition molecules binding to different types of terminal amino acids at the end during the degradation process. In some embodiments, the data serves as an indicator of a luminescence signal generated during the degradation process. In some embodiments, the data serves as an indicator of an electrical signal generated during the degradation process.

[0021] In some embodiments, analyzing the data further includes detecting a series of cleavage events and determining the portion of the data between sequential cleavage events. In some embodiments, analyzing the data further includes determining the type of amino acid for each individual portion. In some embodiments, each individual portion includes a pulse pattern (e.g., a characteristic pattern), and analyzing the data further includes determining the type of amino acid for one or more portions based on their respective pulse patterns. In some embodiments, determining the type of amino acid further includes identifying the time quantity within a portion when the data exceeds a threshold and comparing the time quantity with the duration for a portion. In some embodiments, determining the type of amino acid further includes identifying at least one pulse duration for each of one or more portions. In some embodiments, the pulse pattern includes an average pulse duration of about 1 millisecond to about 10 seconds. In some embodiments, determining the type of amino acid further includes identifying at least one inter-pulse duration for each of one or more portions. In some embodiments, the amino acid sequence includes a series of amino acids corresponding to a portion.

[0022] In some embodiments, the present application provides a polypeptide sequencing method comprising contacting a single polypeptide molecule with one or more amino acid recognition molecules (e.g., one or more terminal amino acid recognition molecules). In some embodiments, the method further comprises obtaining sequence information about a single polypeptide molecule by detecting a series of signal pulses that indicate the association of one or more amino acid recognition molecules with sequential amino acids exposed at the terminals of the single polypeptide molecule while the single polypeptide is being degraded. In some embodiments, the amino acid sequence of almost all or all of the single polypeptide molecule is determined. In some embodiments, the series of signal pulses is a series of real-time signal pulses.

[0023] In some embodiments, the association of one or more amino acid recognition molecules with each type of amino acid exposed at the terminus generates a characteristic pattern of a series of signal pulses distinct from other types of amino acids exposed at the terminus. In some embodiments, the signal pulses of the characteristic pulses have an average pulse duration of about 1 millisecond to about 10 seconds. In some embodiments, the signal pulses of the characteristic pattern correspond to individual association events between the amino acid recognition molecules and the amino acids exposed at the terminus. In some embodiments, the characteristic pattern corresponds to a series of reversible amino acid recognition molecule binding interactions with the amino acids exposed at the terminus of a single polypeptide. In some embodiments, the characteristic pattern serves as an indicator of the amino acids exposed at the terminus of a single polypeptide molecule and the amino acids at consecutive positions (e.g., amino acids of the same type or different types from each other).

[0024] In some embodiments, a single polypeptide molecule is degraded by a cleavage reagent that removes at least one amino acid from the terminus of the single polypeptide molecule. In some embodiments, the method further includes detecting a signal that indicates the association of the cleavage reagent with the terminus. In some embodiments, the cleavage reagent includes a detectable label (e.g., an luminescent label, a conductivity label). In some embodiments, the single polypeptide molecule is immobilized on a surface. In some embodiments, the single polypeptide molecule is immobilized on a surface via a distal terminus from the terminus where one or more amino acid recognition molecules associate. In some embodiments, the single polypeptide molecule is immobilized on a surface via a linker (e.g., a solubilizing linker containing a biomolecule).

[0025] In some embodiments, the present application provides a method for sequencing a polypeptide, comprising contacting a single polypeptide molecule in a reaction mixture with a composition comprising one or more amino acid recognition molecules (e.g., one or more terminal amino acid recognition molecules) and a cleavage reagent. In some embodiments, the method further comprises detecting a series of signal pulses that indicate the association of one or more amino acid recognition molecules with the terminals of the single polypeptide molecule in the presence of the cleavage reagent. In some embodiments, the series of signal pulses are indicators of a series of amino acids that are exposed at the terminals over time as a result of terminal amino acid cleavage by the cleavage reagent.

[0026] In some embodiments, the present application provides a method for sequencing a polypeptide, comprising: (a) identifying a terminal first amino acid of a single polypeptide molecule; (b) removing the first amino acid to expose a terminal second amino acid of the single polypeptide molecule; and (c) identifying the terminal second amino acid of the single polypeptide molecule. In some embodiments, (a) to (c) are carried out in a single reaction mixture. In some embodiments, (a) to (c) occur sequentially. In some embodiments, (c) occurs before (a) and (b). In some embodiments, the single reaction mixture comprises one or more amino acid recognition molecules (e.g., one or more terminal amino acid recognition molecules). In some embodiments, the single reaction mixture comprises a cleavage reagent. In some embodiments, the first amino acid is removed by the cleavage reagent. In some embodiments, the method further comprises identifying the sequence (e.g., a partial or complete sequence) of a single polypeptide molecule by repeating the step of removing and identifying one or more terminal amino acids of the single polypeptide molecule.

[0027] In some embodiments, the present application provides a method for identifying amino acids in a polypeptide, comprising contacting a single polypeptide molecule with one or more amino acid recognition molecules bound to the single polypeptide molecule. In some embodiments, the method further comprises detecting a series of signal pulses that indicate the association of one or more amino acid recognition molecules with the single polypeptide molecule under polypeptide degradation conditions. In some embodiments, the method further comprises identifying a first type of amino acid in the single polypeptide molecule based on a first characteristic pattern of the series of signal pulses. In some embodiments, the signal pulses of the characteristic pulses include an average pulse duration of about 1 millisecond to about 10 seconds.

[0028] In some embodiments, the present application provides a method for identifying terminal amino acids (e.g., N-terminal or C-terminal amino acids) of a polypeptide. In some embodiments, the method involves contacting the polypeptide with one or more labeled recognition molecules that selectively bind to one or more types of terminal amino acids at the terminus of the polypeptide. In some embodiments, the method further includes identifying the terminal amino acids at the terminus of the polypeptide by detecting the interaction between the polypeptide and one or more labeled recognition molecules.

[0029] In yet another embodiment, the present application provides a method for polypeptide sequencing by an Edman-type degradation reaction. In some embodiments, the Edman-type degradation reaction may be carried out by contacting a polypeptide with various reaction mixtures for either detection or cleavage purposes (compared to, for example, dynamic sequencing reactions which may include detection and cleavage using a single reaction mixture).

[0030] Therefore, in some embodiments, the present application provides a method for determining the amino acid sequence of a polypeptide, comprising (i) contacting the polypeptide with one or more labeled recognition molecules that selectively bind to one or more terminal amino acids of the polypeptide's terminal ends. In some embodiments, the method further comprises (ii) identifying the terminal amino acids of the polypeptide (e.g., N-terminal or C-terminal amino acids) by detecting the interaction between the polypeptide and one or more labeled recognition molecules. In some embodiments, the method further comprises (iii) removing the terminal amino acids. In some embodiments, the method further comprises (iv) determining the amino acid sequence of the polypeptide by repeating (i) to (iii) one or more times at the end of the polypeptide.

[0031] In some embodiments, the method includes removing one or more labeled recognition molecules that do not selectively bind to terminal amino acids after (i) and before (ii). In some embodiments, the method includes removing one or more labeled recognition molecules that selectively bind to terminal amino acids after (ii) and before (iii).

[0032] In some embodiments, removing terminal amino acids (e.g., (iii)) includes modifying terminal amino acids by contacting them with an isothiocyanate (e.g., phenylisothiocyanate), and contacting the modified terminal amino acids with a protease that specifically binds to and removes the modified terminal amino acids. In some embodiments, cleaving terminal amino acids (e.g., (iii)) includes modifying terminal amino acids by contacting them with an isothiocyanate, and exposing the modified terminal amino acids to conditions that are sufficiently acidic or basic to remove the modified terminal amino acids.

[0033] In some embodiments, identifying a terminal amino acid includes identifying the terminal amino acid as one of one or more types of terminal amino acids to which one or more labeled recognition molecules bind. In some embodiments, identifying a terminal amino acid includes identifying the terminal amino acid as a type other than one or more types of terminal amino acids to which one or more labeled recognition molecules bind.

[0034] In some embodiments, the present application provides a method for identifying a target protein in a mixed sample. In some embodiments, the method comprises cleaving a mixed protein sample to generate a plurality of polypeptide fragments. In some embodiments, the method further comprises determining the amino acid sequence of at least one of the plurality of polypeptide fragments in the manner of the method of the present application. In some embodiments, the method further comprises identifying a target protein in a mixed sample if the amino acid sequence is uniquely identifiable to the target protein.

[0035] In some embodiments, a method for identifying a target protein in a mixed sample includes cleaving the mixed protein sample to generate multiple polypeptide fragments. In some embodiments, the method further includes labeling one or more types of amino acids in the multiple polypeptide fragments with one or more different luminescent labels. In some embodiments, the method further includes measuring the luminescence over time for at least one of the multiple labeled polypeptides. In some embodiments, the method further includes determining the amino acid sequence of at least one labeled polypeptide based on the detected luminescence. In some embodiments, the method further includes identifying a target protein in a mixed sample if the amino acid sequence is uniquely identifiable to the target protein.

[0036] Therefore, in some embodiments, the polypeptide molecule or protein of interest to be analyzed according to the present invention may be a mixed or purified sample. In some embodiments, the polypeptide molecule or protein of interest is obtained from a biological sample (e.g., blood, tissue, saliva, urine, or other biological source). In some embodiments, the polypeptide molecule or protein of interest is obtained from a patient sample (e.g., a human sample).

[0037] In some embodiments, the present application provides a system comprising at least one hardware processor and at least one non-temporary computer-readable storage medium storing processor-executable instructions that cause the at least one hardware processor to perform the method described in the specification when executed by the at least one hardware processor. In some embodiments, the present application provides at least one non-temporary computer-readable storage medium storing processor-executable instructions that cause the at least one hardware processor to perform the method described in the specification when executed by the at least one hardware processor.

[0038] Details of certain embodiments of the present invention are shown in the detailed description of a particular embodiment below. Other features, purposes, and advantages of the present invention will become apparent from the examples, drawings, and claims. [Brief explanation of the drawing]

[0039] The accompanying drawings, which constitute part of this specification, illustrate several embodiments of the present invention and, together with the description, help to illustrate the principles of the present invention. [Figure 1A] This shows an example of polypeptide sequencing by detecting (Figure 1A) and analyzing (Figure 1B) single-molecule bond interactions. [Figure 1B] This shows an example of polypeptide sequencing by detecting (Figure 1A) and analyzing (Figure 1B) single-molecule bond interactions. [Figure 2]This diagram illustrates examples of labeled recognition molecules, including labeled enzymes and labeled aptamers that selectively bind to one or more terminal amino acids. [Figure 3A] This shows a non-limiting example of an amino acid recognition molecule labeled via a shielding element. Figure 3A illustrates single-molecule peptide sequencing using a recognition molecule labeled via conventional covalent linkage. [Figure 3B] A non-limiting example of an amino acid recognition molecule labeled via a shielding element is shown. Figure 3B illustrates single-molecule peptide sequencing using a recognition molecule containing a shielding element. [Figure 3C] Non-limiting examples of amino acid recognition molecules labeled via a shielding element are shown. Figures 3C-3E illustrate various examples of shielding elements according to the present invention. [Figure 3D] Non-limiting examples of amino acid recognition molecules labeled via a shielding element are shown. Figures 3C-3E illustrate various examples of shielding elements according to the present invention. [Figure 3E] Non-limiting examples of amino acid recognition molecules labeled via a shielding element are shown. Figures 3C-3E illustrate various examples of shielding elements according to the present invention. [Figure 4] This describes the degradation-based process of polypeptide sequencing using labeled recognition molecules. [Figure 5] An example of real-time polypeptide sequencing is shown by evaluating the binding interactions between terminal and / or internal amino acids and labeled recognition molecules and labeled cleavage reagents (e.g., labeled nonspecific exopeptidases). Figure 5 shows an example of real-time sequencing by detecting a series of pulses at the signal output. [Figure 6] An example of real-time polypeptide sequencing is shown by evaluating the binding interactions between terminal and / or internal amino acids and labeled recognition molecules and labeled cleavage reagents (e.g., labeled nonspecific exopeptidases). Figure 6 schematically illustrates the temperature-dependent sequencing process. [Figure 7]An example of real-time polypeptide sequencing is shown by evaluating the binding interactions between terminal and / or internal amino acids and labeled recognition molecules and labeled cleavage reagents (e.g., labeled nonspecific exopeptidases). Figure 7 shows an example of real-time polypeptide sequencing by evaluating the binding interactions between terminal and internal amino acids and labeled recognition molecules and labeled nonspecific exopeptidases. [Figure 8] This paper shows various examples of sample and sample well surface preparations for the analysis of polypeptides and proteins related to this application. Figure 8 generally illustrates an example of the process for preparing terminally modified polypeptides from a protein sample. [Figure 9] This paper shows examples of various preparations of samples and sample well surfaces for the analysis of polypeptides and proteins related to this application. Figure 9 generally illustrates an example of a process for conjugating a polypeptide with a solubilizing linker. [Figure 10] This paper shows various preparation examples of samples and sample well surfaces for the analysis of polypeptides and proteins according to the present invention. Figure 10 shows a schematic example of a sample well having a modified surface that can be used to promote single molecule fixation to the bottom surface. [Figure 11] This is a diagram illustrating an exemplary sequence data processing pipeline for analyzing data obtained during a polypeptide decomposition process according to some embodiments of the technology described herein. [Figure 12] This is a flowchart illustrating an exemplary process for determining the amino acid sequence of a polypeptide molecule according to some embodiments of the technology described herein. [Figure 13] This is a flowchart illustrating an exemplary process for determining the amino acid sequence of a polypeptide corresponding to some embodiment of the technology described herein. [Figure 14] This is a block diagram of an exemplary computer system that may be used to implement some embodiments of the technology described herein. [Figure 15A]This section presents experimental data of selected peptide-linker conjugates prepared and evaluated for solubility enhancement provided by various solubilizing linkers. Figure 15A shows structural examples of synthesized and evaluated peptide-linker conjugates. [Figure 15B] Experimental data of selected peptide-linker conjugates prepared and evaluated for solubility enhancement provided by various solubilizing linkers are shown. Figure 15B shows LCMS results demonstrating N-terminal peptide cleavage. [Figure 15C-1] Experimental data of selected peptide-linker conjugates prepared and evaluated for solubility enhancement provided by various solubilizing linkers are shown. Figure 15C shows the results of loading experiments. [Figure 15C-2] Same as above. [Figure 16] This section outlines the amino acid cleavage activity of exopeptidases selected based on experimental results. [Figure 17A-1] Experimental data from a dye / peptide conjugate assay for the detection and cleavage of terminal amino acids are shown. Figure 17A shows an example scheme and structure used to perform the dye / peptide conjugate assay. [Figure 17A-2] Same as above. [Figure 17B] Experimental data from a dye / peptide conjugate assay for the detection and cleavage of terminal amino acids are shown. Figure 17B shows imaging results of peptide-linker conjugate loading into sample wells in an on-chip assay. [Figure 17C] Experimental data from a dye / peptide conjugate assay for the detection and cleavage of terminal amino acids are shown. Figure 17C shows an example of a signal trace that detected peptide conjugate loading and terminal amino acid cleavage. [Figure 18A-1] Experimental data from a FRET dye / peptide conjugate assay for the detection and cleavage of terminal amino acids are shown. Figure 18A shows an example scheme and structure used to perform the FRET dye / peptide conjugate assay. [Figure 18A-2] Same as above. [Figure 18B] Experimental data from a FRET dye / peptide conjugate assay for the detection and cleavage of terminal amino acids are shown. Figure 18B shows FRET imaging results at various time points. [Figure 18C] Experimental data from a FRET dye / peptide conjugate assay for the detection and cleavage of terminal amino acids are shown. Figure 18C shows the cutting efficiency at various time points. [Figure 18D] Experimental data from a FRET dye / peptide conjugate assay for the detection and cleavage of terminal amino acids are shown. Figure 18D shows the cleavage observed at various time points. [Figure 18E] Experimental data from a FRET dye / peptide conjugate assay for the detection and cleavage of terminal amino acids are shown. Figure 18E shows additional FRET imaging results at various time points using proline iminopeptidase (yPIP) from Yersinia pestis. [Figure 18F] Experimental data from a FRET dye / peptide conjugate assay for the detection and cleavage of terminal amino acids are shown. Figure 18F shows FRET imaging results at various time points using aminopeptidase (VPr) derived from Vibrio proteolyticus. [Figure 19A] Experimental data on terminal amino acid recognition using labeled recognition molecules are shown. Figure 19A shows the crystal structure of the ClpS2 protein labeled for these experiments. [Figure 19B-1] Experimental data on terminal amino acid recognition by labeled recognition molecules are shown. Figure 19B shows a single-molecule intensity trace illustrating N-terminal amino acid recognition by labeled ClpS2 protein. [Figure 19B-2] Same as above. [Figure 19C] Experimental data on terminal amino acid recognition using labeled recognition molecules are shown. Figure 19C is a plot showing the average pulse duration of various terminal amino acids. [Figure 19D]Experimental data on terminal amino acid recognition using labeled recognition molecules are shown. Figure 19D is a plot showing the average inter-pulse duration of various terminal amino acids. [Figure 19E] Experimental data on terminal amino acid recognition using labeled recognition molecules are shown. Figure 19E shows plots further illustrating discriminant pulse durations between various terminal amino acids. [Figure 19F] Experimental data on terminal amino acid recognition by labeled recognition molecules are shown. Figures 19F, 19G, and 19H show examples of residence time analysis results demonstrating leucine recognition by the ClpS protein (teClpS) derived from Thermosynochoccus elongatus. [Figure 19G] Experimental data on terminal amino acid recognition by labeled recognition molecules are shown. Figures 19F, 19G, and 19H show examples of residence time analysis results demonstrating leucine recognition by the ClpS protein (teClpS) derived from Thermosynochoccus elongatus. [Figure 19H] Experimental data on terminal amino acid recognition by labeled recognition molecules are shown. Figures 19F, 19G, and 19H show examples of residence time analysis results demonstrating leucine recognition by the ClpS protein (teClpS) derived from Thermosynochoccus elongatus. [Figure 19I] This shows experimental data on terminal amino acid recognition by labeled recognition molecules. Figure 19I shows an example of residence time analysis results demonstrating the distinguishable recognition of phenylalanine, leucine, tryptophan, and tyrosine by A. tumefaciens ClpS1. [Figure 19J] This shows experimental data on terminal amino acid recognition by labeled recognition molecules. Figure 19J shows an example of residence time analysis results demonstrating leucine recognition by S. elongatus ClpS2. [Figure 19K]Experimental data on terminal amino acid recognition by labeled recognition molecules are shown. Figures 19K-19L show an example of residence time analysis results demonstrating proline recognition by GID4. [Figure 19L] Experimental data on terminal amino acid recognition by labeled recognition molecules are shown. Figures 19K-19L show an example of residence time analysis results demonstrating proline recognition by GID4. [Figure 19M] Experimental data on terminal amino acid recognition by labeled recognition molecules are shown. Figure 19M shows exemplary binding curves for atClpS2-V1 with peptides having various N-terminal amino acids. [Figure 20A] An example of the results of a polypeptide sequencing reaction performed in real time using a labeled ClpS2-recognizing protein and an aminopeptidase cleavage reagent in the same reaction mixture is shown. Figure 20A shows the signal trace data of the first sequencing reaction. [Figure 20B] An example of the results of a polypeptide sequencing reaction performed in real time using a labeled ClpS2-recognizing protein and an aminopeptidase cleavage reagent in the same reaction mixture is shown. Figure 20B shows the pulse duration statistics of the signal trace data shown in Figure 20A. [Figure 20C] An example of the results of a polypeptide sequencing reaction performed in real time using a labeled ClpS2-recognizing protein and an aminopeptidase cleavage reagent in the same reaction mixture is shown. Figure 20C shows the signal trace data of the second sequencing reaction. [Figure 20D] An example of the results of a polypeptide sequencing reaction performed in real time using a labeled ClpS2-recognizing protein and an aminopeptidase cleavage reagent in the same reaction mixture is shown. Figure 20D shows the pulse duration statistics of the signal trace data shown in Figure 20C. [Figure 21A] Experimental data on terminal amino acid identification and cleavage using labeled exopeptidases are shown. Figure 21A shows the crystal structure of site-specifically labeled proline iminopeptidase (yPIP) used for these experiments. [Figure 21B]Experimental data on terminal amino acid identification and cleavage using labeled exopeptidase are shown. Figure 21B shows the degree of labeling of the purified protein product. [Figure 21C] This shows experimental data on terminal amino acid identification and cleavage using labeled exopeptidase. Figure 21C is an image of the SDS page confirming site-specific labeling of yPIP. [Figure 21D] This shows experimental data on terminal amino acid identification and cleavage using labeled exopeptidase. Figure 21D is an overexposed image of the SDS page gel confirming site-specific labeling. [Figure 21E] This shows experimental data on terminal amino acid identification and cleavage using labeled exopeptidase. Figure 21E is an image of a Coomassie-stained gel confirming the purity of the labeled protein product. [Figure 21F] This shows experimental data on terminal amino acid identification and cleavage using labeled exopeptidase. Figure 21F is an HPLC trace demonstrating the cleavage activity of labeled exopeptidase. [Figure 22A] This section presents experimental data evaluating the recognition of amino acids containing specific post-translational modifications. Figure 22A shows a representative trace demonstrating phosphotyrosine recognition by an SH2 domain-containing protein. [Figure 22B] This figure shows experimental data evaluating the recognition of amino acids containing specific post-translational modifications. Figure 22B shows pulse duration data corresponding to the trace in Figure 22A, and Figure 22C shows the statistics determined for the trace. [Figure 22C] This figure shows experimental data evaluating the recognition of amino acids containing specific post-translational modifications. Figure 22B shows pulse duration data corresponding to the trace in Figure 22A, and Figure 22C shows the statistics determined for the trace. [Figure 22D] This section presents experimental data evaluating the recognition of amino acids containing specific post-translational modifications. Figures 22D-22F show representative traces from negative control experiments. [Figure 22E]This section presents experimental data evaluating the recognition of amino acids containing specific post-translational modifications. Figures 22D-22F show representative traces from negative control experiments. [Figure 22F] This section presents experimental data evaluating the recognition of amino acids containing specific post-translational modifications. Figures 22D-22F show representative traces from negative control experiments. [Figure 23] This plot shows the median pulse duration in experiments evaluating the effect of penaltymate amino acids on pulse duration. [Figure 24A] This shows experimental data evaluating simultaneous amino acid recognition by differentially labeled recognition molecules. Figure 24A shows a representative trace. [Figure 24B] This shows experimental data evaluating simultaneous amino acid recognition by differentially labeled recognition molecules. Figure 24B is a plot comparing the pulse duration data obtained during these experiments for each recognition molecule. [Figure 24C] This shows experimental data evaluating simultaneous amino acid recognition by differentially labeled recognition molecules. Figure 24C shows pulse duration statistics for these experiments. [Figure 25A] This section presents experimental data evaluating the photostability of peptides during single-molecule recognition. Figure 25A shows a representative trace of recognition using atClpS2-V1, which is labeled with a dye approximately 2 nm from the amino acid binding site. [Figure 25B] This shows experimental data evaluating the photostability of peptides during single-molecule recognition. Figure 25B shows a visualization of the structure of the ClpS2 protein used in these experiments. [Figure 25C] This shows experimental data evaluating the photostability of peptides during single-molecule recognition. Figure 25C shows a representative trace of recognition using ClpS2 labeled with a dye >10 nm from the amino acid binding site via a DNA / protein linker. [Figure 26A-1] This shows a representative trace of a polypeptide sequencing reaction performed in real time on a complementary metal-oxide-semiconductor (CMOS) chip using a ClpS2-recognizing protein labeled via a DNA / streptavidin linker in the presence of an aminopeptidase cleavage reagent. [Figure 26A-2] Same as above. [Figure 26B-1] This shows a representative trace of a polypeptide sequencing reaction performed in real time on a complementary metal-oxide-semiconductor (CMOS) chip using a ClpS2-recognizing protein labeled via a DNA / streptavidin linker in the presence of an aminopeptidase cleavage reagent. [Figure 26B-2] Same as above. [Figure 26C-1] This shows a representative trace of a polypeptide sequencing reaction performed in real time on a complementary metal-oxide-semiconductor (CMOS) chip using a ClpS2-recognizing protein labeled via a DNA / streptavidin linker in the presence of an aminopeptidase cleavage reagent. [Figure 26C-2] Same as above. [Figure 26D-1] This shows a representative trace of a polypeptide sequencing reaction performed in real time on a complementary metal-oxide-semiconductor (CMOS) chip using a ClpS2-recognizing protein labeled via a DNA / streptavidin linker in the presence of an aminopeptidase cleavage reagent. [Figure 26D-2] Same as above. [Figure 27] This shows a representative trace of a polypeptide sequencing reaction performed in real time using atClpS2-V1 recognition protein labeled via a DNA / streptavidin linker in the presence of Pyrococcus horikoshii TET aminopeptidase cleavage reagent. [Figure 28A-1] Representative trace data of polypeptide sequencing reactions performed in real time using multiple types of exopeptidases with differential cleavage specificity are shown. Figure 28A shows a representative trace of a reaction performed with hTET exopeptidase, along with the expanded pulse pattern region shown in Figure 28B. [Figure 28A-2] Same as above. [Figure 28B]Representative trace data of polypeptide sequencing reactions performed in real time using multiple types of exopeptidases with differential cleavage specificity are shown. Figure 28A shows a representative trace of a reaction performed with hTET exopeptidase, along with the expanded pulse pattern region shown in Figure 28B. [Figure 28C-1] Representative trace data of polypeptide sequencing reactions performed in real time using multiple types of exopeptidases with differential cleavage specificity are shown. Figure 28C shows representative traces of reactions performed with both hTET and yPIP exopeptidases, along with the expanded pulse pattern region shown in Figure 28D and the additional representative traces shown in Figure 28E. [Figure 28C-2] Same as above. [Figure 28D] Representative trace data of polypeptide sequencing reactions performed in real time using multiple types of exopeptidases with differential cleavage specificity are shown. Figure 28C shows representative traces of reactions performed with both hTET and yPIP exopeptidases, along with the expanded pulse pattern region shown in Figure 28D and the additional representative traces shown in Figure 28E. [Figure 28E] Representative trace data of polypeptide sequencing reactions performed in real time using multiple types of exopeptidases with differential cleavage specificity are shown. Figure 28C shows representative traces of reactions performed with both hTET and yPIP exopeptidases, along with the expanded pulse pattern region shown in Figure 28D and the additional representative traces shown in Figure 28E. [Figure 28F] Representative trace data of polypeptide sequencing reactions performed in real time using multiple types of exopeptidases with differential cleavage specificity are shown. Figure 28F shows representative traces of further reactions performed with both hTET and yPIP exopeptidases, along with the expanded pulse pattern region shown in Figure 28G and additional representative traces shown in Figure 28H. [Figure 28G]Representative trace data of polypeptide sequencing reactions performed in real time using multiple types of exopeptidases with differential cleavage specificity are shown. Figure 28F shows representative traces of further reactions performed with both hTET and yPIP exopeptidases, along with the expanded pulse pattern region shown in Figure 28G and additional representative traces shown in Figure 28H. [Figure 28H] Representative trace data of polypeptide sequencing reactions performed in real time using multiple types of exopeptidases with differential cleavage specificity are shown. Figure 28F shows representative traces of further reactions performed with both hTET and yPIP exopeptidases, along with the expanded pulse pattern region shown in Figure 28G and additional representative traces shown in Figure 28H. [Figure 28I] Representative trace data of polypeptide sequencing reactions performed in real time using multiple types of exopeptidases with differential cleavage specificity are shown. Figure 28I shows representative traces of reactions performed with both PfuTET and yPIP exopeptidases, along with the expanded pulse pattern region shown in Figure 28J. [Figure 28J] Representative trace data of polypeptide sequencing reactions performed in real time using multiple types of exopeptidases with differential cleavage specificity are shown. Figure 28I shows representative traces of reactions performed with both PfuTET and yPIP exopeptidases, along with the expanded pulse pattern region shown in Figure 28J. [Figure 29A] Data from experiments evaluating newly identified ClpS homologs are shown. Figure 29A shows SDS-PAGE gel imaging of purified ClpS protein. [Figure 29B-1] Data from experiments evaluating newly identified ClpS homologs are shown. Figure 29B shows the results of biolayer interference screening of ClpS homologs. [Figure 29B-2] Same as above. [Figure 29C]Data from experiments evaluating newly identified ClpS homologs are shown. Figure 29C shows the screening selection results. [Figure 29D] This shows experimental data evaluating newly identified ClpS homologs. Figure 29D shows the reaction curves between the ClpS protein (PS372) and the LA, IA, and VA peptides. [Figure 29E] Data from experiments evaluating newly identified ClpS homologs are shown. Figure 29E shows the polarization response of PS372 and four other homologs, as well as a protein-free control. [Figure 29F] Data from experiments evaluating newly identified ClpS homologs are shown. Figure 29F shows the biolayer interference reaction curves of IA, IR, IQ, VA, and VR peptides. [Figure 29G] This shows experimental data evaluating the newly identified ClpS homolog. Figure 29G shows the pulse width histogram and representative trace of PS372 with IR peptide (top panel) and LF peptide (bottom panel). [Figure 30A] This section presents experimental data evaluating terminal amino acid recognition by newly identified ClpS homologs. Figure 30A shows in vivo biotinylation of PS372 by SDS-PAGE. [Figure 30B] This shows experimental data evaluating terminal amino acid recognition by newly identified ClpS homologs. Figure 30B shows the purification profile of PS372 conjugated with SV- dye. [Figure 30C] This shows experimental data evaluating terminal amino acid recognition by newly identified ClpS homologs. Figure 30C shows the SDS-PAGE of purified PS372 conjugated with SV dye. [Figure 30D] This shows experimental data evaluating terminal amino acid recognition by newly identified ClpS homologs. Figure 30D shows representative traces illustrating transitions from I to L(a) and from L to I(b). [Figure 30E]This section presents experimental data evaluating terminal amino acid recognition by newly identified ClpS homologs. Figure 30E shows illustrative data from a real-time dynamic peptide sequencing assay using dye-labeled PS327. [Figure 31A] Engineering data for the methionine-binding ClpS protein are shown. Figure 31A shows the results of selection performed via fluorescence-activated cell sorting (FACS). [Figure 31B] Engineering data for the methionine-binding ClpS protein are shown. Figure 31B shows an exemplary reaction curve for the methionine-binding ClpS protein having a peptide with an N-terminal LA. [Figure 31C] Engineering data for the methionine-binding ClpS protein are shown. Figure 31C shows an exemplary reaction curve for the methionine-binding ClpS protein having a peptide with an N-terminal MA. [Figure 31D] Engineering data for the methionine-binding ClpS protein are shown. Figure 31D shows an exemplary reaction curve for the methionine-binding ClpS protein having a peptide with an N-terminal MR. [Figure 31E] Engineering data for the methionine-binding ClpS protein are shown. Figure 31E shows an exemplary reaction curve for the methionine-binding ClpS protein having a peptide with an N-terminal FA. [Figure 31F] Engineering data for the methionine-binding ClpS protein are shown. Figure 31F shows an exemplary reaction curve for the methionine-binding ClpS protein with a peptide having an N-terminal MQ. [Figure 32A] The data from experiments evaluating UBR box domain homologs are shown. Figures 32A-32B show exemplary binding curves for UBR box homologs PS535 (Figure 32A) and PS522 (Figure 32B) that bind to 14 peptidases containing an N-terminal R before the penaltymate amino acid. [Figure 32B]The data from experiments evaluating UBR box domain homologs are shown. Figures 32A-32B show exemplary binding curves for UBR box homologs PS535 (Figure 32A) and PS522 (Figure 32B) that bind to 14 peptidases containing an N-terminal R before the penaltymate amino acid. [Figure 32C] This shows experimental data for evaluating UBR box domain homologs. Figure 32C is a heatmap showing the results of measuring 24 UBR box homologs that bind to the N-terminal R peptide. [Figure 32D-1] The data from experiments evaluating UBR box domain homologs are shown. Figure 32D is a heatmap showing the results of measuring an expanded set of UBR box homologs that bind to polypeptides containing R, K, or H at the N-terminus. [Figure 32D-2] Same as above. [Figure 32E-1] The data from experiments evaluating UBR box domain homologs are shown. Figure 32E is a heatmap showing the results of measuring an expanded set of UBR box homologs that bind to 14 polypeptides containing the N-terminal R before various amino acids at penalty mate positions. [Figure 32E-2] Same as above. [Figure 32F] The data from experiments evaluating UBR box domain homologs are shown. Figure 32F shows the results of a single-point fluorescence polarization assay. [Figure 32G] This shows experimental data for evaluating UBR box domain homologs. Figure 32G shows the analysis of polarization results for binding affinity determination. [Figure 32H] The data from experiments evaluating UBR box domain homologs are shown. Figure 32H shows a representative trace of PS621 in the recognition assay. [Figure 32I] This shows experimental data evaluating UBR box domain homologs. Figure 32I shows an exemplary sequencing trace of a three-binder dynamic sequencing reaction. [Figure 33A]The data from experiments evaluating P372 homolog proteins are shown. Figures 33A–33E show exemplary binding curves of PS372 (Figure 33A) and its homologs PS545 (Figure 33B), PS551 (Figure 33C), PS557 (Figure 33D), and PS558 (Figure 33E) to four polypeptides containing various N-terminal amino acids (I, V, L, F). [Figure 33B] The data from experiments evaluating P372 homolog proteins are shown. Figures 33A–33E show exemplary binding curves of PS372 (Figure 33A) and its homologs PS545 (Figure 33B), PS551 (Figure 33C), PS557 (Figure 33D), and PS558 (Figure 33E) to four polypeptides containing various N-terminal amino acids (I, V, L, F). [Figure 33C] The data from experiments evaluating P372 homolog proteins are shown. Figures 33A–33E show exemplary binding curves of PS372 (Figure 33A) and its homologs PS545 (Figure 33B), PS551 (Figure 33C), PS557 (Figure 33D), and PS558 (Figure 33E) to four polypeptides containing various N-terminal amino acids (I, V, L, F). [Figure 33D] The data from experiments evaluating P372 homolog proteins are shown. Figures 33A–33E show exemplary binding curves of PS372 (Figure 33A) and its homologs PS545 (Figure 33B), PS551 (Figure 33C), PS557 (Figure 33D), and PS558 (Figure 33E) to four polypeptides containing various N-terminal amino acids (I, V, L, F). [Figure 33E] The data from experiments evaluating P372 homolog proteins are shown. Figures 33A–33E show exemplary binding curves of PS372 (Figure 33A) and its homologs PS545 (Figure 33B), PS551 (Figure 33C), PS557 (Figure 33D), and PS558 (Figure 33E) to four polypeptides containing various N-terminal amino acids (I, V, L, F). [Figure 33F] The data from experiments evaluating P372 homolog proteins are shown. Figure 33F is a heatmap showing the results of measuring 34 PS372 homologs. [Figure 34A] This shows experimental data evaluating an engineered polyvalent amino acid binder (PS610) generated as a single polypeptide containing a serial copy of atClpS2-V1. Figure 34A shows representative trace data from a peptide-on-chip recognition assay. [Figure 34B] This shows experimental data evaluating an engineered polyvalent amino acid binder (PS610) generated as a single polypeptide containing a serial copy of atClpS2-V1. Figure 34B is a plot showing the average pulse velocity as a function of binder concentration. [Figure 34C] This shows experimental data evaluating an engineered polyvalent amino acid binder (PS610) generated as a single polypeptide containing a serial copy of atClpS2-V1. Figure 34C shows illustrative data from real-time dynamic peptide sequencing assays using dye-labeled P610 and PS327. [Figure 34D] This shows experimental data evaluating an engineered polyvalent amino acid binder (PS610) generated as a single polypeptide containing a serial copy of atClpS2-V1. Figure 34D shows representative trace data from a binder-on-chip assay. [Figure 35A] This shows experimental data evaluating series ClpS2 constructs containing various linkers. Figure 35A shows an exemplary bond curve for the monovalent binder atClpS2-V1. [Figure 35B] This shows experimental data evaluating series ClpS2 constructs containing various linkers. Figures 35B–35D show exemplary coupling curves for series constructs having two copies of atClpS2-V1 separated by linker 1 (Figure 35B), linker 2 (Figure 35C), or linker 3 (Figure 35D). [Figure 35C] This shows experimental data evaluating series ClpS2 constructs containing various linkers. Figures 35B–35D show exemplary coupling curves for series constructs having two copies of atClpS2-V1 separated by linker 1 (Figure 35B), linker 2 (Figure 35C), or linker 3 (Figure 35D). [Figure 35D] This shows experimental data evaluating series ClpS2 constructs containing various linkers. Figures 35B–35D show exemplary coupling curves for series constructs having two copies of atClpS2-V1 separated by linker 1 (Figure 35B), linker 2 (Figure 35C), or linker 3 (Figure 35D). [Figure 36A] This shows experimental data evaluating engineered polyvalent amino acid binders, generated as single polypeptides with serial copies of the same or different ClpS proteins. Figure 36A shows exemplary binding curves for the monovalent binder atClpS2-V1 (left plot) and the monovalent binder PS372 (right plot). [Figure 36B] This shows experimental data evaluating engineered polyvalent amino acid binders generated as single polypeptides containing serial copies of the same or different ClpS proteins. Figure 36B shows exemplary binding curves for polyvalent polypeptides containing serial copies of atClpS2-V1 and PS372. [Figure 36C] This shows experimental data evaluating engineered polyvalent amino acid binders, generated as single polypeptides with serial copies of the same or different ClpS proteins. Figure 36C shows an exemplary binding curve for the monovalent binder PS372. [Figure 36D] This shows experimental data evaluating engineered polyvalent amino acid binders, generated as single polypeptides with serial copies of the same or different ClpS proteins. Figure 36D shows an exemplary binding curve for a polyvalent polypeptide with two serial copies of S372. [Figure 36E] This shows experimental data evaluating engineered polyvalent amino acid binders, generated as single polypeptides with serial copies of the same or different ClpS proteins. Figure 36E shows an exemplary binding curve for the monovalent binder PS557. [Figure 36F]The following are experimental data evaluating engineered polyvalent amino acid binders, generated as single polypeptides with serial copies of the same or different ClpS proteins. Figures 36F–36H show exemplary binding curves for serial constructs with two copies of PS557 separated by linker 1 (Figure 36F), linker 2 (Figure 36G), or linker 3 (Figure 36H). [Figure 36G] The following are experimental data evaluating engineered polyvalent amino acid binders, generated as single polypeptides with serial copies of the same or different ClpS proteins. Figures 36F–36H show exemplary binding curves for serial constructs with two copies of PS557 separated by linker 1 (Figure 36F), linker 2 (Figure 36G), or linker 3 (Figure 36H). [Figure 36H] The following are experimental data evaluating engineered polyvalent amino acid binders, generated as single polypeptides with serial copies of the same or different ClpS proteins. Figures 36F–36H show exemplary binding curves for serial constructs with two copies of PS557 separated by linker 1 (Figure 36F), linker 2 (Figure 36G), or linker 3 (Figure 36H). [Figure 37A] This figure shows data from stopped-flow rapid kinetic analysis for determining the kon rate constant and koff rate of binders and fusion proteins driven by C-terminal addition of a protein shield. Figure 37A shows a schematic diagram illustrating the assay design (top panel) and plots showing experimental results and analyses for determining the association rate constant (kon) (middle and bottom panels). [Figure 37B] This figure shows data from stopped-flow rapid kinetic analysis to determine the kon rate constant and koff rate of binders and fusion proteins driven by C-terminal addition of a protein shield. Figure 37B shows a schematic diagram illustrating the assay design (top panel) and a plot showing experimental results and analysis for measuring the dissociation rate (koff) (bottom panel). [Modes for carrying out the invention]

[0040] Aspects of the present application relate to a method for protein sequencing and identification, a method for polypeptide sequencing and identification, a method for amino acid identification, and compositions for carrying out such methods.

[0041] In some embodiments, the present application relates to the discovery of polypeptide sequencing techniques that can be implemented using existing analytical instruments with little or no device modification. For example, previous polypeptide sequencing strategies involved the repeated cycling of various reagent mixtures via a reaction vessel containing the polypeptide to be analyzed. Such strategies may require modifications to existing analytical instruments, such as nucleic acid sequencing instruments, which may not have flow cells or similar devices capable of reagent cycling. The inventors recognized and appreciated that certain polypeptide sequencing techniques of the present application do not require repeated reagent cycling, thus enabling the use of existing instruments without significant modifications that may increase instrument size. Therefore, in some embodiments, the present application provides polypeptide sequencing methods that enable the use of smaller sequencing instruments. In some embodiments, the present application relates to the discovery of polypeptide sequencing techniques that enable both genomic and proteomic analyses performed using the same sequencing instrument.

[0042] The inventors have further recognized and appreciated that differential binding interactions can provide an additional or alternative approach to conventional labeling strategies for polypeptide sequencing. Conventional polypeptide sequencing may involve labeling each type of amino acid with a uniquely identifiable label. Given the numerous post-translational mutations and at least 20 different types of naturally occurring amino acids, such processes can be laborious and error-prone. In some embodiments, the present application relates to the discovery of techniques involving the use of amino acid recognition molecules that differentially associate with different types of amino acids to generate detectable characteristic signatures that serve as indicators of the polypeptide amino acid sequence. Thus, embodiments of the present application provide techniques that do not require polypeptide labeling and / or strong chemical reagents used in certain conventional polypeptide sequencing, thereby increasing the throughput and / or accuracy of sequence information obtained from samples.

[0043] In some embodiments, the present invention relates to the discovery that polypeptide sequencing reactions can be monitored in real time using only a single reaction mixture (for example, without requiring repeated reagent cycling via a reaction vessel). As detailed above, conventional polypeptide sequencing reactions may involve exposing polypeptides to various reagent mixtures in a cycle between an amino acid detection step and an amino acid cleavage step. Therefore, in some embodiments, the present invention relates to advances in next-generation sequencing that enable polypeptide analysis by real-time amino acid detection throughout the entire decomposition reaction in progress. Such approaches to polypeptide analysis by dynamic sequencing are described below.

[0044] As described herein, in some embodiments, the present application provides a polypeptide sequencing method comprising obtaining data during a polypeptide degradation process and analyzing the data to determine portions of the data corresponding to amino acids sequentially exposed at the polypeptide termini during the degradation process. In some embodiments, the portions of the data include a series of signal pulses that indicate the association of one or more amino acid recognition molecules with sequential amino acids exposed at the polypeptide termini (e.g., during degradation). In some embodiments, the series of signal pulses correspond to a series of reversible single-molecule bonding interactions at the polypeptide termini during the degradation process.

[0045] A non-limiting example of polypeptide sequencing by detecting single-molecule binding interactions during polypeptide degradation processes is schematically illustrated in Figure 1A. An example of signal tracing (I) is shown in a series of panels (II) depicting various association events in response to changes in the signal. As shown, association events between amino acid recognition molecules (indicated as speckles) and terminal amino acids of polypeptides (indicated as beads-on-a-strings) generate changes in the magnitude of the signal that persist over a certain duration.

[0046] Panels (A) and (B) depict various association events between an amino acid recognition molecule and a first amino acid (e.g., the first terminal amino acid) exposed at the terminus of a polypeptide. Each association event generates a change in the signal trace (I), characterized by a change in the magnitude of the signal that persists over the duration of the association event. Therefore, the durations between association events in panels (A) and (B) correspond to the duration during which the polypeptide is not detectably associated with the amino acid recognition molecule.

[0047] Panels (C) and (D) depict various association events between an amino acid recognition molecule and a second amino acid (e.g., a second terminal amino acid) exposed at the end of a polypeptide. As described herein, an amino acid “exposed” at the end of a polypeptide is an amino acid that remains attached to the polypeptide and becomes a terminal amino acid upon degradation by the removal of a preceding terminal amino acid (e.g., either alone or with one or more additional amino acids). Thus, the first and second amino acids in a series of panels (II) provide exemplary examples of sequential amino acids exposed at the end of a polypeptide, in which case the second amino acid is the terminal amino acid that has become the terminal amino acid by the removal of the first amino acid.

[0048] As is generally depicted, the association events of panels (C) and (D) produce changes in signal trace (I) characterized by changes in magnitude that persist for a relatively shorter duration than those of panels (A) and (B), and the duration between the association events of panels (C) and (D) is relatively shorter than that between panels (A) and (B). As described herein, in some embodiments, one or both of these identifiable changes in the signal may be used to determine a characteristic pattern of the signal trace (I) that is identifiable between different types of amino acids. In some embodiments, the transition from one characteristic pattern to the other is an indicator of amino acid cleavage. As used herein, in some embodiments, amino acid cleavage means the removal of at least one amino acid from the terminal of a polypeptide (e.g., the removal of at least one terminal amino acid from a polypeptide). In some embodiments, amino acid cleavage is determined by inference based on the duration between characteristic patterns. In some embodiments, amino acid cleavage is determined by detecting a change in the signal produced by the association of a labeled cleavage reagent with terminal amino acids of a polypeptide. Since amino acids are sequentially cleaved from the end of the polypeptide during degradation, a series of size changes or a series of signal pulses can be detected. In some embodiments, the signal pulse data can be analyzed as illustrated in Figure 1B.

[0049] In some embodiments, the signal data can be analyzed to extract signal pulse information by applying threshold levels to one or more parameters of the signal data. For example, panel (III) applies a threshold magnitude level ("M") to the signal data of signal trace example (I). L In some embodiments, M L This is the minimum difference between the signal detected at a given point in time and the baseline determined for a given dataset. In some embodiments, the signal pulse ("sp") is M L It serves as an indicator of a change exceeding a certain magnitude and is assigned to each portion of the data that persists for a certain duration. In some embodiments, the threshold duration is M L The data can be applied to the portion that satisfies the condition and determine whether a signal pulse is assigned to that portion. For example, experimental artifacts may not persist for a sufficient duration to assign a signal pulse with the desired reliability. L This can cause changes in magnitude exceeding a certain threshold (for example, transient association events that may not allow identification of amino acid types, or nonspecific detection events such as diffusion into or fixation of reagents within the observation area). Therefore, in some embodiments, signal pulses are extracted from signal data based on threshold magnitude levels and threshold durations.

[0050] The extracted signal pulse information is shown in panel (III) along with an example of a superimposed signal trace (I) for illustrative purposes. In some embodiments, the peak of the signal pulse magnitude is M L The magnitude is determined by averaging the detected magnitudes over a duration in which the threshold persists. It should be recognized that in some embodiments, the term “signal pulse” as used herein may mean a change in signal data that persists above a baseline over a certain duration (e.g., raw signal data as illustrated by signal trace example (I)) or signal pulse information extracted therefrom (e.g., processed signal data as illustrated in panel (IV)).

[0051] Panel (IV) shows signal pulse information extracted from signal trace example (I). In some embodiments, the signal pulse information can be analyzed to identify different types of amino acids based on various characteristic patterns in a series of signal pulses. For example, as shown in Panel (IV), the signal pulse information can be an indicator of a first type of amino acid based on a first characteristic pattern ("CP1") and a second type of amino acid based on a second characteristic pattern ("CP2"). For example, two signal pulses detected at an earlier time point provide information that indicates the first terminal amino acid of the polypeptide based on CP1, and two signal pulses detected at a later time point provide information that indicates the second terminal amino acid of the polypeptide based on CP2.

[0052] Furthermore, as shown in panel (IV), each signal pulse includes a pulse duration ("pd") corresponding to the association event between the amino acid recognition molecule and the amino acids of the characteristic pattern. In some embodiments, the pulse duration is specific to the rate of binding dissociation. Also, as shown, each signal pulse of the characteristic pattern is separated from other signal pulses of the characteristic pattern by an interpulse duration ("ipd"). In some embodiments, the interpulse duration is specific to the rate of binding association. In some embodiments, the magnitude change ("ΔM") can be determined for the signal pulse based on the difference between the baseline and the peak of the signal pulse. In some embodiments, the characteristic pattern is determined based on the pulse duration. In some embodiments, the characteristic pattern is determined based on the pulse duration and the interpulse duration. In some embodiments, the characteristic pattern is determined based on one or more of the pulse duration, interpulse duration, and magnitude change.

[0053] Therefore, as illustrated by Figures 1A-1B, in some embodiments, polypeptide sequencing is performed by detecting a series of signal pulses that indicate the association of one or more amino acid recognition molecules with sequential amino acids exposed at the termini of the polypeptide during the ongoing degradation reaction. The series of signal pulses can be analyzed to determine a characteristic pattern within the series of signal pulses, and the change in the characteristic pattern over time can be used to determine the amino acid sequence of the polypeptide.

[0054] In some embodiments, a series of signal pulses includes a series of changes in the magnitude of an optical signal over time. In some embodiments, a series of changes in the optical signal includes a series of changes in luminescence generated during an association event. In some embodiments, the luminescence is generated by a detectable label associated with one or more reagents in the sequencing reaction. For example, in some embodiments, each of one or more amino acid recognition molecules includes a luminescent label. In some embodiments, a cleavage reagent includes a luminescent label. Examples of luminescent labels and their uses according to this application are provided elsewhere in this specification.

[0055] In some embodiments, a series of signal pulses includes a series of changes in the magnitude of an electrical signal over time. In some embodiments, a series of changes in the electrical signal includes a series of changes in conductance generated at an association event. In some embodiments, conductivity is generated by a detectable label associated with one or more reagents in a sequencing reaction. For example, in some embodiments, each of one or more amino acid recognition molecules includes a conductivity label. Conductivity labels relating to this application and examples of their use are provided elsewhere in this specification. Methods for identifying single molecules using conductivity labels are described (see, for example, U.S. Patent Application Publication No. 2017 / 0037462).

[0056] In some embodiments, a series of conductance changes includes a series of conductance changes via a nanopore. For example, a method for evaluating receptor-ligand interactions using a nanopore has been described (see, for example, Thakur, AK and Movileanu, L., Nature Biotechnology, Vol. 37, No. 1, 2019). The inventors recognize and appreciate that such nanopores can be used to monitor polypeptide sequencing reactions according to the present invention. Therefore, in some embodiments, the present invention provides a polypeptide sequencing method comprising contacting a single polypeptide molecule with one or more amino acid recognition molecules, in which case the single polypeptide molecule is immobilized on a nanopore. In some embodiments, the method further comprises sequencing a single polypeptide molecule by detecting a series of conductance changes via a nanopore that indicate the association of one or more terminal amino acid recognition molecules with sequential amino acids exposed at the terminals of the single polypeptide while the single polypeptide is being degraded.

[0057] In some embodiments, the present application provides a method for sequencing and / or identifying individual proteins in a complex mixture of proteins by identifying one or more amino acids of a polypeptide from the mixture. In some embodiments, one or more amino acids of a polypeptide (e.g., terminal amino acids and / or internal amino acids) are labeled (e.g., directly or indirectly, using a binder such as an amino acid recognition molecule), and the relative position of the labeled amino acid in the polypeptide is determined. In some embodiments, the relative position of the amino acid in the polypeptide is determined by a series of amino acid labeling and cleavage steps. However, in some embodiments, the relative position of the labeled amino acid in the polypeptide can be determined by translocating the labeled polypeptide through a pore (e.g., a protein channel) without removing the amino acid from the polypeptide, and by detecting a signal (e.g., a FRET signal) from the labeled amino acid during translocation through the pore to determine the relative position of the labeled amino acid in the polypeptide molecule.

[0058] In some embodiments, the identity of a terminal amino acid (e.g., an N-terminal or C-terminal amino acid) is evaluated, then the terminal amino acid is removed, and the identity of the next amino acid at the terminal is evaluated, and this process is repeated until multiple sequential amino acids in the polypeptide are evaluated. In some embodiments, evaluating the identity of an amino acid includes determining the type of amino acid present. In some embodiments, determining the type of amino acid includes determining the identity of the actual amino acid by determining, for example, which of the 20 naturally occurring amino acids it is (e.g., using a binder specific to the individual terminal amino acid). In some embodiments, the type of amino acid is selected from alanine, arginine, asparagine, aspartic acid, cysteine, glutamine, glutamic acid, glycine, histidine, isoleucine, leucine, lysine, methionine, phenylalanine, proline, selenocysteine, threonine, tryptophan, tyrosine, and valine.

[0059] However, in some embodiments, assessing the identity of a terminal amino acid type may involve determining a subset of potential amino acids that can be present at the end of a polypeptide. In some embodiments, this can be achieved by determining that an amino acid is not one or more specific amino acids (and therefore could be any of the other amino acids). In some embodiments, this can be achieved by determining which of a particular subset of amino acids can be at the end of a polypeptide (for example, based on size, charge, hydrophobicity, post-translational modification, and binding affinity) (for example, using a binder that binds to a particular subset of two or more terminal amino acids).

[0060] In some embodiments, assessing the identity of a terminal amino acid type involves determining whether an amino acid has post-translational modifications. Non-limiting examples of post-translational modifications include acetylation, ADP-ribosylation, caspase cleavage, citrullination, formylation, N-linked glycosylation, O-linked glycosylation, hydroxylation, methylation, myristoylation, neddylation, nitration, oxidation, palmitoylation, phosphorylation, prenylation, S-nitrosylation, sulfation, sumoylation, and ubiquitination.

[0061] In some embodiments, assessing the identity of a terminal amino acid type involves determining whether an amino acid contains a side chain characterized by one or more biochemical properties. For example, an amino acid may contain a nonpolar aliphatic side chain, a positively charged side chain, an uncharged side chain, a nonpolar aromatic side chain, or a polar uncharged side chain. Non-limiting examples of amino acids containing a nonpolar aliphatic side chain include alanine, glycine, valine, leucine, methionine, and isoleucine. Non-limiting examples of amino acids containing a positively charged side chain include lysine, arginine, and histidine. Non-limiting examples of amino acids containing an uncharged side chain include aspartic acid and glutamic acid. Non-limiting examples of amino acids containing a nonpolar aromatic side chain include phenylalanine, tyrosine, and tryptophan. Non-limiting examples of amino acids containing a polar uncharged side chain include serine, threonine, cysteine, proline, asparagine, and glutamine.

[0062] In some embodiments, a protein or polypeptide can be digested into a plurality of smaller polypeptides, and sequence information can be obtained from one or more of these smaller polypeptides (for example, by a method comprising sequentially evaluating the terminal amino acids of a polypeptide and removing that amino acid to expose the next amino acid at the end).

[0063] In some embodiments, the polypeptide is sequenced from its amino (N) terminus. In some embodiments, the polypeptide is sequenced from its carboxy (C) terminus. In some embodiments, the first terminus of the polypeptide (e.g., the N or C terminus) is immobilized, and the other terminus (e.g., the C or N terminus) is sequenced as described herein.

[0064] As used herein, sequencing a polypeptide means determining the sequence information of the polypeptide. In some embodiments, this may include determining the identity of each sequential amino acid in a portion (or all) of the polypeptide. However, in some embodiments, this may include evaluating the identity of a subset of amino acids within the polypeptide (e.g., determining the relative positions of one or more amino acid types without determining the identity of each amino acid in the polypeptide). However, in some embodiments, it is possible to obtain amino acid content information from a polypeptide without directly determining the relative positions of various types of amino acids in the polypeptide. Amino acid content alone may be used to infer the identity of a present polypeptide (e.g., by comparing the amino acid content with a polypeptide information database and determining which polypeptides have the same amino acid content).

[0065] In some embodiments, sequence information of multiple polypeptide products obtained from a longer polypeptide or protein (e.g., via enzymatic and / or chemical cleavage) can be analyzed to reconstruct or infer the sequence of the longer polypeptide or protein.

[0066] Therefore, in some embodiments, one or more types of amino acids are identified by detecting the luminescence of one or more labeled recognition molecules that selectively bind to one or more types of amino acids. In some embodiments, one or more types of amino acids are identified by detecting the luminescence of labeled polypeptides.

[0067] The inventors have further recognized and highly valued that the polypeptide sequencing techniques described herein, in particular in contrast to conventional polypeptide sequencing techniques, may include the generation of novel polypeptide sequencing data. Therefore, conventional techniques for analyzing polypeptide sequencing data may be insufficient when applied to data generated using the polypeptide sequencing techniques described herein.

[0068] For example, conventional polypeptide sequencing techniques, including repeated reagent cycling, may generate data associated with individual amino acids of the polypeptide being sequenced. In such cases, since the detected data corresponds to only one amino acid, analyzing the generated data would simply involve determining which amino acids are detected at a particular time. In contrast, the polypeptide sequencing techniques described herein may generate data during the polypeptide degradation process while multiple amino acids of the polypeptide molecule are detected, potentially resulting in data where it is difficult to distinguish between sections of data corresponding to different amino acids of the polypeptide. Therefore, the inventors have developed a novel computational technique for analyzing such generated data using the polypeptide sequencing techniques described herein, which include determining sections of data corresponding to individual amino acids, by segmenting the data into portions corresponding to each amino acid association event, for example. These sections can then be further analyzed to identify the amino acids detected at the time of those individual sections.

[0069] As another example, conventional sequencing techniques, which involve using uniquely identifiable labels for each type of amino acid, may merely involve analyzing which labels are detected at a given time without considering the dynamics of how individual amino acids interact with other molecules. In contrast, the polypeptide sequencing techniques described herein generate data that indicates how amino acids interact with recognition molecules. As discussed above, the data may include a set of characteristic patterns corresponding to association events between amino acids and their respective recognition molecules. Therefore, we have developed a novel computational technique to analyze characteristic patterns to determine the type of amino acid corresponding to that portion of the data, enabling the determination of the amino acid sequence of a polypeptide by analyzing a set of various characteristic patterns.

[0070] In some embodiments, the polypeptide sequencing techniques described herein generate data that indicate how a polypeptide interacts with a binding means while the polypeptide is being degraded by a cleaving means. As discussed above, the data may include a set of characteristic patterns corresponding to the terminal association events of the polypeptide between terminal cleaving means. In some embodiments, the sequencing method described herein involves contacting a single polypeptide molecule with a binding means and a cleaving means, wherein the binding means and the cleaving means are configured to achieve at least 10 association events before a cleavage event. In some embodiments, the means are configured to achieve at least 10 association events between two cleavage events.

[0071] As described herein, in some embodiments, multiple single-molecule sequencing reactions are carried out in parallel in an array of sample wells. In some embodiments, the array includes about 10,000 to about 1,000,000 sample wells. In some realizations, the volume of the sample wells is about 10 -21 Liters ~ approximately 10 -15The volume can be in liters. Since the sample wells have a small volume, it is possible to detect single-molecule events because only about one polypeptide may be present in a sample well at any given time. Statistically, some sample wells may not contain single-molecule sequencing reactions, and some may contain more than one single polypeptide molecule. However, since a considerable number of sample wells may each contain a single-molecule reaction (e.g., at least 30% in some embodiments), single-molecule analysis can be performed in parallel across a large number of sample wells. In some embodiments, the binding and cleaving means are configured to achieve at least 10 association events before cleavage events in at least 10% (e.g., 10–50%, more than 50%, 25–75%, at least 80%, or more) of the sample wells where single-molecule reactions occur. In some embodiments, the binding and cleaving means are configured to achieve at least 10 association events before cleavage events to at least 50% (e.g., more than 50%, at least 80%, or more) of the amino acids of the polypeptide in the single-molecule reaction.

[0072] amino acid recognition molecules In some embodiments, the methods provided herein involve contacting a polypeptide with an amino acid recognition molecule which may or may not contain a label that selectively binds to at least one type of terminal amino acid. In some embodiments, as used herein, a terminal amino acid may mean an amino-terminal amino acid of a polypeptide or a carboxy-terminal amino acid of a polypeptide. In some embodiments, the labeled recognition molecule selectively binds to one type of terminal amino acid rather than other types of terminal amino acids. In some embodiments, the labeled recognition molecule selectively binds to one type of terminal amino acid rather than internal amino acids of the same type. In yet another embodiment, the labeled recognition molecule selectively binds to one type of amino acid at any position on the polypeptide, for example, an amino acid of the same type as a terminal amino acid or an internal amino acid.

[0073] In some embodiments, as used herein, an amino acid type means one of the 20 naturally occurring amino acids or a subset of that type. In some embodiments, an amino acid type means a modified variant of one of the 20 naturally occurring amino acids or a subset of its unmodified and / or modified variants. Examples of modified amino acid variants include, but are not limited to, post-translational modification variants (e.g., acetylation, ADP-ribosylation, caspase cleavage, citrullination, formylation, N-linked glycosylation, O-linked glycosylation, hydroxylation, methylation, myristoylation, NEDDation, nitration, oxidation, palmitoylation, phosphorylation, prenylation, S-nitrosylation, sulfation, SUMOylation, and ubiquitination), chemical modification variants, unnatural amino acids, and proteogenic amino acids, such as selenocysteine ​​and pyrrolicin. In some embodiments, a subset of an amino acid type includes more than 1 and less than 20 amino acids having one or more similar biochemical properties. For example, in some embodiments, the type of amino acid means one type selected from amino acids having charged side chains (e.g., positive and / or uncharged side chains), amino acids having polar side chains (e.g., polar uncharged side chains), amino acids having nonpolar side chains (e.g., nonpolar aliphatic and / or aromatic side chains), and amino acids having hydrophobic side chains.

[0074] In some embodiments, the methods provided herein involve contacting a polypeptide with one or more labeled recognition molecules that selectively bind to one or more types of terminal amino acids. In exemplary and non-limiting examples, if four labeled recognition molecules are used in the methods of the present application, any one of the recognition molecules selectively binds to one type of terminal amino acid different from the other one type to which any of the other three selectively bind (for example, the first recognition molecule binds to the first type, the second recognition molecule binds to the second type, the third recognition molecule binds to the third type, and the fourth recognition molecule binds to the fourth type of terminal amino acid). For the purposes of this discussion, one or more labeled recognition molecules in relation to the methods described herein may be substituted for a set of labeled recognition molecules.

[0075] In some embodiments, the set of labeled recognition molecules includes at least one and up to six labeled recognition molecules. For example, in some embodiments, the set of labeled recognition molecules includes one, two, three, four, five, or six labeled recognition molecules. In some embodiments, the set of labeled recognition molecules includes ten or fewer labeled recognition molecules. In some embodiments, the set of labeled recognition molecules includes eight or fewer labeled recognition molecules. In some embodiments, the set of labeled recognition molecules includes six or fewer labeled recognition molecules. In some embodiments, the set of labeled recognition molecules includes four or fewer labeled recognition molecules. In some embodiments, the set of labeled recognition molecules includes three or fewer labeled recognition molecules. In some embodiments, the set of labeled recognition molecules includes two or fewer labeled recognition molecules. In some embodiments, the set of labeled recognition molecules includes four labeled recognition molecules. In some embodiments, the set of labeled recognition molecules includes at least two and up to 20 (e.g., at least two and up to 10, at least two and up to 8, at least four and up to 20, at least four and up to 10) labeled recognition molecules. In some embodiments, the set of labeled recognition molecules includes more than 20 (e.g., 20-25, 20-30) recognition molecules. However, it should be recognized that any number of recognition molecules can be used according to the methods of the present application adapted to the desired use.

[0076] According to the present invention, in some embodiments, one or more types of amino acids are identified by detecting the luminescence of a labeled recognition molecule. In some embodiments, the labeled recognition molecule includes a recognition molecule that selectively binds to one type of amino acid, and a luminescent label having luminescence associated with the recognition molecule. Thus, luminescence (e.g., luminescence lifetime, luminescence intensity, and other luminescence properties as described elsewhere in this specification) can be associated with the selective binding of the recognition molecule to identify amino acids of a polypeptide. In some embodiments, multiple types of labeled recognition molecules can be used in the manner described herein, each type including a luminescent label having luminescence that is uniquely identifiable among multiple types. Suitable luminescent labels may include luminescent molecules such as fluorophore dyes and are described elsewhere in this specification.

[0077] In some embodiments, one or more types of amino acids are identified by detecting one or more electrical properties of a labeled recognition molecule. In some embodiments, the labeled recognition molecule includes a recognition molecule that selectively binds to one type of amino acid and a conductivity label that associates with the recognition molecule. Thus, one or more electrical properties (e.g., charge, current oscillation color, and other electrical properties) can be associated with the selective binding of the recognition molecule to identify amino acids in a polypeptide. In some embodiments, multiple types of labeled recognition molecules may be used in the manner of the present invention, and each type includes a conductivity label that produces a change in electrical signal (e.g., a change in conductance, e.g., a change in the magnitude of the conductivity pattern and a transition in conductivity) that is uniquely identifiable among the multiple types. In some embodiments, each of the multiple types of labeled recognition molecules includes a conductivity label having a different number of charged groups (e.g., a different number of negative and / or positive charged groups). Therefore, in some embodiments, the conductivity label is a charge label. Examples of charge labels include dendrimers, nanoparticles, nucleic acids, and other polymers having multiple charged groups. In some embodiments, conductivity labels can be uniquely identified by their net charge (e.g., net positive charge or net negative charge), their charge density, and / or their charge base number.

[0078] In some embodiments, the amino acid recognition molecule can be engineered by those skilled in the art using conventionally known techniques. In some embodiments, desirable properties may include the ability to selectively and with high affinity bind to one type of amino acid only when located at the terminal (e.g., N-terminus or C-terminus) of a polypeptide. In yet other embodiments, desirable properties may include the ability to selectively and with high affinity bind to one type of amino acid both when located at the terminal (e.g., N-terminus or C-terminus) and when located internally within the polypeptide. In some embodiments, desirable properties may include the ability to selectively and with low affinity bind to more than one type of amino acid (e.g., with KD of about 50 nM or more, e.g., about 50 nM to about 50 μM, about 100 nM to about 10 μM, about 500 nM to about 50 μM). For example, in some aspects, the present application provides a sequencing method by detecting reversible binding interactions during a polypeptide degradation process. Advantageously, such a method can be carried out using a recognition molecule that reversibly binds with low affinity to more than one type of amino acid (e.g., a subset of amino acid types).

[0079] As used herein, in some embodiments, the terms “selective” and “specific” (and their variations, e.g., selectively, specifically, selectivity, specificity) mean preferential binding interactions. For example, in some embodiments, an amino acid recognition molecule that selectively binds to one type of amino acid preferentially binds to that one type of amino acid over other types of amino acids. A selective binding interaction will typically distinguish one type of amino acid (e.g., one terminal amino acid) from another type of amino acid (e.g., another terminal amino acid) by about 10 to 100 times (e.g., about 1,000 times or more or 10,000 times). Therefore, it should be recognized that a selective binding interaction may mean any binding interaction that makes one type of amino acid more uniquely identifiable than other types of amino acids. For example, in some embodiments, the present application provides a polypeptide sequencing method by obtaining data that indicates the association of one or more amino acid recognition molecules with polypeptide molecules. In some embodiments, the data includes a series of signal pulses corresponding to a series of reversible amino acid recognition molecule binding interactions between the polypeptide molecule and amino acids, and the data can be used to determine the identity of the amino acids. Therefore, in some embodiments, “selective” or “specific” binding interaction means a detection binding interaction that distinguishes one type of amino acid from another type of amino acid.

[0080] In some embodiments, the amino acid recognition molecule does not significantly bind to other types of amino acids for about 10 -6 Less than M (for example, about 10 -7 Less than M, approximately 10 -8 Less than M, approximately 10 -9 Less than M, approximately 10 -10 Less than M, approximately 10 -11 Less than M, approximately 10 -12 Less than M, 10 -16 Dissociation constant (K up to approximately M) D ) binds to one type of amino acid. In some embodiments, the amino acid recognition molecule has a K content of less than about 100 nM, less than about 50 nM, less than about 25 nM, less than about 10 nM, or less than about 1 nM.D It binds to one type of amino acid (for example, one type of terminal amino acid). In some embodiments, the amino acid recognition molecule has a K content of about 50 nM to about 50 μM (for example, about 50 nM to about 500 nM, about 50 nM to about 5 μM, about 500 nM to about 50 μM, about 5 μM to about 50 μM, or about 10 μM to about 50 μM). D It binds to one type of amino acid. In some embodiments, the amino acid recognition molecule has a K250 nM D It then binds to one type of amino acid.

[0081] In some embodiments, the amino acid recognition molecule is approximately 10 -6 Less than M (for example, about 10 -7 Less than M, approximately 10 -8 Less than M, approximately 10 -9 Less than M, approximately 10 -10 Less than M, approximately 10 -11 Less than M, approximately 10 -12 Less than M, 10 -16 K (up to about M size) D It binds to two or more types of amino acids. In some embodiments, the amino acid recognition molecule has a K content of less than about 100 nM, less than about 50 nM, less than about 25 nM, less than about 10 nM, or less than about 1 nM. D It binds to two or more types of amino acids. In some embodiments, the amino acid recognition molecule has a K content of about 50 nM to about 50 μM (for example, about 50 nM to about 500 nM, about 50 nM to about 5 μM, about 500 nM to about 50 μM, about 5 μM to about 50 μM, or about 10 μM to about 50 μM). D It binds to two or more types of amino acids. In some embodiments, the amino acid recognition molecule has a K250 nM D It binds to two or more types of amino acids.

[0082] In some embodiments, the amino acid recognition molecule lasts for at least 0.1 seconds. -1 Dissociation rate (K off ) binds to at least one type of amino acid. In some embodiments, the dissociation rate is about 0.1 s -1 ~about 1,000s -1 (For example, about 0.5s) -1 ~about 500s-1 , about 0.1 s -1 ~ about 100 s -1 , about 1 s -1 ~ about 100 s -1 , or about 0.5 s -1 ~ about 50 s -1 ) is. In some embodiments, the dissociation rate is about 0.5 s -1 ~ about 20 s -1 . In some embodiments, the dissociation rate is about 2 s -1 ~ about 20 s -1 . In some embodiments, the dissociation rate is about 0.5 s -1 ~ about 2 s -1 is.

[0083] In some embodiments, K D or k off 's value may be the value of known literature, or the value may be determined experimentally. For example, K D or k off 's value can be measured in a single-molecule assay or an ensemble assay (see, for example, Example 4 and FIG. 19M). In some embodiments, the value of k off can be determined experimentally based on the signal pulse information obtained in a single-molecule assay as described elsewhere in this specification. For example, the value of k off can be estimated by the reciprocal of the average pulse duration. In some embodiments, the amino acid recognition molecule binds to two or more types of amino acids with different K D or k off for each of the two or more types. In some embodiments, the first K D or k off for the first type of amino acid is at least 10% (e.g., at least 25%, at least 50%, at least 100%, or more) different from the second K D or k off for the second type of amino acid. In some embodiments, K D or k offThe first and second values ​​differ by approximately 10-25%, 25-50%, 50-75%, 75-100%, or more than 100%, for example, by about 2, 3, 4, 5, or more.

[0084] According to the methods and compositions provided herein, Figure 2 shows various configurations and uses of labeled recognition molecules. In some embodiments, the labeled recognition molecule 200 comprises a luminescent label 210 (e.g., a label) and a recognition molecule (shown as a spotted shape) that selectively binds to one or more terminal amino acids of polypeptide 220. In some embodiments, the recognition molecule is selective for one type of amino acid or a subset of amino acid types (e.g., fewer than 20 common types of amino acids) at the terminal position or both terminal and internal positions.

[0085] As described herein, an "amino acid recognition molecule" can be any biomolecule capable of selectively or specifically binding one molecule to another (for example, one type of amino acid to another type of amino acid). In some embodiments, the recognition molecule is not a peptidase or does not have peptidase activity. For example, in some embodiments, the polypeptide sequencing method of the present application comprises contacting a polypeptide molecule with one or more recognition molecules and a cleavage reagent. In such embodiments, one or more recognition molecules do not have peptidase activity, and the removal of one or more amino acids from the polypeptide molecule (for example, removal of amino acids from the terminus of the polypeptide molecule) is carried out by the cleavage reagent.

[0086] Recognition molecules include, for example, proteins and nucleic acids, which may be synthetic or recombinant. In some embodiments, the recognition molecule is an antibody or the antigen-binding portion of an antibody, an SH2 domain-containing protein or a fragment thereof, or an enzyme biomolecule, such as a peptidase, aminotransferase, ribozyme, aptazyme, or tRNA synthetase. This may include aminoacyl-tRNA synthetases and related molecules described in U.S. Patent Application No. 15 / 255,433, filed September 2, 2016, titled "AND PROCESSING."

[0087] In some embodiments, the Application relates to the discovery and development of amino acid recognition molecules for use according to the methods described herein or methods known in the Art. In some embodiments, the Application provides a binding amino acid-binding protein (e.g., a ClpS protein) that is not previously known to exist between other homologous members of a protein family. In some embodiments, the Application provides an engineered amino acid-binding protein. For example, in some embodiments, the Application provides a fusion construct comprising a single polypeptide having serial copies of two or more amino acid-binding proteins.

[0088] The inventors recognize and appreciate that the fusion constructs of the present invention enable an effective increase in the concentration of the recognition molecule without increasing the labeled background noise (e.g., background fluorescence). The inventors also recognize and appreciate that the fusion constructs of the present invention increase the accuracy of sequencing reactions and / or reduce the amount of time required to carry out sequencing reactions. Furthermore, by providing fusion constructs having two or more different serial copies of amino acid-binding proteins, fewer reagents are required for the reaction, thus providing a more efficient and inexpensive approach to sequencing.

[0089] In some embodiments, the recognition molecule of the present invention is a degradation pathway protein. Examples of degradation pathway proteins suitable for use as a recognition molecule include, but are not limited to, N-terminal rule pathway proteins, such as Arg / N-terminal rule pathway proteins, Ac / N-terminal rule pathway proteins, and Pro / N-terminal rule pathway proteins. In some embodiments, the recognition molecule is an N-terminal rule pathway protein selected from Gid proteins (e.g., Gid4 or Gid10 proteins), UBR box proteins (e.g., UBR1, UBR2) or protein fragments containing the UBR box domain, p62 proteins or fragments containing the ZZ domain, and ClpS proteins (e.g., ClpS1, ClpS2). Therefore, in some embodiments, the labeled recognition molecule 200 comprises a degradation pathway protein. In some embodiments, the labeled recognition molecule 200 is a ClpS protein.

[0090] In some embodiments, the recognition molecule of the present invention is a ClpS protein, for example, Agrobacterium tumifaciens ClpS1, Agrobacterium tumifaciens The recognition molecules are tumifaciens)ClpS2, Synechococcus elongatus ClpS1, Synechococcus elongatus ClpS2, Thermosynechococcus elongatus ClpS, Escherichia coli ClpS, or Plasmodium falciparum ClpS. In some embodiments, the recognition molecule is an L / F transferase, for example, Escherichia coli leucyl / phenylalanyl tRNA protein transferase. In some embodiments, the recognition molecule is a D / E leucyl transferase, for example, Vibrio vulnificus aspartate / glutamate leucyl transferase Bpt. In some embodiments, the recognition molecule is a UBR protein or UBR box domain, for example, a UBR protein or the UBR box domain of human UBR1 and UBR2 or Saccharomyces cerevisiae UBR1. In some embodiments, the recognition molecule is a p62 protein, for example, H. sapiens p62 protein or Rattus norvegicus p62 protein or their truncated variants containing at least a ZZ domain. In some embodiments, the recognition molecule is a Gid4 protein, for example, H. sapiens GID4 or Saccharomyces cerevisiae GID4. In some embodiments, the recognition molecule is a Gid10 protein, for example, Saccharomyces cerevisiae GID10.In some embodiments, the recognition molecule is an N-meristoyltransferase, such as Leishmania major N-meristoyltransferase or H. sapiens N-meristoyltransferase (NMT1). In some embodiments, the recognition molecule is a BIR2 protein, such as Drosophila melanogaster BIR2. In some embodiments, the recognition molecule is a tyrosine kinase or the SH2 domain of a tyrosine kinase, such as the H. sapiens FynSH2 domain, the H. sapiens Src tyrosine kinase SH2 domain, or a variant thereof, such as the H. sapiens FynSH2 domain triple mutant superbinder. In some embodiments, the recognition molecule is an antibody or antibody fragment, such as a single-chain antibody variable fragment (scFv) against phosphotyrosine or other post-translational modified amino acid variants described herein.

[0091] Tables 1 and 2 provide a list of example sequences of amino acid recognition molecules. Unless otherwise specified in Tables 1 and 2, the amino acid binding priority of each molecule to the amino acid identity at the terminal position of the polypeptide is also indicated. These sequences and other examples described herein are intended to be non-limiting, and it should be recognized that the recognition molecules relating to this application may include homologs, variants, or fragments thereof that contain at least the peptide recognition domain or subdomain.

[0092] [Table 1-1]

[0093] [Table 1-2]

[0094] Table 1-3

[0095] Table 1-4

[0096] Table 1-5

[0097] Table 1-6

[0098] Table 1-7

[0099] Table 1-8

[0100] Table 2-1

[0101] Table 2-2

[0102] Table 2-3

[0103] Table 2-4

[0104] [Table 2-5]

[0105] [Table 2-6]

[0106] [Table 2-7]

[0107] Therefore, in some embodiments, the present application provides an amino acid recognition molecule having an amino acid sequence selected from Table 1 or Table 2 (or an amino acid sequence having at least 50%, at least 60%, at least 70%, at least 80%, 80-90%, 90-95%, 95-99%, or more amino acid sequence identity with an amino acid sequence selected from Table 1 or Table 2). In some embodiments, the amino acid recognition molecule has 25-50%, 50-60%, 60-70%, 70-80%, 80-90%, 90-95%, or 95-99%, or more amino acid sequence identity with respect to the amino acid recognition molecules listed in Table 1 or Table 2. In some embodiments, the amino acid recognition molecule is a modified amino acid recognition molecule and includes the deletion, addition, or mutation of one or more amino acids with respect to the sequences shown in Table 1 or Table 2. In some embodiments, the modified amino acid recognition molecule includes deletions, additions, or mutations of 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50 or more amino acids (which may or may not be consecutive) to the sequences shown in Table 1 or Table 2.

[0108] In some embodiments, an amino acid recognition molecule comprises a single polypeptide having two or more series copies of amino acid-binding proteins (e.g., two or more binders). As used herein, in some embodiments, the series arrangement or orientation of elements within a molecule means that each element is terminally linked to the next element in a linear fusion so that the elements fuse sequentially. For example, in some embodiments, a polypeptide having two series copies of binders means a fusion polypeptide in which the C-terminus of one binder is fused to the N-terminus of the other binder. Similarly, a polypeptide having two or more series copies of binders means a fusion polypeptide in which the C-terminus of a first binder is fused to the N-terminus of a second binder, the C-terminus of the second binder is fused to the N-terminus of a third binder, and so on. Such a fusion polypeptide may comprise multiple copies of the same binder or multiple copies of different binders. In some embodiments, the fusion polypeptide of the present invention has at least two and up to ten binders (for example, at least two binders and up to eight, six, five, four, or three binders). In some embodiments, the fusion polypeptide of the present invention has five or fewer binders (for example, two, three, four, or five binders). Therefore, in some embodiments, the labeled recognition molecule 200 comprises the fusion polypeptide of the present invention.

[0109] In some embodiments, if the expression of a single coding sequence produces a single full-length popipeptide having two or more independent binding sites, the fusion polypeptide is provided by the expression of a single coding sequence that includes a segment encoding a monomeric binder subunit separated by a segment encoding a flexible linker. In some embodiments, one or more monomeric subunits (e.g., binders) are ClpS proteins. In some embodiments, the ClpS subunits may be identical or non-identical. If non-identical, the ClpS subunits may be identifiable variants of the same parental ClpS protein, or they may be derived from different parental ClpS proteins. In some embodiments, the fusion polypeptide comprises one or more ClpS monomers and one or more non-ClpS monomers. In some embodiments, the monomeric subunits include non-ClpS monomers. In some embodiments, the monomeric subunits include one or more degradation pathway proteins. For example, in some embodiments, the monomeric subunit includes one or more of the following: Gid protein, UBR box protein or a protein fragment containing its UBR box domain, p62 protein or a protein fragment containing its ZZ domain, and ClpS protein (e.g., ClpS1, ClpS2).

[0110] In some embodiments, at least one binder of the fusion polypeptide has an amino acid sequence selected from Table 1 or Table 2 (or has an amino acid sequence having at least 50%, at least 60%, at least 70%, at least 80%, 80-90%, 90-95%, 95-99%, or more amino acid sequence identity with an amino acid sequence selected from Table 1 or Table 2). In some embodiments, each binder of the fusion polypeptide has at least 80% (e.g., 80-90%, 90-95%, 95-99%, or more) amino acid sequence identity with an amino acid sequence selected from Table 1 or Table 2 (or has an amino acid sequence having at least 50%, at least 60%, at least 70%, at least 80%, 80-90%, 90-95%, 95-99%, or more amino acid sequence identity with an amino acid sequence selected from Table 1 or Table 2). In some embodiments, the binder of the fusion polypeptide is modified and includes one or more amino acid deletions, additions, or mutations relative to the sequences shown in Table 1 or Table 2. In some embodiments, the binder of the fusion polypeptide includes deletions, additions, or mutations of 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50 or more amino acids (which may be consecutive or not) relative to the sequences shown in Table 1 or Table 2.

[0111] In some embodiments, the binder of the fusion polypeptide recognizes an identical set of one or more amino acids. In some embodiments, the binder of the fusion polypeptide recognizes a identifiable set of one or more amino acids. In some embodiments, the binder of the fusion polypeptide recognizes a duplicate set of amino acids. In some embodiments, when the binders of the fusion polypeptide recognize the same amino acids, they may recognize amino acids having the same characteristic pulse pattern or amino acids having different characteristic pulse patterns.

[0112] In some embodiments, the binders of the fusion polypeptide are terminally linked either by a covalent bond or by a linker that covalently links the C-terminus of one binder to the N-terminus of the other binder. In relation to the fusion polypeptide of the present application, the linker means one or more amino acids in the fusion polypeptide that link the two binders but do not form a portion of the polypeptide sequence corresponding to either of the two binders. In some embodiments, the linker comprises at least two amino acids (e.g., at least two, three, four, five, six, eight, ten, 15, 25, 50, 100, or more amino acids). In some embodiments, the linker comprises up to five, up to ten, up to fifteen, up to 25, up to fifty, or up to 100 amino acids. In some embodiments, the linker contains about 2 to about 200 amino acids (for example, about 2 to about 100, about 5 to about 50, about 2 to about 20, about 5 to about 20, or about 2 to about 30 amino acids).

[0113] In some embodiments, the application provides nucleic acids encoding a single polypeptide having two or more serial copies of amino acid-binding proteins. In some embodiments, the nucleic acid is an expression construct encoding the fusion polypeptide of the application. In some embodiments, the expression construct encodes a fusion polypeptide having at least two and up to ten binders (e.g., at least two binders and up to eight, six, five, four, or three binders). In some embodiments, the expression construct encodes a fusion polypeptide having five or fewer binders (e.g., two, three, four, or five binders).

[0114] In some embodiments, the amino acid recognition molecule comprises one or more labels. In some embodiments, the one or more labels include luminescence labels or conductivity labels as described elsewhere in the specification. In some embodiments, the one or more labels comprise one or more polyol moieties (e.g., one or more moieties selected from dextran, polyvinylpyrrolidone, polyethylene glycol, polypropylene glycol, polyoxyethylene glycol, and polyvinyl alcohol). For example, in some embodiments, the amino acid recognition molecule is PEGylated. In some embodiments, polyol modification (e.g., PEGylation) may limit nonspecific adhesion to the substrate surface (e.g., sequencing chip). In some embodiments, polyol modification may limit aggregation or interaction between the amino acid recognition molecule and other recognition molecules, between the amino acid recognition molecule and cleavage reagents, or between the amino acid recognition molecule and other species present in the sequencing reaction mixture. PEGylation can be carried out by incubating the recognition molecule (e.g., an amino acid-binding protein such as ClpS protein) with an mPEG4-NHS ester, thereby labeling a primary amine such as a surface-exposed lysine side chain. Other types of PEGs and other methods of polyol modification are known in this field.

[0115] In some embodiments, one or more labels include tag sequences. For example, in some embodiments, an amino acid recognition molecule includes a tag sequence that provides one or more functions other than amino acid binding. In some embodiments, the tag sequence includes at least one biotin ligase recognition sequence that enables biotinylation of the recognition molecule (e.g., incorporation of one or more biotin molecules, including biotin and bisbiotin moieties). In some embodiments, the tag sequence includes two biotin ligase recognition sequences oriented in series. In some embodiments, a biotin ligase recognition sequence means an amino acid sequence recognized by a biotin ligase that catalyzes covalent bonding between the sequence and biotin molecules. Since each biotin ligase recognition sequence in a tag sequence can be covalently bonded to a biotin moiety, a tag sequence having multiple biotin ligase recognition sequences can be covalently bonded to multiple biotin molecules. A region of a tag sequence having one or more biotin ligase recognition sequences can generally be referred to as a biotinylated tag or biotinylated sequence. In some embodiments, bisbiotin or bisbiotin moiety may mean two biotins bound to two biotin ligase recognition sequences oriented in series.

[0116] Examples of additional functional sequences in tag sequences include purification tags, cleavage sites, and other parts useful for the purification and / or modification of recognition molecules. Table 3 provides a non-limiting list of tag sequences, one or more of which may be used in combination with any one of the amino acid recognition molecules of the present invention (for example, in combination with the sequences shown in Table 1). It should be recognized that the tag sequences shown in Table 3 are intended to be non-limiting, and that the recognition molecules of the present invention may contain one or more tag sequences (e.g., His tags and / or biotinylated tags) at the N-terminus or C-terminus of the recognition molecule polypeptide or at an internal position, or the N-terminus and C-terminus may be split, or they may be reconfigured in other ways as practical in the art.

[0117] [Table 3]

[0118] In some embodiments, the recognition molecule of the present invention is an amino acid-binding protein that can be used in conjunction with other types of amino acid-binding molecules, such as peptidases and / or nucleic acid aptomers, in sequencing methods. Peptidases, also called proteases or proteinases, are enzymes that catalyze the hydrolysis of peptide bonds. Peptidases can generally be classified into endopeptidases and exopeptidases, which digest polypeptides into shorter fragments and cleave polypeptide chains internally or at the terminal, respectively. In some embodiments, the labeled recognition molecule 200 includes a peptidase modified to inactivate exopeptidase activity or endopeptidase activity. Thus, the labeled recognition molecule 200 selectively binds to the polypeptide without cleaving amino acids. In yet another embodiment, a peptidase that is not modified to inactivate exopeptidase or endopeptidase activity may be used in conjunction with the amino acid-binding protein of the present invention. For example, in some embodiments, the labeled recognition molecule includes labeled exopeptidase 202.

[0119] According to certain embodiments of the present application, a protein sequencing method may include repeated detection and cleavage of the ends of a polypeptide. In some embodiments, labeled exopeptidase 202 may be used as a single reagent to perform both the detection and cleavage steps of an amino acid. As generally depicted, in some embodiments, labeled exopeptidase 202 has aminopeptidase activity or carboxypeptidase activity to selectively bind to and cleave the N-terminal or C-terminal amino acid of a polypeptide, respectively. In certain embodiments, it should be recognized that labeled exopeptidase 202 may be catalytically inactivated to retain selective binding ability for use as a non-cleaving labeled recognition molecule 200, as described herein.

[0120] Exopeptidases generally require a polypeptide substrate containing at least one free amino group at its amino terminus or a free carboxyl group at its carboxy terminus. In some embodiments, the exopeptidases according to the present application hydrolyze bonds at or near the terminus of a polypeptide. In some embodiments, the exopeptidases hydrolyze bonds of three residues or less from the polypeptide terminus. For example, in some embodiments, a single hydrolysis reaction catalyzed by an exopeptidase cleaves a single amino acid, dipeptide, or tripeptide from the polypeptide terminus.

[0121] In some embodiments, the exopeptidase according to the present invention is an aminopeptidase or carboxypeptidase that cleaves a single amino acid from the amino terminus or carboxy terminus, respectively. In some embodiments, the exopeptidase according to the present invention is a dipeptidylpeptidase or peptidyldipeptidase that cleaves a dipeptide from the amino terminus or carboxy terminus, respectively. In yet another embodiment, the exopeptidase according to the present invention is a tripeptidylpeptidase that cleaves a tripeptide from the amino terminus. The classification and activity of each class or subclass of peptidases are well known and documented in the literature (see, for example, Gurupriya, VS and Roy, SC, "Proteases and Protease Inhibitors in Male Reproduction," Proteases in Physiology and Pathology, pp. 195-216, 2017, and Brix, K. and Stocker, W., "Proteases: Structure and Function," Chapter 1). In some embodiments, the peptidases of this application have more than three amino acids removed from the polypeptide terminus. Therefore, in some embodiments, the peptidase is an endopeptidase that preferentially cleaves at a specific site (e.g., before or after a specific amino acid). In some embodiments, the size of the polypeptide cleavage product of the endopeptidase activity will depend on the distribution of cleavage sites (e.g., amino acids) within the polypeptide being analyzed.

[0122] The exopeptidases relating to this application may be selected or engineered based on the direction of the sequencing reaction. For example, in embodiments of sequencing a polypeptide from the amino terminus to the carboxy terminus, the exopeptidase may include aminopeptidase activity. Conversely, in embodiments of sequencing a polypeptide from the carboxy terminus to the amino terminus, the exopeptidase may include carboxypeptidase activity. Examples of carboxypeptidases that recognize specific carboxy-terminal amino acids and can be used as labeled exopeptidases or inactivated to be used as non-cleaved labeled recognition molecules as described herein are described in the literature (see, for example, Garcia-Guerrero, MC et al., 2018, Proceedings of the National Academy of Sciences (PNAS), Vol. 115, No. 17).

[0123] Peptidases suitable for use as cleavage reagents and / or recognition molecules include aminopeptidases that selectively bind to one or more types of amino acids. In some embodiments, the aminopeptidase recognition molecule is modified to inactivate aminopeptidase activity. In some embodiments, the aminopeptidase cleavage reagent is non-specific and cleaves most or all types of amino acids from the end of the polypeptide. In some embodiments, the aminopeptidase cleavage reagent cleaves one or more types of amino acids from the end of the polypeptide more efficiently compared to other types of amino acids at the end of the polypeptide. For example, the aminopeptidase according to the present application specifically cleaves alanine, arginine, asparagine, aspartic acid, cysteine, glutamine, glutamic acid, glycine, histidine, isoleucine, leucine, lysine, methionine, phenylalanine, proline, selenocysteine, serine, threonine, tryptophan, tyrosine, and / or valine. In some embodiments, the aminopeptidase is proline aminopeptidase. In some embodiments, the aminopeptidase is proline iminopeptidase. In some embodiments, the aminopeptidase is glutamic acid / aspartic acid specific aminopeptidase. In some embodiments, the aminopeptidase is methionine specific aminopeptidase. In some embodiments, the aminopeptidase is the aminopeptidase shown in Table 4. In some embodiments, the aminopeptidase cleavage reagent cleaves the peptide substrate shown in Table 4.

[0124] In some embodiments, the aminopeptidase is a non-specific aminopeptidase. In some embodiments, the non-specific aminopeptidase is a zinc metalloprotease. In some embodiments, the non-specific aminopeptidase is the aminopeptidase shown in Table 5. In some embodiments, the non-specific aminopeptidase cleaves the peptide substrate shown in Table 5.

[0125] Therefore, in some embodiments, the present application provides an aminopeptidase (e.g., an aminopeptidase recognition molecule, an aminopeptidase cleavage reagent) having an amino acid sequence selected from Table 4 or Table 5 (or having an amino acid sequence having at least 50%, at least 60%, at least 70%, at least 80%, 80-90%, 90-95%, 95-99%, or more amino acid sequence identity to the amino acid sequence selected from Table 4 or Table 5). In some embodiments, the aminopeptidase has 25-50%, 50-60%, 60-70%, 70-80%, 80-90%, 90-95%, or 95-99%, or more amino acid sequence identity to the aminopeptidase listed in Table 4 or Table 5. In some embodiments, the aminopeptidase is a modified aminopeptidase and contains one or more amino acid mutations relative to the sequence shown in Table 4 or Table 5.

[0126]

Table 4-1

[0127]

Table 4-2

[0128]

Table 5-1

[0129]

Table 5-2

[0130]

Table 5-3

[0131]

Table 5-4

[0132] [Table 5-5]

[0133] For the purpose of comparing two or more amino acid sequences, the percentage of "sequence identity" (also referred to herein as "amino acid identity") between a first amino acid sequence and a second amino acid sequence can be calculated by dividing [the number of amino acid residues in the first amino acid sequence that are identical to the amino acid residues at the corresponding positions in the second amino acid sequence] by [the total number of amino acid residues in the first amino acid sequence] and multiplying by

[0100] . In this case, each deletion, insertion, substitution, or addition of an amino acid residue in the second amino acid sequence compared to the first amino acid sequence is considered a difference of a single amino acid residue (position). Alternatively, the degree of sequence identity between two amino acid sequences can be calculated using known computer algorithms (for example, by the local homology algorithm in Smith and Waterman, 1970, Advances in Applied Mathematics (Adv.Appl.Math.), Vol. 2, p. 482c; and Needleman and Wunsch, Journal of Molecular Biology (J.Mol.Biol.), 1970). By the homology alignment algorithm in 1998, Vol. 48, p. 443, by the similarity search method in Pearson and Lipman, Proceedings of the National Academy of Sciences (Proc. Natl. Acad. Sci. USA), Vol. 85, p. 2444, or by running any algorithm available as Blast, Clustal, Omega, or other sequence alignment algorithms on a computer, for example, using standard settings. Typically, for the purpose of determining the percentage of "sequence identity" between two amino acid sequences according to the calculation methods outlined above, the amino acid sequence with the largest number of amino acid residues would be considered the "first" amino acid sequence, and the other amino acid sequences would be considered the "second" amino acid sequence.

[0134] Additionally or alternatively, two or more sequences may be evaluated for identity between them. The terms “identical” or “percent “identical” in relation to two or more nucleic acid sequences or amino acid sequences mean two or more identical sequences or subsequences. Two sequences are “substantially identical” if, when compared using one of the above sequence comparison algorithms or by manual alignment and visual inspection across a comparison window or specified region to obtain the greatest possible correspondence, the two sequences have a specific percentage of identical amino acid residues or nucleotides across a specified region or across the entire sequence (for example, at least 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, 99.5%, 99.6%, 99.7%, 99.8%, or 99.9% identical). Selectively, identity exists over regions of at least approximately 25, 50, 75, or 100 amino acids in length, or over regions of 100–150, 150–200, 100–200, or 200 amino acids or longer.

[0135] Additionally or alternatively, two or more sequences may be evaluated for alignment between them. The terms “alignment” or “percent alignment” in relation to two or more nucleic acid sequences or amino acid sequences mean two or more identical sequences or subsequences. Two sequences are “substantially aligned” if, when aligned to obtain the greatest possible correspondence by comparing them across a comparison window or specified region measured using one of the above sequence comparison algorithms or by manual alignment and visual inspection, the two sequences have a specific percentage of identical amino acid residues or nucleotides across a specified region or across the entire sequence (for example, at least 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, 99.5%, 99.6%, 99.7%, 99.8%, or 99.9% identical). Selectively, alignments may be present over regions of at least approximately 25, 50, 75, or 100 amino acids in length, or over regions of 100–150, 150–200, 100–200, or 200 amino acids or longer.

[0136] In addition to protein molecules, nucleic acid molecules possess various properties that make them advantageous for use as amino acid recognition molecules according to the present invention. Nucleic acid aptamers are nucleic acid molecules that have been engineered to bind to a desired target with high affinity and selectivity. Therefore, nucleic acid aptamers can be engineered to selectively bind to a desired type of amino acid using selection and / or enrichment techniques known in the art. Thus, in some embodiments, the recognition molecule includes a nucleic acid aptamer (e.g., a DNA aptamer, an RNA aptamer). As shown in Figure 2, in some embodiments, the labeled recognition molecule is a labeled aptamer 204 that selectively binds to one type of terminal amino acid. For example, in some embodiments, the labeled aptamer 204 selectively binds to one type of amino acid at the terminal of a polypeptide (e.g., a single type of amino acid or a subset of amino acid types) as described herein. It should be recognized that, although not shown, the labeled aptamer 204 can be engineered to selectively bind to one type of amino acid at any position in a polypeptide (e.g., the terminal position of the polypeptide or both terminal and internal positions) according to the methods of this application.

[0137] In some embodiments, the labeled recognition molecule comprises a label having binding-induced luminescence. For example, in some embodiments, the labeled aptamer 206 comprises a donor label 212 and an acceptor label 214 and functions as illustrated in panels (I) and (II) of FIG. 2. As depicted in panel (I), the labeled aptamer 206 as a free molecule adopts a conformation in which the donor label 212 and the acceptor label 214 are separated by a distance (e.g., about 10 nm or more) that restricts detectable FRET between the labels. As depicted in panel (II), the labeled aptamer 206 as a selective binding molecule adopts a conformation in which the donor label 212 and the acceptor label 214 are within a distance (e.g., about 10 nm or less) that promotes detectable FRET between the labels. In still other embodiments, the labeled aptamer 206 comprises a quenching moiety and functions similar to a molecular beacon, where the luminescence of the labeled aptamer 206 is internally quenched as a free molecule and restored as a selective binding molecule (see, e.g., Hamaguchi et al., Analytical Biochemistry, Vol. 294, pp. 126-131, 2001). Without wishing to be bound by theory, these and other types of mechanisms of binding-induced luminescence are thought to advantageously reduce or eliminate background luminescence and increase the overall sensitivity and accuracy of the methods described herein.

[0138] Shielded Recognition Molecules According to the embodiments described herein, a single molecule polypeptide sequencing method can be performed by illuminating a surface-immobilized polypeptide with excitation light and detecting luminescence generated by a label attached to an amino acid recognition molecule. In some cases, the radiative and / or non-radiative decay generated by the label can cause photo-damage to the polypeptide. For example, FIG. 3A illustrates an example of a sequencing reaction in which a recognition molecule associates with a polypeptide immobilized on a surface.

[0139] In the presence of excitation illumination, the label can generate fluorescence via radiative decay, leading to a detectable association event. However, in some cases, the label produces non-radiative decay, which may lead to the formation of reactive oxygen species 300. Reactive oxygen species 300 can ultimately damage the immobilized peptide, in which case the reaction will terminate before obtaining the complete sequence information of the polypeptide. This photodamage can occur, for example, at the exposed polypeptide end (top hollow arrow), at internal positions (middle hollow arrow), and at the surface linker that attaches the polypeptide to the surface (bottom hollow arrow).

[0140] The inventors have found that photodamage can be reduced and recognition time can be extended by incorporating the shielding element into the amino acid recognition molecule. Figure 3B illustrates an example of a sequencing reaction using a shield recognition molecule containing shield 302. Shield 302 forms a covalent or non-covalent linking group that provides an increased distance between the label and the polypeptide, thereby reducing the damaging effect from reactive oxygen species 300 due to free radical decay over the label-polypeptide separation distance. Shield 302 can also provide a steric barrier that shields the polypeptide from the label by absorbing damage from reactive oxygen species 300 and radiative and / or non-radiative decay.

[0141] While not intended to be constrained by theory, it is conceivable that a shield positioned between the recognition component and the labeling component can absorb, deflect, or block the radiation and / or non-radiative attenuation emitted by the labeling component. In some embodiments, the shield prevents or limits the interaction between one or more labels (e.g., luminescent labels) and one or more amino acid recognition molecules. In some embodiments, the shield prevents or limits the interaction between one or more labels and one or more molecules that associate with the amino acid recognition molecules (e.g., polypeptides that associate with the recognition molecules). Therefore, in some embodiments, the term shield may generally mean any protective or shielding effect provided by any portion of the linking group formed between the recognition component and the labeling component.

[0142] In some embodiments, the shield is attached to one or more amino acid-recognizing molecules (e.g., recognition components) and one or more labels (e.g., labeling components). In some embodiments, the recognition components and labeling components are attached to non-proximate sites on the shield. For example, one or more amino acid-recognizing molecules can be attached to a first side of the shield, and one or more labels can be attached to a second side of the shield, in which case the first and second sides of the shield are separated from each other. In some embodiments, the attachment sites are approximately on both sides of the shield.

[0143] The distance between the site where the shield is attached to the recognition molecule and the site where the shield is attached to the label can be measured linearly across space or nonlinearly across the entire surface of the shield. The distance between the recognition molecule binding site and the label binding site on the shield can be measured by modeling the three-dimensional structure of the shield. In some embodiments, this distance can be at least 2 nm, at least 4 nm, at least 6 nm, at least 8 nm, at least 10 nm, at least 12 nm, at least 15 nm, at least 20 nm, at least 30 nm, at least 40 nm, or more. Alternatively, the relative positions of the recognition molecule and the label on the shield can be described by treating the structure of the shield as a quadratic surface (e.g., ellipsoid, ellipsoidal cylinder). In some embodiments, the recognition molecule binding site and the label binding site are separated by a distance of at least 1 / 8 of the distance around the ellipsoidal shape representing the shield. In some embodiments, the recognition molecule and the label are separated by a distance of at least 1 / 4 of the distance around the ellipsoidal shape representing the shield. In some embodiments, the recognition molecule and the label are separated by a distance of at least 1 / 3 of the distance around the ellipsoidal shape representing the shield. In some embodiments, the recognition molecule and the label are separated by a distance of half the distance around the ellipsoidal shape representing the shield.

[0144] The size of the shield should be such that the label cannot or is unlikely to come into direct contact with the polypeptide when the amino acid recognition molecule associates with the polypeptide. The size of the shield should also be such that the attached label is detectable when the amino acid recognition molecule associates with the polypeptide. For example, the size should be such that the attached luminescent label is within the illumination volume from which it is excited.

[0145] It should be recognized that there are various parameters that allow the implementer to evaluate the shielding effect. In general, the effect of a shielding element can be evaluated by performing a comparative assessment between a composition having a shielding element and a composition lacking a shielding element. For example, a shielding element can increase the recognition time of an amino acid recognition molecule. In some embodiments, recognition time means the length of time during which an association event between the recognition molecule and the polypeptide is observable during the polypeptide sequencing reaction described herein. In some embodiments, the recognition time is increased by about 10-25%, 25-50%, 50-75%, 75-100%, or more than 100%, for example, by about 2, 3, 4, 5, or more, compared to a polypeptide sequencing reaction carried out under identical conditions except that the amino acid recognition molecule lacks a shielding element but is otherwise similar or identical. In some embodiments, the shielding element can increase sequencing accuracy and / or sequence read length (for example, by at least 5%, at least 10%, at least 15%, at least 25%, or more compared to sequencing reactions performed under the comparative conditions described above).

[0146] Therefore, in some embodiments, the present application provides a shielded recognition molecule comprising at least one amino acid recognition molecule, at least one detectable label, and a shielding element (e.g., “shield”) that forms a covalent or non-covalent linking group between the recognition molecule and the label. In some embodiments, the shielding element is at least 2 nm, at least 5 nm, at least 10 nm, at least 12 nm, at least 15 nm, at least 20 nm, or longer (e.g., in an aqueous solution). In some embodiments, the shielding element is about 2 nm to about 100 nm in length (e.g., about 2 nm to about 50 nm, about 10 nm to about 50 nm, about 20 nm to about 100 nm).

[0147] In some embodiments, the shield (e.g., shielding element) forms covalent or non-covalent linking groups between one or more amino acid recognition molecules (e.g., recognition components) and one or more labels (e.g., labeling components). As used herein, in some embodiments, covalent and non-covalent linking or linking groups refer to the nature of the attachment of the recognition components and labeling components to the shield.

[0148] In some embodiments, a covalent linkage or covalent linking group refers to a shield attached to each of the recognition and labeling components via a covalent bond or a series of contig covalent bonds. Covalent attachment of one or both components can be achieved by covalent conjugation methods known in the art. For example, in some embodiments, click chemistry techniques (e.g., copper catalysts, strain acceleration, copper-free click chemistry, etc.) can be used to attach one or both components to the shield. Such methods generally involve conjugating one reactive part to another reactive part so as to form one or more covalent bonds between the reactive parts. Thus, in some embodiments, a first reactive part of the shield is brought into contact with a second reactive part of the recognition or labeling component to form a covalent attachment. Examples of reactive parts include, but are not limited to, reactive amines, azides, alkynes, nitrones, alkenes (e.g., cycloalkenes), tetrazines, tetrazoles, and other reactive parts suitable for click reactions and similar coupling techniques.

[0149] In some embodiments, non-covalent linkage or non-covalent linking group means one or more non-covalent coupling means, for example, a shield attached to one or both of the recognition and labeling components via receptor-ligand interactions and oligonucleotide chain hybridization. Examples of receptor-ligand interactions are provided herein, but are not limited to, protein-protein complexes, protein-ligand complexes, protein-aptamer complexes, and aptamer-nucleic acid complexes. Various configurations and strategies for oligonucleotide chain hybridization are described herein and are known in the art (see, for example, U.S. Patent Application Publication No. 2019 / 0024168).

[0150] In some embodiments, the shield 302 includes a polymer such as a biomolecule or a dendritic polymer. Figure 3C shows an example of a polymer shield and the composition of the shield recognition molecule of the present application. The first shield construct 304 shows an example of a protein shield 330. In some embodiments, the protein shield 330 forms a covalent linking group between the recognition molecule and the label. For example, in some embodiments, the protein shield 330 is attached to each of the recognition molecule and the label via one or more covalent bonds, for example, by covalent attachment via the side chains of native or unnatural amino acids of the protein shield 330. In some embodiments, the amino acid recognition molecule includes a single polypeptide having at least one amino acid-binding protein and the protein shield 330 linked terminally.

[0151] Therefore, in some embodiments, the present application provides a shield-recognizing molecule comprising a fusion polypeptide having an amino acid-binding protein and a protein shield linked at their ends (e.g., C-terminus and N-terminus fusion). In some embodiments, the binder and the protein shield are linked at their ends either by covalent bonds or by a linker that covalently links the C-terminus of one protein to the N-terminus of the other protein. In some embodiments, the linker in relation to the fusion polypeptide means one or more amino acids in the fusion polypeptide that link the binder and the protein shield but do not form a portion of the polypeptide sequence corresponding to either the binder or the protein shield. In some embodiments, the linker comprises at least two amino acids (e.g., at least two, three, four, five, six, eight, ten, 15, 25, 50, 100, or more amino acids). In some embodiments, the linker comprises up to five, up to ten, up to fifteen, up to 25, up to fifty, or up to 100 amino acids. In some embodiments, the linker contains about 2 to about 200 amino acids (for example, about 2 to about 100, about 5 to about 50, about 2 to about 20, about 5 to about 20, or about 2 to about 30 amino acids).

[0152] In some embodiments, the protein shield of the fusion polypeptide is a protein having a molecular weight of at least 10 kDa. For example, in some embodiments, the protein shield is a protein having a molecular weight of at least 10 kDa and up to 500 kDa (e.g., about 10 kDa to about 250 kDa, about 10 kDa to about 150 kDa, about 10 kDa to about 100 kDa, about 20 kDa to about 80 kDa, about 15 kDa to about 100 kDa, or about 15 kDa to about 50 kDa). In some embodiments, the protein shield of the fusion polypeptide is a protein containing at least 25 amino acids. For example, in some embodiments, the protein shield is a protein containing at least 25 to 1,000 amino acids (e.g., about 100 to about 1,000 amino acids, about 100 to about 750 amino acids, about 500 to about 1,000 amino acids, about 250 to about 750 amino acids, about 50 to about 500 amino acids, about 100 to about 400 amino acids, or about 50 to about 250 amino acids).

[0153] In some embodiments, the protein shield is a polypeptide comprising one or more tag proteins. In some embodiments, the protein shield is a polypeptide comprising at least two tag proteins. In some embodiments, at least two tag proteins are identical (for example, the polypeptide comprises at least two copies of the tag protein sequence). In some embodiments, at least two tag proteins are different from each other (for example, the polypeptide comprises at least two different tag protein sequences). Examples of tagged proteins are not limited to, but include: Fasciola hepatica 8kDa antigen (Fh8), maltose-binding protein (MBP), N-utilization substance (NusA), thioredoxin (Trx), ubiquitin-like factor (SUMO), glutathione S-transferase (GST), soluble enhancer peptide sequence (SET), IgG domain B1 of protein G (GB1), IgG repeat domain ZZ of protein A (ZZ), mutant dehalogenase (HaloTag), soluble enhancing ubiquitous tag (SNUT), 17-kilodalton protein (Skp), T7 phage protein kinase (T7PK), E. coli selective protein A (EspA), and monomeric bacteriophage T7. Examples include 0.3 protein (Orc protein; Mocr), E. coli trypsin inhibitor (Ecotin), calcium-binding protein (CaBP), stress-responsive arsenate reductase (ArsC), the N-terminus of translation initiation factor IF2 (IF2 domain I), stress-responsive proteins (e.g., RpoA, SlyD, Tsf, RpoS, PotD, Crr), and E. coli acidic proteins (e.g., msyB, msyD, rpoD). For example, see Costa, S. et al., "Fusion tags for protein solubility, purification, and immunogenicity in Escherichia coli: a novel Fh8 system (Fusion)." See "Tags for protein solubility, purification and immunogenicity in Escherichia coli: the novel Fh8 system," Frontier Microbiol., February 19, 2014, Vol. 5, p. 63 (relevant content is incorporated herein by reference).

[0154] As described herein, the shielding elements of this application can favorably absorb, deflect, or block the radiation and / or non-radiative attenuation emitted by the labeling component of the amino acid recognition molecule. Therefore, it should be recognized that suitable protein shields for fusion polypeptides are readily selectable by those skilled in the art. For example, the inventors have demonstrated the use of various types of protein shields in connection with fusion polypeptides, including polypeptides having amino acid-binding proteins fused to enzymes (e.g., DNA polymerase, glutathione S-transferase), transport proteins (e.g., maltose-binding proteins), fluorescent proteins (e.g., GFP), and commercially available tag proteins (SNAP-tag®). The inventors have further demonstrated the use of fusion polypeptides having multiple copies of a series-oriented protein shield.

[0155] Therefore, in some embodiments, the present application provides a fusion polypeptide having one or more series-oriented amino acid-binding proteins fused to one or more series-oriented protein shields. In some embodiments, if the fusion polypeptide comprises two or more series-oriented binders and / or two or more series-oriented shields, one end of the two or more binders is terminally linked to one end of the two or more shields. Fusion polypeptides having series copies of two or more binders are described elsewhere in this specification, and in some embodiments, such fusion further comprises one of the two or more binders and a terminally linked protein shield.

[0156] In some embodiments, the protein shield 330 forms a non-covalent linking group between the recognition molecule and the label. For example, in some embodiments, the protein shield 330 is a monomeric or multimeric protein containing one or more ligand-binding sites. In some embodiments, the non-covalent group is formed via one or more ligand moieties bound to one or more ligand-binding sites. Further examples of non-covalent bonds formed by the protein shield are described elsewhere in this specification.

[0157] The second shield construct 306 illustrates an example of a double-stranded nucleic acid shield comprising a first oligonucleotide chain 332 hybridized to a second oligonucleotide chain 334. As shown, in some embodiments, the double-stranded nucleic acid shield may include a recognition molecule attached to the first oligonucleotide chain 332 and a label attached to the second oligonucleotide chain 334. Thus, the double-stranded nucleic acid shield forms a non-covalent linking group between the recognition molecule and the label via oligonucleotide chain hybridization. In some embodiments, the recognition molecule and the label can be attached to the same oligonucleotide chain, which can provide a single-stranded nucleic acid shield or a double-stranded nucleic acid shield via hybridization with another oligonucleotide chain. In some embodiments, chain hybridization can provide increased rigidity within the linking group to further enhance the separation between the recognition molecule and the label.

[0158] If the shielding element 302 contains nucleic acid, the separation distance between the label and the recognition molecule can be measured by the distance between attachment sites on the nucleic acid (e.g., direct attachment or indirect attachment via one or more additional shielding polymers). In some embodiments, the distance between attachment sites on the nucleic acid can be measured by the number of nucleotides in the nucleic acid present between the label and the recognition molecule. It should be understood that the number of nucleotides can mean either the number of nucleotide bases in a single-stranded nucleic acid or the number of nucleotide base pairs in a double-stranded nucleic acid.

[0159] Therefore, in some embodiments, the attachment sites for the recognition molecule and the attachment sites for the label may be separated by only 5 to 200 nucleotides (e.g., 5 to 150 nucleotides, 5 to 100 nucleotides, 5 to 50 nucleotides, 10 to 100 nucleotides). It should be recognized that any position in the nucleic acid can function as an attachment site for the recognition molecule, the label, or one or more additional polymer shields. In some embodiments, the attachment sites may be at the 5' end, the 3' end, or approximately therein, or at an internal position along the nucleic acid strand.

[0160] A non-limiting configuration of the second shield construct 306 illustrates an example of a shield that forms a non-covalent linkage via chain hybridization. Further examples of non-covalent linkages are illustrated by the third shield construct 308, which includes an oligonucleotide shield 336. In some embodiments, the oligonucleotide shield 336 is a nucleic acid aptamer that binds to a recognition molecule to form a non-covalent linkage. In some embodiments, the recognition molecule is a nucleic acid aptamer, and the oligonucleotide shield 336 includes an oligonucleotide chain that hybridizes to the aptamer to form a non-covalent linkage.

[0161] A fourth shield structure 310 illustrates an example of a dendritic polymer shield 338. As used herein, in some embodiments, dendritic polymer generally means polyol or dendrimer. Polyols and dendrimers are described in the Art and may include branched dendritic structures optimized for specific configurations. In some embodiments, the dendritic polymer shield 338 includes polyethylene glycol, tetraethylene glycol, poly(amidoamine), poly(propyleneimine), poly(propyleneamine), carbosilane, poly(L-lysine), or a combination of one or more thereof.

[0162] Dendrimers or dendrons are repeating, branched molecules that typically have a symmetrical core and can take on a spherical three-dimensional form. See, for example, Astruc et al., 2010, Chem. Rev., Vol. 110, p. 1857. The incorporation of such structures into the shield of this application can provide a protective effect through steric inhibition of contact between the label and one or more biomolecules associated with it (e.g., the recognition molecule and / or the polypeptide associated with the recognition molecule). Refinement of the chemical and physical properties of the dendrimer through variation in the primary structure of the molecule allows for the adjustment of the desired shielding effect, including the potential functionalization of the dendrimer surface. Dendrimers can be synthesized by various techniques using a wide range of materials and branching reactions known in the art. Such synthetic variations allow for the customization of the properties of the dendrimer as needed. Examples of polyol and dendrimer compounds that can be used in accordance with the shield of this application include, but are not limited to, the compounds described in U.S. Patent Application Publication No. 20180346507.

[0163] Figure 3D shows further configuration examples of the shield recognition molecule of the present invention. Protein-nucleic acid construct 312 shows an example of a shield comprising a polymer of one or more protein forms and a double-stranded nucleic acid. In some embodiments, the protein portion of the shield is attached to the nucleic acid portion of the shield via covalent linkages. In some embodiments, the attachment is via non-covalent linkages. For example, in some embodiments, the protein portion of the shield is a monovalent or polyvalent protein that forms at least one non-covalent linkage via a ligand portion attached to a ligand-binding site of the monovalent or polyvalent protein. In some embodiments, the protein portion of the shield comprises an avidin protein.

[0164] In some embodiments, the shield recognition molecule of the present application is an avidin nucleic acid construct 314. In some embodiments, the avidin nucleic acid construct 314 comprises a shield comprising an avidin protein 340 and a double-stranded nucleic acid. As described herein, the avidin protein 340 may be used to form a non-covalent linkage between one or more amino acid recognition molecules and one or more labels, either directly or indirectly, such as through one or more additional shield polymers described herein.

[0165] Avidin proteins are biotin-binding proteins, and generally each of the four subunits of an avidin protein has a biotin-binding site. Examples of avidin proteins include avidin, streptavidin, traptabidin, tamavidin, bladavidin, xenavidin, and their homologs and variants. In some cases, monomeric, dimer, or tetramerized avidin proteins can be used. In some embodiments, the avidin protein of the avidin protein complex is streptavidin in tetramerized form (e.g., homotetramer). In some embodiments, the biotin-binding site of the avidin protein provides attachment sites for one or more amino acid recognition molecules, one or more labels, and / or one or more additional shield polymers as described herein.

[0166] An exemplary diagram of the avidin protein complex is shown in the inset panel of Figure 3D. As shown in the inset panel, the avidin protein 340 may contain binding sites 342 in each of the four subunits of the protein that can bind to the biotin moiety (shown as white circles). The polyvalence of the avidin protein 340 may enable various binding configurations, which are generally shown for illustrative purposes. For example, in some embodiments, a biotin binding portion 344 may be used to provide a single attachment site to the avidin protein 340. In some embodiments, a bisbiotin binding portion 346 may be used to provide two attachment sites to the avidin protein 340. As exemplified by the avidin nucleic acid construct 314, the avidin protein complex may be formed by two bisbiotin binding portions that form a trans configuration to provide an increased separation distance between the recognition molecule and the label.

[0167] Various further examples of avidin protein shield configurations are shown. The first avidin construct 316 shows an example of an avidin shield attached to a recognition molecule via a bisbiotin ligation site and to two labels via individual biotin ligations. The second avidin construct 318 shows an example of an avidin shield attached to two recognition molecules via individual biotin ligations and to a label via a bisbiotin ligation site. The third avidin construct 320 shows an example of an avidin shield attached to two recognition molecules via individual biotin ligations and to a labeled nucleic acid via biotin ligations on each strand of the nucleic acid. The fourth avidin construct 322 shows an example of an avidin shield attached to a recognition molecule and a labeled nucleic acid via individual bisbiotin ligations. As shown, the label is further shielded from the recognition molecule by a dendritic polymer between the label and the nucleic acid. The fifth avidin construct 324 shows an example of an internal label 325 attached to two avidin-shielded recognition molecules. As shown, each recognition molecule is attached to a different avidin protein via a bisbiotin ligation site, and internal label 326 is attached to both avidin proteins via separate bisbiotin ligation sites.

[0168] It should be noted that the examples of shield recognition molecule configurations shown in Figures 3A to 3D are provided for illustrative purposes only. The inventors have conceived of various other shield configurations using one or more different polymers that form covalent or non-covalent links between the recognition component and the labeling component of the shield recognition molecule. As an example, Figure 3E illustrates the modularity of the shield configuration according to the present invention.

[0169] As shown at the top of Figure 3E, a shield recognition molecule generally comprises a recognition component 350, a shielding element 352, and a labeling component 354. For ease of illustration, the recognition component 350 is depicted as a single amino acid recognition molecule, and the labeling component 354 is depicted as a single label.

[0170] It should be recognized that the shield recognition molecule of the present application may comprise a shielding element 352 to which one or more amino acid recognition molecules and one or more labels are attached. If the recognition component 350 comprises more than one recognition molecule, each recognition molecule is attachable to the shielding element 352 at one or more attachment sites on the shielding element 352. In some embodiments, the recognition component 350 comprises a single polypeptide fusion construct having two or more serial copies of amino acid-binding proteins, as described elsewhere in this specification. If the labeling component 354 comprises more than one label, each label is attachable to the shielding element 352 at one or more attachment sites on the shielding element 352. The labeling component 354 is generally shown to have a single attachment point, but is not limited thereto. For example, in some embodiments, an internal label having one or more attachment points may be used to link to one or more recognition components 350 and / or shielding elements 352, as exemplified by the avidin construct 324.

[0171] In some embodiments, the shielding element 352 comprises a protein 360. In some embodiments, the protein 360 is a monovalent or polyvalent protein. In some embodiments, the protein 360 is a monomer or multimer protein, such as a protein homodimer, protein heterodimer, protein oligomer, or other protein molecule. In some embodiments, the shielding element 352 comprises a protein complex formed by proteins non-covalently bound to at least one other molecule. For example, in some embodiments, the shielding element 352 comprises a protein-protein complex 362. In some embodiments, the protein-protein complex 362 comprises one protein molecule specifically bound to another protein molecule. In some embodiments, the protein-protein complex 362 comprises an antibody or antibody fragment (e.g., scFv) bound to an antigen. In some embodiments, the protein-protein complex 362 comprises a receptor bound to a protein ligand. Additional examples of protein-protein complexes, but not limited to, include trypsin-aprotinin, barnase-barster, and cholin E9-Im9 immunoprotein.

[0172] In some embodiments, the shielding element 352 includes a protein-ligand complex 364. In some embodiments, the protein-ligand complex 364 includes a monovalent protein and a non-protein ligand portion. For example, in some embodiments, the protein-ligand complex 364 includes an enzyme bound to a small molecule inhibitor portion. In some embodiments, the protein-ligand complex 364 includes a receptor bound to a non-protein ligand portion.

[0173] In some embodiments, the shielding element 352 includes a multivalent protein complex formed by one or more non-covalently bound multivalent proteins to a ligand moiety. In some embodiments, the shielding element 352 includes an avidin protein complex formed by one or more avidin proteins non-covalently bound to a biotin linkage moiety. Constructions 366, 368, 370, and 372 provide exemplary examples of avidin protein complexes, one or more of which can be incorporated into the shielding element 352.

[0174] In some embodiments, the shielding element 352 includes a two-way avidin complex 366 containing avidin protein bound to two bisbiotin ligators. In some embodiments, the shielding element 352 includes a three-way avidin complex 368 containing avidin protein bound to two biotin ligators and a bisbiotin ligator. In some embodiments, the shielding element 352 includes a four-way avidin complex 370 containing avidin protein bound to four biotin ligators.

[0175] In some embodiments, the shielding element 352 comprises an avidin protein containing one or two non-functional binding sites introduced by engineering. For example, in some embodiments, the shielding element 352 comprises a divalent avidin complex 372 comprising an avidin protein bound to a biotin ligation site in each of two subunits, in which case the avidin protein contains a non-functional ligand binding site 348 in each of the other two subunits. As shown, in some embodiments, the divalent avidin complex 372 comprises a trans-divalent avidin protein, but a cis-divalent avidin protein may be used depending on the desired realization. In some embodiments, the avidin protein is a trivalent avidin protein. In some embodiments, the trivalent avidin protein contains a non-functional ligand binding site 348 in one subunit and is bound to three biotin ligation sites, or to one biotin ligation site and a bis-biotin ligation site in the other subunits.

[0176] In some embodiments, the shielding element 352 comprises a dendritic polymer 374. In some embodiments, the dendritic polymer 374 is a polyol or dendrimer as described elsewhere in this specification. In some embodiments, the dendritic polymer 374 is a branched polyol or branched dendrimer. In some embodiments, the dendritic polymer 374 comprises a monosaccharide-TEG, a disaccharide, an N-acetyl monosaccharide, TEMPO-TEG, Trolox-TEG, or a glycerol dendrimer. Examples of useful polyols relating to the shield recognition molecule of this application include polyether polyols and polyester polyols, such as polyethylene glycol, polypropylene glycol, and similar polymers well known in the art. In some embodiments, the dendritic polymer 374 is of the following formula:-(CH2CH2O) n -The compound comprises (wherein n is an integer from 1 to 500 (including the values ​​at both ends)). In some embodiments, the dendritic polymer 374 is of the following formula:-(CH2CH2O)n - (where n is an integer from 1 to 100 (including both end values)) and includes a compound of.

[0177] In some embodiments, the shielding element 352 includes a nucleic acid. In some embodiments, the nucleic acid is single-stranded. In some embodiments, the labeling component 354 is attached directly or indirectly to one end (e.g., the 5' end or the 3' end) of the single-stranded nucleic acid, and the recognition component 350 is attached directly or indirectly to the other end (e.g., the 3' end or the 5' end) of the single-stranded nucleic acid. For example, the single-stranded nucleic acid may include a label attached to the 5' end of the nucleic acid and an amino acid recognition molecule attached to the 3' end of the nucleic acid.

[0178] In some embodiments, the shielding element 352 includes a double-stranded nucleic acid 376. As shown, in some embodiments, the double-stranded nucleic acid 376 can form a non-covalent linkage between the recognition component 350 and the labeling component 354 via strand hybridization. However, in some embodiments, the double-stranded nucleic acid 376 can form a covalent linkage between the recognition component 350 and the labeling component 354 via attachment to the same oligonucleotide strand. In some embodiments, the labeling component 354 is attached directly or indirectly to one end of the double-stranded nucleic acid, and the recognition component 350 is attached directly or indirectly to the other end of the double-stranded nucleic acid. For example, the double-stranded nucleic acid may include a label attached to the 5' end of one strand and an amino acid recognition molecule attached to the 5' end of the other strand.

[0179] In some embodiments, the shielding element 352 includes a nucleic acid that forms one or more structural motifs that may be useful for increasing the steric bulk of the shield. Examples of nucleic acid structural motifs include, but are not limited to, stem-loops, three-way junctions (e.g., formed by two or more stem-loop motifs), four-way junctions (e.g., Holliday junctions), and bulge loops.

[0180] In some embodiments, the shielding element 352 contains nucleic acids that form a stem-loop 378. A stem-loop, or hairpin loop, is an unpaired loop of nucleotides on an oligonucleotide chain that is formed when the oligonucleotide chain folds to form base pairs with other sections of the same chain. In some embodiments, the unpaired loop of the stem-loop 378 contains 3 to 10 nucleotides. Therefore, the stem-loop 378 can be formed by two regions of an oligonucleotide chain having an inverted complementary sequence that hybridizes to form a stem, in which case the two regions are separated by 3 to 10 nucleotides that form an unpaired loop. In some embodiments, the stem of the stem-loop 378 can be designed to have one or more G / C nucleotides that can provide increased stability due to additional hydrogen bonding interactions that occur compared to A / T nucleotides. In some embodiments, the stem of the stem-loop 378 contains G / C nucleotides in direct proximity to the unpaired loop sequence. In some embodiments, the stem of the stem-loop 378 contains G / C nucleotides within the first 2, 3, 4, or 5 nucleotides adjacent to the unpaired loop sequence. In some embodiments, the unpaired loop of the stem-loop 378 includes one or more attachment sites. In some embodiments, the attachment sites occur at debasement sites in the unpaired loop. In some embodiments, the attachment sites occur at bases in the unpaired loop.

[0181] In some embodiments, the stem-loop 378 is formed by double-stranded nucleic acids. As described herein, in some embodiments, the double-stranded nucleic acids can form a non-covalent linking group via chain hybridization of the first and second oligonucleotide chains. However, in some embodiments, the shielding element 352 provides a covalent linking group comprising single-stranded nucleic acids that form a stem-loop motif. In some embodiments, the shielding element 352 comprises nucleic acids that form two or more stem-loop motifs. For example, in some embodiments, the nucleic acid comprises two stem-loop motifs. In some embodiments, the stems of one stem-loop motif are adjacent to the other stem so that the motifs come together to form a three-way junction. In some embodiments, the shielding element 352 comprises nucleic acids that form a four-way junction 378. In some embodiments, the four-way junction 378 is formed via hybridization of two or more oligonucleotide chains (e.g., two, three, or four oligonucleotide chains).

[0182] In some embodiments, the shielding element 352 comprises one or more polymers selected from 360, 362, 364, 366, 368, 370, 372, 374, 376, 378, and 380 in Figure 3E. It should be recognized that the connecting portions and attachment portions shown for each of 360, 362, 364, 366, 368, 370, 372, 374, 376, 378, and 380 are shown for illustrative purposes only and are not intended to depict preferred configurations of connecting or attachment portions.

[0183] In some embodiments, the present application is based on formula (II) A-(Y) n -D (II) (In the formula, A is an amino acid binding component containing at least one amino acid recognition molecule; each of Y is a polymer that forms a covalent or non-covalent linking group; n is an integer from 1 to 10 (including the values ​​at both ends); and D is a labeling component containing at least one detectable label.) The present application provides an amino acid recognition molecule. In some embodiments, the present application provides a composition comprising a soluble amino acid recognition molecule of formula (II).

[0184] In some embodiments, A comprises a plurality of amino acid recognition molecules. In some embodiments, each of the plurality of amino acid recognition molecules is attached to a different attachment site on Y. In some embodiments, at least two of the plurality of amino acid recognition molecules are attached to a single attachment site on Y. In some embodiments, the amino acid recognition molecules are recognition proteins or nucleic acid aptamers, for example, those described elsewhere in this specification.

[0185] In some embodiments, the detectable label is an luminescent label or a conductivity label. In some embodiments, the luminescent label comprises at least one fluorophore dye molecule. In some embodiments, D comprises 20 or fewer fluorophore dye molecules. In some embodiments, the ratio of the number of fluorophore dye molecules to the number of amino acid recognition molecules is 1:1 to 20:1. In some embodiments, the luminescent label comprises at least one FRET pair comprising a donor label and an acceptor label. In some embodiments, the ratio of the donor label to the acceptor label is 1:1, 2:1, 3:1, 4:1, or 5:1. In some embodiments, the ratio of the acceptor label to the donor label is 1:1, 2:1, 3:1, 4:1, or 5:1.

[0186] In some embodiments, D has a diameter of less than 20 nm (200 Å). In some embodiments, -(Y) n - is at least 2 nm in length. In some embodiments, -(Y) n - is at least 5 nm in length. In some embodiments, -(Y) n- is at least 10 nm in length. In some embodiments, each of Y is independently a biomolecule, polyol, or dendrimer. In some embodiments, the biomolecule is a nucleic acid, polypeptide, or polysaccharide.

[0187] In some embodiments, the amino acid recognition molecule is given by the following formula: AY 1 -(Y) m -D or A-(Y) m -Y 1 -D (In the formula, Y 1 (where m is a nucleic acid or polypeptide, and m is an integer between 0 and 10 (including the values ​​at both ends)) It is one of them.

[0188] In some embodiments, the nucleic acid comprises a first oligonucleotide chain. In some embodiments, the nucleic acid comprises a second oligonucleotide chain hybridized to the first oligonucleotide chain. In some embodiments, the nucleic acid forms a covalent linkage via the first oligonucleotide chain. In some embodiments, the nucleic acid forms a non-covalent linkage via the hybridized first and second oligonucleotide chains.

[0189] In some embodiments, the polypeptide is a monovalent or polyvalent protein. In some embodiments, the monovalent or polyvalent protein forms at least one non-covalent linkage via a ligand moiety attached to a ligand-binding site of the monovalent or polyvalent protein. In some embodiments, A, Y, or D comprises a ligand moiety.

[0190] In some embodiments, the amino acid recognition molecule is given by the following formula: A-(Y) m -Y 2 -D or AY 2 -(Y) m -D (In the formula, Y 2 (where m is a polyol or dendrimer, and m is an integer between 0 and 10 (including the values ​​at both ends)) It is one of the following. In some embodiments, the polyol or dendrimer includes polyethylene glycol, tetraethylene glycol, poly(amidoamine), poly(propyleneimine), poly(propyleneamine), carbosilane, poly(L-lysine), or one or more of these.

[0191] In some embodiments, the present application is formula (III): AY 1 -D (III) (wherein A is an amino acid binding component containing at least one amino acid recognition molecule; Y 1 (where D is a nucleic acid or polypeptide; D is a labeled component containing at least one detectable label.) It provides an amino acid recognition molecule. In some embodiments, Y 1 When is a nucleic acid, the nucleic acid forms covalent or non-covalent linking groups. In some embodiments, Y 1 When it is a polypeptide, the polypeptide is 50 × 10 -9 Dissociation constant less than M (K D It forms a non-covalent linking group characterized by ).

[0192] In some embodiments, Y 1 This is a nucleic acid comprising a first oligonucleotide chain. In some embodiments, the nucleic acid comprises a second oligonucleotide chain hybridized to the first oligonucleotide chain. In some embodiments, A is attached to the first oligonucleotide chain and D is attached to the second oligonucleotide chain. In some embodiments, A is attached to a first attachment site on the first oligonucleotide chain and D is attached to a second attachment site on the first oligonucleotide chain. In some embodiments, each oligonucleotide chain of the nucleic acid contains less than 150, less than 100, or less than 50 nucleotides.

[0193] In some embodiments, Y 1A is a monovalent or polyvalent protein. In some embodiments, the monovalent or polyvalent protein forms at least one non-covalent linkage via a ligand moiety attached to the ligand-binding site of the monovalent or polyvalent protein. In some embodiments, at least one of A and D includes a ligand moiety. In some embodiments, the polypeptide is an avidin protein (e.g., avidin, streptavidin, traptabidine, tamavidin, bladavidin, xenavidin, or homologs or variants thereof). In some embodiments, the ligand moiety is a biotin moiety.

[0194] In some embodiments, the amino acid recognition molecule is given by the following formula: AY 1 -(Y) n -D or A-(Y) n -Y 1 -D (In the formula, each of Y is a polymer that forms a covalent or non-covalent linking group, and n is an integer between 1 and 10 (including the values ​​at both ends).) It is one of these. In some embodiments, each of Y is independently a biomolecule, a polyol, or a dendrimer.

[0195] In other embodiments, the present application provides an amino acid recognition molecule comprising a nucleic acid, at least one amino acid recognition molecule attached to a first attachment site on the nucleic acid, and at least one detectable label attached to a second attachment site on the nucleic acid. In some embodiments, the nucleic acid forms a covalent or non-covalent linkage group between the at least one amino acid recognition molecule and the at least one detectable label.

[0196] In some embodiments, the nucleic acid is a double-stranded nucleic acid comprising a first oligonucleotide chain hybridized to a second oligonucleotide chain. In some embodiments, the first attachment site is on the first oligonucleotide chain, and the second attachment site is on the second oligonucleotide chain. In some embodiments, at least one amino acid recognition molecule is attached to the first attachment site via a protein that forms a covalent or non-covalent linkage between at least one amino acid recognition molecule and the nucleic acid. In some embodiments, at least one detectable label is attached to the second attachment site via a protein that forms a covalent or non-covalent linkage between at least one detectable label and the nucleic acid. In some embodiments, the first and second attachment sites are isolated from only 5 to 100 nucleotide bases or nucleotide base pairs on the nucleic acid.

[0197] In yet another embodiment, the present invention provides an amino acid recognition molecule comprising a polyvalent protein having at least two ligand-binding sites, at least one amino acid recognition molecule attached to the protein via a first ligand moiety bound to a first ligand-binding site on the protein, and at least one detectable label attached to the protein via a second ligand moiety bound to a second ligand-binding site on the protein.

[0198] In some embodiments, the polyvalent protein is an avidin protein containing four ligand-binding sites. In some embodiments, the ligand-binding sites are biotin-binding sites, and the ligand moiety is a biotin moiety. In some embodiments, at least one of the biotin moieties is a bisbiotin moiety, and the bisbiotin moiety is bound to two biotin-binding sites on the avidin protein. In some embodiments, at least one amino acid recognition molecule is attached to the protein via a nucleic acid containing the first ligand moiety. In some embodiments, at least one detectable label is attached to the protein via a nucleic acid containing the second ligand moiety.

[0199] As described elsewhere in this specification, the shield recognition molecules of this application may be used in the polypeptide sequencing method relating to this application or in any method known in the art. For example, in some embodiments, the shield recognition molecules provided herein may be used in Edman-type degradation reactions provided herein or conventionally known in the art, which may involve repeated cycling of multiple reaction mixtures in a polypeptide sequencing reaction. In some embodiments, the shield recognition molecules provided herein may be used in the dynamic sequencing reaction of this application, which involves amino acid recognition and degradation in a single reaction mixture.

[0200] Polypeptide sequencing In addition to methods for identifying the terminal amino acids of a polypeptide, this application provides a method for sequencing a polypeptide using a labeled recognition molecule. In some embodiments, the sequencing method may include exposing the polypeptide ends to a repeated cycle of terminal amino acid detection and terminal amino acid cleavage. For example, in some embodiments, this application provides a method for determining the amino acid sequence of a polypeptide, comprising contacting the polypeptide with one or more labeled recognition molecules described herein and exposing the polypeptide to Edman degradation.

[0201] Conventional Edman degradation involves a repetitive cycle of modifying and cleaving the terminal amino acids of a polypeptide, with each sequentially cleaved amino acid being identified to determine the polypeptide's amino acid sequence. As in exemplary examples of conventional Edman degradation, the N-terminal amino acid of a polypeptide is modified using phenyl isothiocyanate (PITC) to form a PITC-induced N-terminal amino acid. The PITC-induced N-terminal amino acid is then cleaved using acidic, basic, and / or high temperatures. It has also been shown that the cleavage of the PITC-induced N-terminal amino acid may be achieved enzymatically using a modified cysteine ​​protease derived from the protozoan Trypanosoma cruzi, including relatively mild cleavage conditions at neutral or near-neutral pH. A non-limiting example of useful enzymes is described in "Molecules and Methods for Repetitive Polypeptide Analysis and Processing." This is described in U.S. Patent Application No. 15 / 255,433, filed on September 2, 2016, entitled "METHODS FOR ITERATIVE POLYPEPTIDE ANALYSIS AND PROCESSING".

[0202] An example of sequencing by Edman degradation using the labeled recognition molecule according to this application is shown in Figure 4. In some embodiments, sequencing by Edman degradation involves providing polypeptide 420 immobilized on the surface 430 of a solid support (e.g., the bottom surface or sidewall surface of a sample well) via a linker 424. In some embodiments, as described herein, polypeptide 420 is immobilized at one end (e.g., an amino-terminal amino acid or a carboxy-terminal amino acid), so that the other end is free for detection and cleavage of the terminal amino acid. Therefore, in some embodiments, the reagents used in the Edman degradation method described herein selectively interact with the terminal amino acid at the unimmobilized (e.g., free) end of polypeptide 420. Thus, polypeptide 420 remains immobilized over repeated detection and cleavage cycles. For this purpose, in some embodiments, the linker 424 may be designed according to a desired set of conditions used for detection and cleavage, for example, to restrict the detachment of polypeptide 420 from the surface 430 under chemical cleavage conditions. Suitable linker compositions and techniques for immobilizing polypeptides on a surface are described elsewhere in this specification.

[0203] According to the present invention, in some embodiments, the sequencing method by Edman degradation comprises step (I) of contacting peptide 420 with one or more labeled recognition molecules that selectively bind to one or more terminal amino acids. As shown, in some embodiments, the labeled recognition molecule 400 interacts with polypeptide 420 by selectively binding to terminal amino acids. In some embodiments, step (I) further comprises removing one or more labeled recognition molecules that do not selectively bind to terminal amino acids (e.g., free amino acids) of polypeptide 420.

[0204] In some embodiments, the method further includes identifying the terminal amino acids of polypeptide 420 by detecting the labeled recognition molecule 400. In some embodiments, the detection includes detecting luminescence from the labeled recognition molecule 400. As described herein, in some embodiments, the luminescence is uniquely associated with the labeled recognition molecule 400, and therefore the luminescence is associated with the type of amino acid to which the labeled recognition molecule 400 selectively binds. Thus, in some embodiments, this type of amino acid is identified by determining the luminescence of one or more labeled recognition molecules 400.

[0205] In some embodiments, the sequencing method by Edman degradation includes step (2) of removing terminal amino acids from polypeptide 420. In some embodiments, step (2) includes removing labeled recognition molecule 400 (e.g., any one or more labeled recognition molecules that selectively bind to terminal amino acids) from polypeptide 420. In some embodiments, step (2) includes modifying terminal amino acids (e.g., favorable amino acids) of polypeptide 420 by contacting the terminal amino acids with an isothiocyanate (e.g., PITC) to form isothiocyanate-modified terminal amino acids. In some embodiments, isothiocyanate-modified terminal amino acids are more susceptible to removal by cleavage reagents (e.g., chemical or enzymatic cleavage reagents) than unmodified terminal amino acids.

[0206] In some embodiments, step (2) includes removing terminal amino acids by contacting polypeptide 420 with protease 440 which specifically binds to and cleaves isothiocyanate-modified terminal amino acids. In some embodiments, protease 440 includes modified cysteine ​​proteases. In some embodiments, protease 440 includes modified cysteine ​​proteases such as cysteine ​​protease derived from Trypanosoma cruzi (see, for example, Borgo et al., 2015, Protein Science, vol. 24, pp. 571-579). In yet other embodiments, step (2) includes removing terminal amino acids by exposing polypeptide 420 to chemical conditions (e.g., acidic, basic) sufficient to cleave the isothiocyanate-modified terminal amino acids.

[0207] In some embodiments, the sequencing method by Edman degradation includes step (3), which involves washing polypeptide 420 after terminal amino acid cleavage. In some embodiments, the washing includes removing protease 440. In some embodiments, the washing includes restoring polypeptide 420 to neutral pH conditions (for example, after chemical cleavage under acidic or basic conditions). In some embodiments, the sequencing method by Edman degradation includes repeating steps (1) to (3) for multiple cycles.

[0208] In some embodiments, the present application provides a method for real-time polypeptide sequencing by evaluating the binding interaction between terminal amino acids and labeled amino acid recognition molecules and labeled cleavage molecules (e.g., labeled nonspecific exopeptidases). Figure 5 shows an example of a sequencing method in which individual association events generate signal pulses of signal output 500. The inset panel of Figure 5 illustrates a general scheme of real-time sequencing by this approach. As shown, the labeled recognition molecule 510 selectively associates (e.g., binds) with a terminal amino acid (indicated as lysine in the specification) and dissociates from the terminal amino acid. This generates a series of pulses in signal pulse 500 that can be used to identify the terminal amino acid. In some embodiments, the series of pulses provides a pulse pattern (e.g., a characteristic pattern) that can serve as a diagnosis of the identity of the corresponding terminal amino acid.

[0209] Although we do not wish to be constrained by theory, the labeled recognition molecule 510 has an association rate, i.e., a binding "on" rate (k on ) and the dissociation rate, i.e., the "off" rate of the bond (k off ) is defined by binding affinity (K D It selectively couples according to the rate constant k. off and k on These are important determinants of pulse duration (e.g., the time corresponding to a detectable meeting event) and inter-pulse duration (e.g., the time between detectable meeting events). In some embodiments, these rates can be engineered to achieve pulse duration and pulse rate (e.g., frequency of signal pulses) that give the best sequencing accuracy.

[0210] As shown in the inset panel, the sequencing reaction mixture further comprises labeled nonspecific exopeptidase 520, which has a different luminescence label than that of labeled recognition molecule 510. In some embodiments, labeled nonspecific exopeptidase 520 is present in the mixture at a lower concentration than that of labeled recognition molecule 510. In some embodiments, labeled specific exopeptidase 520 exhibits broad specificity to cleave almost or all types of amino acids. Therefore, a dynamic sequencing approach may involve monitoring the recognition molecule bound to the polypeptide terminal throughout the entire process of the degradation reaction catalyzed by exopeptidase cleavage activity.

[0211] As illustrated by the progression of signal output 500, in some embodiments, terminal amino acid cleavage by labeled nonspecific exopeptidase 520 generates signal pulses, and these events occur less frequently than binding pulses of labeled recognition molecule 510. Thus, the amino acids of the polypeptide can be counted and / or identified in a real-time sequencing process. As further illustrated in signal output 500, in some embodiments, multiple labeled recognition molecules may be used, each having a diagnostic pulse pattern (e.g., a characteristic pattern) that can be used to identify the corresponding terminal amino acids. For example, in some embodiments, different characteristic patterns (as illustrated by lysine, phenylalanine, and glutamine, respectively, in signal output 500) correspond to the association of one or more labeled recognition molecules having different types of terminal amino acids. It should be recognized that, as described herein, single recognition molecules that associate with more than one type of amino acid may be used in accordance with the application. Therefore, in some embodiments, different characteristic patterns correspond to the association of one labeled recognition molecule with different types of terminal amino acids.

[0212] As described herein, signal pulse information can be used to identify amino acids based on a characteristic pattern of a series of signal pulses. In some embodiments, the characteristic pattern includes a plurality of signal pulses, each of which includes a pulse duration. In some embodiments, the plurality of signal pulses can be characterized by key statistics (e.g., mean, median, time decay constant) of the distribution of pulse durations in the characteristic pattern. In some embodiments, the average pulse duration of the characteristic pattern is about 1 millisecond to about 10 seconds (e.g., about 1 ms to 1 s, about 1 ms to about 100 ms, about 1 ms to about 10 ms, about 10 ms to about 10 s, about 100 ms to about 10 s, about 1 s to about 10 s, about 10 ms to about 100 ms, or about 100 ms to about 500 ms). In some embodiments, the average pulse duration is about 50 milliseconds to about 2 seconds, about 50 milliseconds to about 500 milliseconds, or about 500 milliseconds to about 2 seconds.

[0213] In some embodiments, different characteristic pulse patterns corresponding to different types of amino acids in a single polypeptide are distinguishable from each other based on statistically significant differences in key statistics. For example, in some embodiments, one characteristic pattern is distinguishable from the other characteristic pattern based on a difference in average pulse duration of at least 10 milliseconds (e.g., about 10 ms to about 10 s, about 10 ms to about 1 s, about 10 ms to about 100 ms, about 100 ms to about 10 s, about 1 s to about 10 s, or about 100 ms to about 1 s). In some embodiments, the difference in average pulse duration is at least 50 ms, at least 100 ms, at least 250 ms, at least 500 ms, or more. In some embodiments, the difference in average pulse duration is about 50 ms to about 1 s, about 50 ms to about 500 ms, about 50 ms to about 250 ms, about 100 ms to about 500 ms, about 250 ms to about 500 ms, or about 500 ms to about 1 s. In some embodiments, the average pulse duration of one characteristic pattern differs from that of the other characteristic pattern by approximately 10–25%, 25–50%, 50–75%, 75–100%, or more than 100%, for example, by approximately 2, 3, 4, 5, or more. In some embodiments, when the difference in average pulse duration between different characteristic patterns is smaller, a greater number of pulse durations may be required within each characteristic pattern to distinguish one from the other with statistical reliability.

[0214] In some embodiments, the characteristic pattern generally refers to multiple association events between the amino acids of the polypeptide and means for binding to the amino acids (e.g., an amino acid recognition molecule). In some embodiments, the characteristic pattern includes at least 10 association events (e.g., at least 25, at least 50, at least 75, at least 100, at least 250, at least 500, at least 1,000, or more association events). In some embodiments, the characteristic pattern includes about 10 to about 100 association events (e.g., about 10 to about 500 association events, about 10 to about 250 association events, about 10 to about 100 association events, or about 50 to about 500 association events). In some embodiments, multiple association events are detected as multiple signal pulses.

[0215] In some embodiments, the characteristic pattern means a plurality of signal pulses that can be characterized by key statistics as described herein. In some embodiments, the characteristic pattern includes at least 10 signal pulses (e.g., at least 25, at least 50, at least 75, at least 100, at least 250, at least 500, at least 1,000, or more signal pulses). In some embodiments, the characteristic pattern includes about 10 to about 100 signal pulses (e.g., about 10 to about 500 signal pulses, about 10 to about 250 signal pulses, about 10 to about 100 signal pulses, or about 50 to about 500 signal pulses).

[0216] In some embodiments, the characteristic pattern refers to multiple association events occurring between the amino acid recognition molecule and the amino acids of the polypeptide during the time interval prior to the removal of the amino acids. In some embodiments, the characteristic pattern refers to multiple association events occurring during the time interval between two cleavage events (e.g., before the removal of the amino acid and after the removal of the previously terminally exposed amino acid). In some embodiments, the time interval of the characteristic pattern is about 1 minute to about 30 minutes (e.g., about 1 minute to about 20 minutes, about 1 minute to about 10 minutes, about 5 minutes to about 20 minutes, about 5 minutes to about 15 minutes, or about 5 minutes to about 10 minutes).

[0217] In some embodiments, polypeptide sequencing reaction conditions can be configured to achieve time intervals that allow sufficient association events to provide a desired level of confidence with characteristic patterns. This can be achieved by configuring reaction conditions based on various properties, including, for example, reagent concentrations, molar ratios of one reagent to another (e.g., molar ratio of amino acid recognition molecules to cleavage reagents, ratio of one recognition molecule to another, ratio of one cleavage reagent to another), number of different reagent types (e.g., number of recognition molecules and / or cleavage reagents of different types from each other, number of recognition molecule types relative to number of cleavage reagent types), cleavage activity (e.g., peptidase activity), binding properties (e.g., kinetic and / or thermodynamic binding parameters of recognition molecule binding), reagent modifications (e.g., polyols and other protein modifications that can alter interaction dynamics), reaction mixture composition (e.g., one or more compositions such as pH, buffers, salts, divalent cations, surfactants, and other reaction mixture compositions described herein), reaction temperature, and various other parameters apparent to those skilled in the art, as well as combinations thereof. The reaction conditions may be based on one or more embodiments described herein, including, for example, signal pulse information (e.g., pulse duration, inter-pulse duration, magnitude change), labeling strategy (e.g., number and / or type of fluorophores, whether or not the linker has a shielding element), surface modification (e.g., modification of the sample well surface including polypeptide immobilization), sample preparation (e.g., polypeptide fragment size, polypeptide modification for immobilization), and other embodiments described herein.

[0218] In some embodiments, the polypeptide sequencing reaction according to the present invention is carried out under conditions in which amino acid recognition and cleavage can occur simultaneously in a single reaction mixture. For example, in some embodiments, the polypeptide sequencing reaction is carried out in a reaction mixture having a pH in which association and cleavage events can occur. In some embodiments, the polypeptide sequencing reaction is carried out in a reaction mixture with a pH of about 6.5 to about 9.0. In some embodiments, the polypeptide sequencing reaction is carried out in a reaction mixture with a pH of about 7.5 to about 8.0 (e.g., about 7.0 to about 8.0, about 7.5 to about 8.5, about 7.5 to about 8.0, or 8.0 to about 8.5).

[0219] In some embodiments, the polypeptide sequencing reaction is carried out in a reaction mixture containing one or more buffers. In some embodiments, the reaction mixture contains the buffer at a concentration of at least 10 mM (e.g., at least 20 mM and up to 250 mM, at least 50 mM, 10–250 mM, 10–100 mM, 20–100 mM, 50–100 mM, or 100–200 mM). In some embodiments, the reaction mixture contains the buffer at a concentration of about 10 mM to about 50 mM (e.g., about 10 mM to about 25 mM, about 25 mM to about 50 mM, or about 20 mM to about 40 mM). Examples of buffering agents, though not limited to them, include HEPES (4-(2-hydroxyethyl)-1-piperazineethanesulfonic acid), Tris (tris(hydroxymethyl)aminomethane), and MOPS (3-(N-monophorino)propanesulfonic acid).

[0220] In some embodiments, the polypeptide sequencing reaction is carried out in a reaction mixture containing a salt at a concentration of at least 10 mM. In some embodiments, the reaction mixture contains a salt at a concentration of at least 10 mM (e.g., at least 20 mM, at least 50 mM, at least 100 mM, or more). In some embodiments, the reaction mixture contains a salt at a concentration of about 10 mM to about 250 mM (e.g., about 20 mM to about 200 mM, about 50 mM to about 150 mM, about 10 mM to about 50 mM, or about 10 mM to about 100 mM). Examples of salts include, but are not limited to, sodium salts, potassium salts, and acetate salts such as sodium chloride (NaCl), sodium acetate (NaOAc), and potassium acetate (KOAc).

[0221] Additional examples of compositions for use in the reaction mixture include divalent cations (e.g., Mg 2+ Co 2+ Examples include ) and surfactants (polysorbate 20). In some embodiments, the reaction mixture contains a divalent cation at a concentration of about 0.1 mM to about 50 mM (e.g., about 10 mM to about 50 mM, about 0.1 mM to about 10 mM, or about 1 mM to about 20 mM). In some embodiments, the reaction mixture contains a surfactant at a concentration of at least 0.01% (e.g., about 0.01% to about 0.10%). In some embodiments, the reaction mixture contains one or more components useful for single-molecule analysis, such as oxygen exclusion systems (PCA / PCD systems or pyranose / catalase / glucose systems) and / or one or more triplet-state quenchers (e.g., Trolox, COT, and NBA).

[0222] In some embodiments, the polypeptide sequencing reaction is carried out at a temperature at which association and cleavage events can occur. In some embodiments, the polypeptide sequencing reaction is carried out at a temperature of at least 10°C. In some embodiments, the polypeptide sequencing reaction is carried out at a temperature of about 10°C to about 50°C (for example, 15 to 45°C, around 25°C, around 30°C, around 35°C, around 37°C). In some embodiments, the polypeptide sequencing reaction is carried out at or around room temperature.

[0223] As previously detailed, the real-time sequencing process illustrated in Figure 5 typically includes cycles of terminal amino acid recognition and terminal amino acid cleavage, where the relative occurrence of recognition and cleavage can be controlled by the concentration difference between the labeled recognition molecule 510 and the labeled nonspecific exopeptidase 520. In some embodiments, the concentration difference may be optimized so that the number of signal pulses detected during the recognition of individual amino acids provides a desirable confidence interval for identification. For example, if the initial sequencing reaction provides signal data with too few signal pulses between cleavage events to enable the determination of a characteristic pattern with a desirable confidence interval, the sequencing reaction may be repeated using a reduced concentration of nonspecific exopeptidase relative to the recognition molecule.

[0224] In some embodiments, polypeptide sequencing according to the present invention can be carried out by contacting a peptide with a sequencing reaction mixture comprising one or more amino acid recognition molecules and / or one or more cleavage reagents (e.g., peptidases). In some embodiments, the sequencing reaction mixture contains amino acid recognition molecules at a concentration of about 10 nM to about 10 μM. In some embodiments, the sequencing reaction mixture contains cleavage reagents at a concentration of about 500 nM to about 500 μM.

[0225] In some embodiments, the sequencing reaction mixture contains amino acid recognition molecules at concentrations of about 100 nM to about 10 μM, about 250 nM to about 10 μM, about 100 nM to about 1 μM, about 250 nM to about 1 μM, about 250 nM to about 750 nM, or about 500 nM to about 1 μM.

[0226] In some embodiments, the sequencing reaction mixture contains the cleavage reagent at concentrations of approximately 500 nM to approximately 250 nM, approximately 500 nM to approximately 100 nM, approximately 1 μM to approximately 100 μM, approximately 500 nM to approximately 50 μM, approximately 1 μM to approximately 100 μM, approximately 10 μM to approximately 200 μM, or approximately 10 μM to approximately 100 μM. In some embodiments, the sequencing reaction mixture contains the cleavage reagent at concentrations of approximately 1 μM, approximately 5 μM, approximately 10 μM, approximately 30 μM, approximately 50 μM, approximately 70 μM, or approximately 100 μM.

[0227] In some embodiments, the sequencing reaction mixture includes an amino acid recognition molecule at a concentration of about 10 nM to about 10 μM and a cleavage reagent at a concentration of about 500 nM to about 500 μM. In some embodiments, the sequencing reaction mixture includes an amino acid recognition molecule at a concentration of about 100 nM to about 1 μM and a cleavage reagent at a concentration of about 1 μM to about 100 μM. In some embodiments, the sequencing reaction mixture includes an amino acid recognition molecule at a concentration of about 250 nM to about 1 μM and a cleavage reagent at a concentration of about 10 μM to about 100 μM. In some embodiments, the sequencing reaction mixture includes an amino acid recognition molecule at a concentration of about 500 nM and a cleavage reagent at a concentration of about 25 μM to about 75 μM. In some embodiments, the concentrations of the amino acid recognition molecule and / or the cleavage reagent in the reaction mixture are as described elsewhere in this specification.

[0228] In some embodiments, the sequencing reaction mixture contains an amino acid recognition molecule and a cleavage reagent in a molar ratio of about 500:1, about 400:1, about 300:1, about 200:1, about 100:1, about 75:1, about 50:1, about 25:1, about 10:1, about 5:1, about 2:1, or about 1:1. In some embodiments, the sequencing reaction mixture contains an amino acid recognition molecule and a cleavage reagent in a molar ratio of about 10:1 to about 200:1. In some embodiments, the sequencing reaction mixture contains an amino acid recognition molecule and a cleavage reagent in a molar ratio of about 50:1 to about 150:1. In some embodiments, the molar ratio of amino acid recognition molecules to cleavage reagents in the reaction mixture is about 1:1,000 to about 1:1 or about 1:1 to about 100:1 (for example, about 1:1,000, about 1:500, about 1:200, about 1:100, about 1:10, about 1:5, about 1:2, about 1:1, about 5:1, about 10:1, about 50:1, about 100:1). In some embodiments, the molar ratio of amino acid recognition molecules to cleavage reagents in the reaction mixture is about 1:100 to about 1:1 or about 1:1 to about 10:1. In some embodiments, the molar ratio of amino acid recognition molecules to cleavage reagents in the reaction mixture is described elsewhere in this specification.

[0229] In some embodiments, the sequencing reaction mixture comprises one or more amino acid recognition molecules and one or more cleavage reagents. In some embodiments, the sequencing reaction mixture comprises at least three amino acid recognition molecules and at least one cleavage reagent. In some embodiments, the sequencing reaction mixture comprises two or more cleavage reagents. In some embodiments, the sequencing reaction mixture comprises at least one and up to 10 cleavage reagents (e.g., one to three cleavage reagents, two to ten cleavage reagents, one to five cleavage reagents, three to ten cleavage reagents). In some embodiments, the sequencing reaction mixture comprises at least three and up to 30 amino acid recognition molecules (e.g., 3 to 25, 3 to 20, 3 to 10, 3 to 5, 5 to 30, 5 to 20, 5 to 10, or 10 to 20 amino acid recognition molecules). In some embodiments, one or more amino acid recognition molecules comprise at least one recognition molecule selected from Table 1 or Table 2. In some embodiments, one or more cleavage reagents comprise at least one peptidase molecule selected from Table 4.

[0230] In some embodiments, the sequencing reaction mixture includes one or more amino acid recognition molecules and / or one or more cleavage reagents. In some embodiments, a sequencing reaction mixture described as including one or more amino acid recognition molecules (or cleavage reagents) means a mixture having one or more types of amino acid recognition molecules (or cleavage reagents). For example, in some embodiments, the sequencing reaction mixture includes two or more amino acid-binding proteins. In some embodiments, two or more amino acid-binding proteins mean two or more types of amino acid-binding proteins. In some embodiments, one type of amino acid-binding protein has a different amino acid sequence from another type of amino acid-binding protein in the reaction mixture. In some embodiments, one type of amino acid-binding protein has a different label from the label of another type of amino acid-binding protein in the reaction mixture. In some embodiments, one type of amino acid-binding protein associates (e.g., binds) with amino acids different from those associated with another type of amino acid-binding protein in the reaction mixture. In some embodiments, one type of amino acid-binding protein associates (e.g., binds) with a subset of amino acids different from the subset of amino acids associated with another type of amino acid-binding protein in the reaction mixture.

[0231] The example illustrated in Figure 5 relates to a sequencing process using a labeled cleavage reagent, but the sequencing process is not limited to this embodiment. As described elsewhere in this specification, the inventors have demonstrated single-molecule sequencing using unlabeled cleavage molecules. In some embodiments, the approximate frequency at which the cleavage reagent sequentially removes terminal amino acids is known, for example, based on the well-known activity and / or concentration of the enzyme used. In some embodiments, terminal amino acid cleavage by the reagent is inferred, for example, based on a signal detected for amino acid recognition or the absence of a detected signal.

[0232] The inventors are also aware of techniques for controlling real-time sequencing reactions that can be used in combination with or instead of the concentration difference approach described. An example of a temperature-dependent real-time sequencing process is shown in Figure 6. Panels (I) to (III) illustrate a sequencing reaction that includes temperature-dependent terminal amino acid recognition and amino acid cleavage cycles. Each cycle of the sequencing reaction was performed over two temperature ranges. The first temperature range ("T1") is optimized for recognition molecule activity rather than exopeptidase activity (e.g., to promote terminal amino acid recognition), and the second temperature range ("T2") is optimized for exopeptidase activity rather than recognition molecule activity (e.g., to promote terminal amino acid cleavage). The sequencing reaction proceeds by varying the temperature of the reaction mixture between the first temperature range T1 (to initiate amino acid recognition) and the second temperature range T2 (to initiate amino acid cleavage). Therefore, the progress of the temperature-dependent sequencing process is controlled by temperature, and changes between different temperature ranges (e.g., between T1 and T2) can be performed by manual or automated processes. In some embodiments, the recognition molecule activity (e.g., binding affinity to amino acids (K)) in the first temperature range T1 is controlled. D The exopeptidase activity in the second temperature range T2 increases by at least 10 times, at least 100 times, at least 1,000 times, at least 10,000 times, at least 100,000 times, or more, compared to the second temperature range T2. In some embodiments, the exopeptidase activity (e.g., the rate of substrate conversion to cleavage products) in the second temperature range T2 increases by at least 2 times, 10 times, at least 25 times, at least 50 times, at least 100 times, at least 1,000 times, or more, compared to the first temperature range T1.

[0233] In some embodiments, the first temperature range T1 is lower than the second temperature range T2. In some embodiments, the first temperature range T1 is about 15°C to about 40°C (e.g., about 25°C to about 35°C, about 15°C to about 30°C, about 20°C to about 30°C). In some embodiments, the second temperature range T2 is about 40°C to about 100°C (e.g., about 50°C to about 90°C, about 60°C to about 90°C, about 70°C to about 90°C). In some embodiments, the first temperature range T1 is about 20°C to about 40°C (e.g., about 30°C), and the second temperature range T2 is about 60°C to about 100°C (e.g., about 80°C).

[0234] In some embodiments, the first temperature range T1 is higher than the second temperature range T2. In some embodiments, the first temperature range T1 is about 40°C to about 100°C (e.g., about 50°C to about 90°C, about 60°C to about 90°C, about 70°C to about 90°C). In some embodiments, the second temperature range T2 is about 15°C to about 40°C (e.g., about 25°C to about 35°C, about 15°C to about 30°C, about 20°C to about 30°C). In some embodiments, the first temperature range T1 is about 60°C to about 100°C (e.g., about 80°C), and the second temperature range T2 is about 20°C to about 40°C (e.g., about 30°C).

[0235] Panel (I) depicts a sequencing reaction mixture at a temperature within a first temperature range T1 that is optimal for recognition molecule activity rather than exopeptidase activity. For illustrative purposes, a polypeptide with the amino acid sequence "KFVAG..." is shown. When the reaction mixture temperature is within the first temperature range T1, the labeled recognition molecule in the mixture is activated (e.g., restored) and initiates amino acid recognition by association with the polypeptide terminus. Also within the first temperature range T1, the labeled exopeptidase in the mixture is inactivated (e.g., denatured) and inhibits amino acid cleavage during recognition. In Panel (I), the first recognition molecule is shown to reversibly associate with the lysine at the polypeptide terminus, while the labeled exopeptidase (e.g., Pfu aminopeptidase I (Pfu API)) is shown in a denatured state. In some embodiments, amino acid recognition occurs for a predetermined duration before the start of amino acid cleavage. In some embodiments, amino acid recognition occurs for a duration necessary to reach a desired confidence interval for recognition before the start of amino acid cleavage. After amino acid recognition, the reaction proceeds by changing the temperature of the mixture within the second temperature range T2.

[0236] Panel (II) depicts the sequencing reaction mixture at a temperature within a second temperature range T2, which is optimal for exopeptidase activity rather than recognition molecule activity. For illustrative purposes in this example, it should be recognized that although the second temperature range T2 is higher than the first temperature range T1, reagent activity can be optimized to any desired temperature range. Therefore, the progression from Panel (I) to Panel (II) is carried out by raising the temperature of the reaction mixture using a suitable heat source. When the reaction mixture reaches a temperature within the second temperature range T2, the labeled exopeptidase in the mixture is activated (e.g., restored) and begins cleaving terminal amino acids by exopeptidase activity. Also at the second temperature range T2, the labeled recognition molecule in the mixture is inactivated (e.g., denatured) and amino acid recognition during cleavage is inhibited. In Panel (II), the labeled exopeptidase is shown cleaving terminal lysine residues, while the labeled recognition molecule is denatured. In some embodiments, amino acid cleavage occurs for a predetermined duration before the recognition of sequential amino acids at the polypeptide terminus begins. In some embodiments, amino acid cleavage occurs for a duration necessary to detect the cleavage before the recognition of sequential amino acids begins. After amino acid cleavage, the reaction proceeds by changing the temperature of the mixture within a first temperature range T1.

[0237] Panel (III) depicts the start of the next cycle in the sequencing reaction, where the reaction mixture temperature has been reduced to within the first temperature range T1. Therefore, in this example, the progression from Panel (II) to Panel (III) can be achieved by removing the reaction mixture from the heat source or by other means cooling the reaction mixture to within the first temperature range T1 (e.g., actively or passively). As shown, the labeled recognition molecule, including the second recognition molecule that reversibly associates with phenylalanine at the polypeptide terminus, is restored, while the labeled exopeptidase is shown to be denatured. The sequencing reaction is continued by further cycling amino acid recognition and cleavage in a temperature-dependent manner, as illustrated in this example.

[0238] Therefore, dynamic sequencing approaches may involve reaction cycles controlled at the level of protein activity or function of one or more proteins in the reaction mixture. The temperature-dependent polypeptide sequencing process depicted in Figure 6 and described above may exemplify a general approach to polypeptide sequencing with controllable cycles of condition-dependent recognition and cleavage. For example, in some embodiments, the approach provides a luminescence-dependent sequencing process using a luminescence-activating reagent. In some embodiments, the luminescence-dependent sequencing process includes cycles of luminescence-dependent amino acid recognition and cleavage. Each cycle of the sequencing reaction may be performed by exposing the sequencing reaction mixture to two different luminescence conditions. The first luminescence condition is optimized for recognition molecule activity rather than exoceptidase activity (e.g., to promote amino acid recognition), and the second luminescence condition is optimized for exoceptidase activity rather than recognition molecule activity (e.g., to promote amino acid cleavage). The sequencing reaction proceeds by alternating between exposing the reaction mixture to the first luminescence condition (to initiate amino acid recognition) and exposing the reaction mixture to the second luminescence condition (to initiate amino acid cleavage). In some embodiments, the two different emission conditions include a first wavelength and a second wavelength, although this is not limited to the embodiments described above.

[0239] In some embodiments, the present application provides a real-time polypeptide sequencing method by evaluating the binding interactions between one or more labeled recognition molecules and terminal and internal amino acids, as well as the binding interactions between labeled nonspecific exopeptidases and terminal amino acids. Figure 7 shows an example of a sequencing method in which the methods illustrating and illustrating the approaches of Figures 5 and 6 are modified by using labeled recognition amino acids that selectively bind to and dissociate one type of amino acid at both terminal and internal positions. As described in the previous approaches, selective binding results in a series of pulses at signal output 700. However, in this approach, the series of pulses occurs at a rate determined by the number of amino acid types in the entire polypeptide. Therefore, in some embodiments, the rate of pulses corresponding to the association event can be a diagnosis of the number of cognate amino acids currently present in the polypeptide.

[0240] As in previous approaches, the labeled nonspecific peptidase 720 is present at a relatively lower concentration than the labeled recognition molecule 710 to provide, for example, an optimal time window between cleavage events (Figure 7, inset panel). In addition, in certain embodiments, a uniquely identifiable luminescence label on the labeled nonspecific peptidase 720 indicates when a cleavage event occurs. As the polypeptide is repeatedly cleaved, the pulse rate corresponding to binding by the labeled recognition molecule 710 may decrease stepwise each time a terminal amino acid is cleaved by the labeled nonspecific peptidase 720. This concept is illustrated by plot 702, which generally plots the pulse rate as a function of time, with cleavage events occurring at times indicated by arrows. Thus, in some embodiments, amino acids can be identified based on the pulse pattern and / or pulse rate occurring within the pattern detected between cleavage events, thereby allowing the polypeptide to be sequenced using this approach.

[0241] In some embodiments, terminal polypeptide sequence information (for example, determined as described herein) may be combined with polypeptide sequence information obtained from one or more other sources. For example, terminal polypeptide sequence information may be combined with internal polypeptide sequence information. In some embodiments, internal polypeptide sequence information may be obtained using one or more amino acid recognition molecules that associate with internal amino acids, as described herein. Internal or other polypeptide sequence information may be obtained before or during a polypeptide degradation process. In some embodiments, sequence information obtained by these methods may be combined with polypeptide sequence information obtained using other techniques, such as sequence information obtained using one or more internal amino acid recognition molecules.

[0242] Preparation of sequencing samples Polypeptide samples can be modified before sequencing. In some embodiments, the N-terminal or C-terminal amino acids of the polypeptide can be modified. Figure 8 illustrates a non-limiting example of terminal modification for preparing terminally modified polypeptides from a protein sample. In step (1), the protein sample 800 is fragmented to produce polypeptide fragment 802. Polypeptides can be fragmented by cleaving (e.g., chemically) and / or digesting (e.g., enzymatically using a peptidase such as trypsin) the polypeptide of interest. Fragmentation can be performed before or after labeling. In some embodiments, fragmentation is performed after labeling of the entire protein. One or more amino acids can be labeled before or after cleavage to produce labeled polypeptides. In some embodiments, polypeptides are size-selected after chemical or enzymatic fragmentation. In some embodiments, smaller polypeptides (e.g., <2 kDa) are removed, and larger polypeptides are retained for sequence analysis. Size selection can be achieved using techniques such as gel filtration, SEC, dialysis, PAGE gel extraction, microfluidic tension flow, or any other preferred technique. In step (2), the N-terminus or C-terminus of polypeptide fragment 802 is modified to produce a terminally modified polypeptide 804. In some embodiments, the modification includes the addition of a fixed moiety. In some embodiments, the modification includes the addition of a coupling moiety.

[0243] Therefore, provided herein are methods for modifying the ends of proteins and polypeptides with portions that allow for fixation to a surface (for example, the surface of a sample well on a chip used for protein analysis). In some embodiments, such methods include modifying the ends of a labeled polypeptide to be analyzed according to the present application. In yet other embodiments, such methods include modifying the ends of an enzyme that degrades or translocates a protein or polypeptide substrate according to the present application.

[0244] In some embodiments, the carboxyl terminus of a protein or polypeptide is modified in a manner comprising: (i) blocking a free carboxylic acid group of the protein or polypeptide; (ii) denaturing the protein or polypeptide (e.g., by thermal and / or chemical means); (iii) blocking a free thiol group of the protein or polypeptide; (iv) digesting the protein or polypeptide to produce at least one polypeptide fragment containing a free C-terminal carboxylic acid group; and (v) conjugating a functional moiety to the free C-terminal carboxylic acid group (e.g., chemically). In some embodiments, the method further comprises dialysis of the sample containing the protein or polypeptide after (i) and before (ii).

[0245] In some embodiments, the carboxyl terminus of a protein or polypeptide is modified by (i) denaturing the protein or polypeptide (e.g., by thermal and / or chemical means), (ii) blocking a free thiol group of the protein or polypeptide, (iii) digesting the protein or polypeptide to produce at least one polypeptide fragment containing a free C-terminal carboxylic acid group, (iv) blocking the free C-terminal carboxylic acid group to produce at least one polypeptide fragment containing a blocked C-terminal carboxylic acid group, or (v) conjugating a functional moiety to the blocked C-terminal carboxylic acid group (e.g., enzymatically). In some embodiments, the method further comprises dialysis of the sample containing the protein or polypeptide after (iv) and before (v).

[0246] In some embodiments, blocking a free carboxylic acid group means chemically modifying the group in a way that alters its chemical reactivity compared to an unmodified carboxylic acid. Preferred carboxylic acid blocking methods are known in the art and should involve modifying the side-chain carboxylic acid group so that it is chemically distinct from the carboxy-terminal carboxylic acid group of the polypeptide being functionalized. In some embodiments, blocking a free carboxylic acid group includes esterifying or amidating the free carboxylic acid group of the polypeptide. In some embodiments, blocking a free carboxylic acid group includes, for example, methyl esterifying the free carboxylic acid group of the polypeptide by reacting the polypeptide with methanolic HCl. Additional examples of reagents and techniques useful for blocking free carboxylic acid groups include, but are not limited to, 4-sulfo-2,3,5,6-tetrafluorophenol (STP) and / or carbodiimides, such as N-(3-dimethylaminopropyl)-N'-ethylcarbodiimide hydrochloride (EDAC), uronium reagents, diazomethane, alcohols and acids for Fischer esterification, the use of N-hydroxylsuccinimide (NHS) to form NHS esters (potentially as an intermediate to subsequent esterification or amine formation), or any other method of modifying or blocking carboxylic acids via reaction with carbonyldiimidazole (CDI), or formation of mixed anhydrides, or potentially via the formation of either esters or amides.

[0247] In some embodiments, blocking a free thiol group means chemically modifying the group in a way that alters its chemical reactivity compared to an unmodified thiol. In some embodiments, blocking a free thiol group involves reducing and alkylating a free thiol group of a protein or polypeptide. In some embodiments, reduction and alkylation are carried out by contacting the polypeptide with dithiothreitol (DTT) and one or both of iodoacetamide and iodoacetic acid. Examples of additional and alternative cysteine ​​reducing reagents that may be used are well known and not limited to, but include 2-mercaptoethanol, tris(2-carboxyethyl)phosphine hydrochloride (TCEP), tributylphosphine, dithiobutylamine (DTBA), or any reagent capable of reducing thiol groups. Examples of additional and alternative cysteine ​​blocking (e.g., cysteine ​​alkylation) reagents that may be used are well known and not limited to, but include acrylamide, 4-vinylpyridine, N-ethylmaleimide (NEM), N-ε-maleimidocaproic acid (EMCA), or any reagent that modifies cysteine ​​to prevent disulfide bond formation.

[0248] In some embodiments, digestion includes enzymatic digestion. In some embodiments, digestion is carried out by contacting a protein or polypeptide with an endopeptidase (e.g., trypsin) under digestive conditions. In some embodiments, digestion includes chemical digestion. Examples of reagents suitable for chemical and enzymatic digestion, which are known in the art and are not limited to, include trypsin, chemotrypsin, Lys-C, Arg-C, Asp-N, Lys-N, BNPS-skatole, CNBr, caspase, formic acid, glutamyl endopeptidase, hydroxylamine, iodosobenzoic acid, neutrophil elastase, pepsin, proline-endopeptidase, proteinase K, staphylococcal peptidase I, thermolysin, and thrombin.

[0249] In some embodiments, the functional moiety comprises a biotin molecule. In some embodiments, the functional moiety comprises a reactive chemical moiety such as an alkynyl. In some embodiments, conjugation of the functional moiety comprises biotinylation of the carboxy-terminal carboxymethyl ester group by carboxypeptidase Y, as is known in the art.

[0250] In some embodiments, a solubilizing moiety is added to the polypeptide. Figure 9 illustrates, for example, a non-limiting example of a solubilizing moiety added to the terminal amino acids of a polypeptide using a process of conjugating a solubilizing linker to the polypeptide.

[0251] In some embodiments, a terminally modified polypeptide 910 containing a linker-conjugating moiety 912 is conjugated to a solubilizing linker 920 containing a polypeptide-conjugating moiety 922. In some embodiments, the solubilizing linker contains a solubilizing polymer, such as a biomolecule (e.g., shown as a speckled shape). In some embodiments, the resulting linker-conjugated polypeptide 930, containing a linkage 932 formed between 912 and 922, further contains a surface-conjugating moiety 934. Therefore, in some embodiments, the methods and compositions provided herein are useful for modifying the ends of a polypeptide with a moiety that increases its solubility. In some embodiments, the solubilizing moiety is useful for relatively insoluble small polypeptides resulting from fragmentation (e.g., enzymatic fragmentation, such as using trypsin). For example, in some embodiments, short polypeptides in a polypeptide pool can be solubilized by conjugating a polymer (e.g., a short oligo, sugar, or other charged polymer) to the polypeptide.

[0252] In some embodiments, one or more surfaces of the sample well (e.g., the sidewalls of the sample well) can be modified. Non-limiting examples of passivation and / or antifouling of the sample well sidewall are illustrated in Figure 10, which illustrates a schematic example of a sample well with a modified surface that can be used to promote single molecule fixation to the bottom surface. In some embodiments, 1040 is SiO2. In some embodiments, 1042 is a polypeptide conjugating moiety (e.g., TCO, tetrazine, N3, alkynes, aldehydes, NCO, NHS, thiols, alkenes, DBCO, BCN, TPP, biotin, or other preferred conjugating moieties). In some embodiments, 1050 is TiO2 or Al2O3. In some embodiments, 1052 is hydrophobic C 4~18 Molecules, polytetrafluoroethylene compounds (e.g., (CF2)) 4~12 ), polyols, for example, polyethylene glycol (for example, PEG 3~100 ), polypropylene glycol, polyoxyethylene glycol, or combinations or variations thereof, or zwitterions such as sulfobetaine. In some embodiments, 1060 is Si. In some embodiments, 1070 is Al. In some embodiments, 1080 is TiN.

[0253] Luminous marker As used herein, an luminescent label is a molecule that absorbs one or more photons and subsequently emits one or more photons after one or more durations. In some embodiments, this term is used synonymously with “label” or “luminescent molecule,” depending on the context. In certain embodiments described herein, an luminescent label may mean an luminescent label of a labeled recognition molecule, an luminescent label of a labeled peptidase (e.g., a labeled exopeptidase, a labeled nonspecific exopeptidase), an luminescent label of a labeled peptide, an luminescent label of a labeled cofactor, or other labeled compositions described herein. In some embodiments, an luminescent label as described herein means a labeled amino acid of a labeled polypeptide comprising one or more labeled amino acids.

[0254] In some embodiments, the luminescent label may include first and second chromophores. In some embodiments, the excited state of the first chromophore can be relaxed via energy transfer to the second chromophore. In some embodiments, the energy transfer is Förster resonance energy transfer (FRET). Such a FRET pair may be useful in providing a luminescent label having properties that facilitate the distinction between individual labels among multiple luminescent labels in a mixture, for example, as illustrated and described herein for the labeled aptamer 206 in Figure 2. In yet another embodiment, the FRET pair includes a first chromophore of the first luminescent label and a second chromophore of the second luminescent label. In certain embodiments, the FRET pair may absorb excitation energy in a first spectral region and emit emission in a second spectral region.

[0255] In some embodiments, the luminescent label means a fluorophore or dye. Typically, the luminescent label includes aromatic or heteroaromatic compounds and may be pyrene, anthracene, naphthalene, naphthylamine, acridine, stilbene, indole, benzoindole, oxazole, carbazole, thiazole, benzothiazole, benzoxazole, phenanthidine, phenoxazine, porphyrin, quinoline, ethidium, benzamide, cyanine, carbocyanine, salicylate, anthranylate, coumarin, fluoroceine, rhodamine, xanthene, or other similar compounds.

[0256] In some embodiments, the luminescent label is as follows: 5 / 6-carboxyrhodamine 6G, 5-carboxyrhodamine 6G, 6-carboxyrhodamine 6G, 6-TAMRA, Abelia® STAR 440SXP, Abelia® STAR 470SXP, Abelia® STAR 488, Abelia® STAR ST AR)512, Abbelia (registered trademark) STAR 520SXP, Abbelia (registered trademark) STAR 580, Abbelia (registered trademark) STAR 600, Abbelia (registered trademark) STAR 635, Abbelia (registered trademark) STAR 635P, Abbelia (registered trademark) STAR Red (STAR RED), Alexafluor (Alexa Fluor) (registered trademark) 350, Alexafluor (Alexa Fluor) (registered trademark) 405, Alexafluor (Alexa Fluor) (registered trademark) 430, Alexafluor (Alexa Fluor) (registered trademark) 480, Alexafluor (Alexa Fluor) (registered trademark) 488, Alexafluor (Alexa Fluor) (registered trademark) 514, Alexafluor (Alexa Fluor) (registered trademark) 532, Alexafluor (Alexa Fluor) (registered trademark) 546, Alexafluor (Alexa Fluor) (registered trademark) 555, Alexafluor (Alexa Fluor) (registered trademark) 568, Alexafluor (Alexa Fluor) (registered trademark) 594, Alexafluor (Alexa Fluor) (registered trademark) 610×, Alexafluor (Alexa Fluor) (registered trademark) 633, Alexafluor (Alexa Alexafluor (registered trademark) 647, Alexafluor (registered trademark) 660, Alexafluor (registered trademark) 680, Alexafluor (registered trademark) 700, Alexafluor (registered trademark) 750,Alexa Fluor (registered trademark) 790, AMCA, ATTO 390, ATTO 425, ATTO 465, ATTO 488, ATTO 495, ATTO 514, ATTO 520, ATTO 532, ATTO 542, ATTO 550, ATTO 565, ATTO 590, ATTO 610, ATTO 620, ATTO 633, ATTO 647, ATTO 647N, A ATTO 655, ATTO 665, ATTO 680, ATTO 700, ATTO 725, ATTO 740, ATTO Oxa12, ATTO Rho101, ATTO Rho11, ATTO Rho12, ATTO Rho13, ATTO Rho14, ATTO Rho3B, ATTO Rho6G, ATTO Thio12, BD Horizon (trademark) V450, Body Peel ( BODIPY) (Registered Trademark) 493 / 501, BODIPY (Registered Trademark) 530 / 550, BODIPY (Registered Trademark) 558 / 568, BODIPY (Registered Trademark) 564 / 570, BODIPY (Registered Trademark) 576 / 589, BODIPY (Registered Trademark) 581 / 591, BODIPY (Registered Trademark) 630 / 650, BODIPY (Registered Trademark) 650 / 665, BODIPY (Registered Trademark) FL, BODIPY (Registered Trademark) )(Registered Trademark) FL-X, BODIPY(Registered Trademark) R6G, BODIPY(Registered Trademark) TMR, BODIPY(Registered Trademark) TR, CAL Fluor(Registered Trademark) Gold 540, CAL Fluor(Registered Trademark) Green 510, CAL Fluor(Registered Trademark) Orange 560, CAL Fluor(Registered Trademark) Red 590, CAL Fluor(Registered Trademark) Red 610, CAL Fluor(Registered Trademark) Red 615,CAL Fluor (registered trademark) Red 635, Cascade (registered trademark) Blue, CF (trademark) 350, CF (trademark) 405M, CF (trademark) 405S, CF (trademark) 488A, CF (trademark) 514, CF (trademark) 532, CF (trademark) 543, CF (trademark) 546, CF (trademark) 555, CF (trademark) 568, CF (trademark) 594, CF (trademark) 620R, CF (trademark) 633, CF (trademark) 633-V1, CF (trademark) 640R, CF (trademark) 640R-V1, CF (trademark) 640R-V2, CF (trademark) 660C, CF(trademark)660R, CF(trademark)680, CF(trademark)680R, CF(trademark)680R-V1, CF(trademark)750, CF(trademark)770, CF(trademark)790, Chromeo(trademark)642, Chromis425N, Chromis500N, Chromis515N, Chromis530N, Chromis550A, Chromis550C, Chromis550Z, Chromis560N, Chromis(Chr Chromis 570N, Chromis 577N, Chromis 600N, Chromis 630N, Chromis 645A, Chromis 645C, Chromis 645Z, Chromis 678A, Chromis 678C, Chromis 678Z, Chromis 770A, Chromis 770C, Chromis 800A, Chromis 800C, Chromis (Chromis)830A, Chromis 830C, Cy(registered trademark)3, Cy(registered trademark)3.5, Cy(registered trademark)3B, Cy(registered trademark)5, Cy(registered trademark)5.5, Cy(registered trademark)7, DyLight(registered trademark)350, DyLight(registered trademark)405, DyLight(registered trademark)415-Co1, DyLight(registered trademark)425Q, DyLight(registered trademark)485-LS, DyLight(registered trademark)488,DyLight (registered trademark) 504Q, DyLight (registered trademark) 510-LS, DyLight (registered trademark) 515-LS, DyLight (registered trademark) 521-LS, DyLight (registered trademark) 530-R2, DyLight (registered trademark) 543Q, DyLight (registered trademark) 550, DyLight (registered trademark) 554-R0, DyLight (registered trademark) 554-R1, DyLight (Registered Trademark) 590-R2, DyLight (Registered Trademark) 594, DyLight (Registered Trademark) 610-B1, DyLight (Registered Trademark) 615-B2, DyLight (Registered Trademark) 633, DyLight (Registered Trademark) 633-B1, DyLight (Registered Trademark) 633-B2, DyLight (Registered Trademark) 650, DyLight (Registered Trademark) 655-B1, DyLight (Registered Trademark) 655-B2, DyLight DyLight (registered trademark) 655-B3, DyLight (registered trademark) 655-B4, DyLight (registered trademark) 662Q, DyLight (registered trademark) 675-B1, DyLight (registered trademark) 675-B2, DyLight (registered trademark) 675-B3, DyLight (registered trademark) 675-B4, DyLight (registered trademark) 679-C5, DyLight (registered trademark) 680, DyLight ( (Registered Trademark) 683Q, DyLight (Registered Trademark) 690-B1, DyLight (Registered Trademark) 690-B2, DyLight (Registered Trademark) 696Q, DyLight (Registered Trademark) 700-B1, DyLight (Registered Trademark) 700-B1, DyLight (Registered Trademark) 730-B1, DyLight (Registered Trademark) 730-B2, DyLight (Registered Trademark) 730-B3, DyLight (Registered Trademark) 730-B4,DyLight (registered trademark) 747, DyLight (registered trademark) 747-B1, DyLight (registered trademark) 747-B2, DyLight (registered trademark) 747-B3, DyLight (registered trademark) 747-B4, DyLight (registered trademark) 755, DyLight (registered trademark) 766Q, DyLight (registered trademark) 775-B2, DyLight (registered trademark) 775-B3, DyLight (DyLight) ht) (Registered Trademark) 775-B4, DyLight (Registered Trademark) 780-B1, DyLight (Registered Trademark) 780-B2, DyLight (Registered Trademark) 780-B3, DyLight (Registered Trademark) 800, DyLight (Registered Trademark) 830-B2, Dyomics-350, Dyomics-350XL, Dyomics-360XL, Dyomics-370XL, Dyomics (Dyomi cs)-375XL, Diomics-380XL, Diomics-390XL, Diomics-405, Diomics-415, Diomics-430, Diomics-431, Diomics-478, Diomics-480XL, Diomics-481XL, Diomics-485XL, Diomics-490, Diomics (Dyomics)-495, Diomics (Dyomics)-505, Diomics (Dyomics)-510XL, Diomics (Dyomics)-511XL, Diomics (Dyomics)-520XL, Diomics (Dyomics)-521XL, Diomics (Dyomics)-530, Diomics (Dyomics)-547, Diomics (Dyomics)-547P1, Diomics (Dyomics)-548, Diomics (Dyomics)-549, Diomics (Dyomics)-549P1,Dyomics-550, Dyomics-554, Dyomics-555, Dyomics-556, Dyomics-560, Dyomics-590, Dyomics-591, Dyomics-594, Dyomics-Dy, Dyomics)-601XL, Dyomics-605, Dyomics-610, Dyomics-615, Dyomics-630, Dyomics-631, Dyomics-632, Dyomics-633, Dyomics-634, Dyomics-635, Dyomics-636, Dyomics-647, Dyomics (Dyomi cs)-647P1, Dyomics-648, Dyomics-648P1, Dyomics-649, Dyomics-649P1, Dyomics-650, Dyomics-651, Dyomics-652, Dyomics-654, Dyomics-675, Dyomics-676, Dyomics-677, Dyomics-677 cs)-678, Dyomics-679P1, Dyomics-680, Dyomics-681, Dyomics-682, Dyomics-700, Dyomics-701, Dyomics-703, Dyomics-704, Dyomics-730, Dyomics-731, Dyomics-732, Dyomics- 734, Dyomics-749, Dyomics-749P1, Dyomics-750, Dyomics-751, Dyomics-752, Dyomics-754, Dyomics-776, Dyomics-777, Dyomics-778, Dyomics-780, Dyomics-781, Dyomics-782,Dyomics-800, Dyomics-831, eFluor (registered trademark) 450, Eosin, FITC, Fluorescein, HiLyte (trademark) Fluor 405, HiLyte (trademark) Fluor 488, HiLyte (trademark) Fluor 532, HiLyte (trademark) Fluor 555, HiLyte (trademark) Fluor 594, HiLyte (trademark) Fluor 647, HiLyte (trademark) Fluor 680, Ha HiLyte (trademark) Fluor 750, IRDye (registered trademark) 680LT, IRDye (registered trademark) 750, IRDye (registered trademark) 800CW, JOE, LightCycler (registered trademark) 640R, LightCycler (registered trademark) Red 610, LightCycler (registered trademark) Red 640, LightCycler (registered trademark) Red 670, LightCycler (registered trademark) Red 705, Lisamin Rhodamine B, Naphthofluorescein, Oregon Green Green) (Registered Trademark) 488, Oregon Green (Registered Trademark) 514, Pacific Blue (Trademark), Pacific Green (Trademark), Pacific Orange (Trademark), PET, PF350, PF405, PF415, PF488, PF505, PF532, PF546, PF555P, PF568, PF594, PF610, PF633P, PF647P, Quasar (Registered Trademark) 570, Quasar (Registered Trademark) 670, Quasar (Registered Trademark) 705, Rhodamine 123, Rhodamine 6G, Rhodamine B, Rhodamine Green, Rhodamine Green-X, Rhodamine Red, ROX, Seta (Trademark) 375, Seta (Trademark) 470,Seta (trademark) 555, Seta (trademark) 632, Seta (trademark) 633, Seta (trademark) 650, Seta (trademark) 660, Seta (trademark) 670, Seta (trademark) 680, Seta (trademark) 700, Seta (trademark) 750, Seta (trademark) 780, Seta (trademark) APC-780, Seta (trademark) PerCP-680, Seta ( Contains pigments selected from one or more of the following: Seta (trademark) R-PE-670, Seta (trademark) 646, SeTau 380, SeTau 425, SeTau 647, SeTau 405, Square 635, Square 650, Square 660, Square 672, Square 680, Sulforhodamine 101, TAMRA, TET, Texas Red (registered trademark), TMR, TRITC, Yakima Yellow (trademark), Zenon (registered trademark), Zy3, Zy5, Zy5.5, and Zy7.

[0257] Luminous In some embodiments, the present application relates to polypeptide sequencing and / or identification based on one or more luminescence properties of luminescent labels. In some embodiments, luminescent labels are identified based on luminescence lifetime, luminescence intensity, brightness, absorption spectrum, emission spectrum, emission quantum yield, or two or more combinations thereof. In some embodiments, multiple types of luminescent labels are distinguishable from one another based on different luminescence lifetimes, luminescence intensity, brightness, absorption spectra, emission spectra, emission quantum yield, or two or more combinations thereof. Identification may mean assigning a precise identity and / or quantity of one type of amino acid (e.g., a single type or a subset of types) associated with the luminescent label, as well as assigning an amino acid position in the polypeptide compared to other types of amino acids.

[0258] In some embodiments, luminescence is detected by exposing a luminescent label to a series of individual light pulses and evaluating the timing or other properties of each photon emitted from the label. In some embodiments, information from multiple photons sequentially emitted from the label is aggregated and evaluated to identify the type of associated amino acid by identifying the label. In some embodiments, the luminescence lifetime of the label is determined from multiple photons sequentially emitted from the label, and the luminescence lifetime is available for identification of the label. In some embodiments, the luminescence intensity of the label is determined from multiple photons sequentially emitted from the label, and the luminescence intensity is available for identification of the label. In some embodiments, the luminescence lifetime and luminescence intensity of the label are determined from multiple photons sequentially emitted from the label, and the luminescence lifetime and luminescence intensity are available for identification of the label.

[0259] In some embodiments of the present application, a single polypeptide molecule is exposed to a plurality of individual light pulses, and a series of emitted photons are detected and analyzed. In some embodiments, the series of emitted photons provide information about a single polypeptide molecule present in the reaction sample that does not change over time in the experiment. However, in some embodiments, the series of emitted photons provide information about a series of different molecules present in the reaction sample at different times (e.g., as the reaction or process progresses). Such information may be used, but is not limited to, to sequence and / or identify polypeptides subjected to chemical or enzymatic degradation in accordance with the present application.

[0260] In certain embodiments, a light-emitting label absorbs one photon and emits one photon after a certain duration. In some embodiments, the light-emitting lifetime of a label can be determined or estimated by measuring the duration. In some embodiments, the light-emitting lifetime of a label can be determined or estimated by measuring multiple durations for multiple pulse and emission events. In some embodiments, the light-emitting lifetime of a label can be distinguished among the light-emitting lifetimes of multiple types of labels by measuring the duration. In some embodiments, the light-emitting lifetime of a label can be distinguished among the light-emitting lifetimes of multiple types of labels by measuring multiple durations for multiple pulse and emission events. In certain embodiments, a label is identified or distinguished among multiple types of labels by determining or estimating its light-emitting lifetime. In certain embodiments, a label is identified or distinguished among multiple types of labels by distinguishing its light-emitting lifetime among multiple light-emitting lifetimes of multiple types of labels.

[0261] The luminescence lifetime of a luminescent label can be determined using any preferred method (for example, by measuring the lifetime using a preferred technique or by determining the time-dependent characteristics of the emission). In some embodiments, determining the luminescence lifetime of one label includes determining the lifetime in comparison to other labels. In some embodiments, determining the luminescence lifetime of a label includes determining the lifetime in comparison to a reference. In some embodiments, determining the luminescence lifetime of a label includes measuring the lifetime (e.g., fluorescence lifetime). In some embodiments, determining the luminescence lifetime of a label includes determining one or more time characteristics that serve as indicators of lifetime. In some embodiments, the luminescence lifetime of a label can be determined based on the distribution of multiple emission events (e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 30, 40, 50, 60, 70, 80, 90, 100, or more emission events) occurring across one or more time gate windows for an excitation pulse. For example, the luminescence lifetime of a label can be distinguished from multiple labels having different luminescence lifetimes based on the distribution of photon arrival times measured for an excitation pulse.

[0262] It should be recognized that the luminescence lifetime of a luminescent label serves as an indicator of the timing of photons emitted after the label reaches an excited state, and that labels can be distinguished by information that indicates the timing of the photons. Some embodiments may include distinguishing a label from multiple labels based on its luminescence lifetime by measuring the time associated with the photons emitted by the label. The time distribution can provide an indicator of the luminescence lifetime that can be determined from the distribution. In some embodiments, a label can be distinguished from multiple labels based on the time distribution, for example, by comparing the time distribution with a reference distribution corresponding to a known label. In some embodiments, the value of the luminescence lifetime was determined from the time distribution.

[0263] As used herein, in some embodiments, luminescence intensity means the number of emitted photons per unit time emitted by a luminescent label excited by the delivery of pulsed excitation energy. In some embodiments, luminescence intensity means the number of detected emitted photons per unit time emitted by a label excited by the delivery of pulsed excitation energy and detected by a particular sensor or set of sensors.

[0264] As used herein, in some embodiments, luminance refers to a parameter reporting the average luminescence intensity per luminescent label. Thus, in some embodiments, “luminescence intensity” may be used to generally mean the luminance of a composition comprising one or more labels. In some embodiments, the luminance of a label is equal to the product of its quantum yield and its extinction coefficient.

[0265] As used herein, in some embodiments, emission quantum yield means the proportion of excitation events at a given wavelength or within a given spectral region that result in emission events, and is typically less than 1. In some embodiments, the emission quantum yield of the emission labels described herein is 0 to about 0.001, about 0.001 to about 0.01, about 0.01 to about 0.1, about 0.1 to about 0.5, about 0.5 to 0.9, or about 0.9 to 1. In some embodiments, the label is identified by identifying or estimating the emission quantum yield.

[0266] In some embodiments as used herein, the excitation energy is a light pulse from a light source. In some embodiments, the excitation energy is in the visible spectral region. In some embodiments, the excitation energy is in the ultraviolet spectral region. In some embodiments, the excitation energy is in the infrared spectral region. In some embodiments, the excitation energy is at or near the absorption maximum of an emission label where multiple emitted photons are detected. In certain embodiments, the excitation energy is in the range of about 500 nm to about 700 nm (e.g., about 500 nm to about 600 nm, about 600 nm to about 700 nm, about 500 nm to about 550 nm, about 550 nm to about 600 nm, about 600 nm to about 650 nm, or about 650 nm to about 700 nm). In certain embodiments, the excitation energy may be monochromatic or confined to a certain spectral region. In some embodiments, the spectral region has a range of about 0.1 nm to about 1 nm, about 1 nm to about 2 nm, or about 2 nm to about 5 nm. In some embodiments, the spectral region has a range of approximately 5 nm to approximately 10 nm, approximately 10 nm to approximately 50 nm, or approximately 50 nm to approximately 100 nm.

[0267] Sequencing Aspects of this application relate to the sequencing of biopolymers such as polypeptides and proteins. As used herein, terms such as “sequencing,” “sequencing,” and “determining a sequence” include, with respect to polypeptides or proteins, the determination of partial sequence information and even whole sequence information of a polypeptide or protein. That is, the terms include sequence comparison, fingerprinting, stochastic fingerprinting, and equivalent information regarding a target molecule, as well as the clear identification and ordering of each amino acid of a target molecule within a region of interest. In some embodiments, the terms include the identification of a single amino acid of a polypeptide. In yet another embodiment, more than one amino acid of a polypeptide is identified. As used herein, in some embodiments, terms such as “identifying” and “determining an identity” include, with respect to amino acids, the determination of the clear identity of an amino acid, as well as the determination of the probability of the clear identity of an amino acid. For example, in some embodiments, an amino acid is identified by determining the probability (e.g., 0% to 100%) that the amino acid is of a particular type, or by determining the probability of each of several particular types. Therefore, in some embodiments, the terms “amino acid sequence,” “polypeptide sequence,” and “protein sequence” as used herein may mean the polypeptide or protein substance itself, and are not limited to specific sequence information (for example, a series of letters representing the order of amino acids from one end to the other) that biochemically characterizes a particular polypeptide or protein.

[0268] In some embodiments, sequencing of a polypeptide molecule involves identifying at least two amino acids in the polypeptide molecule (for example, at least three, at least four, at least five, at least six, at least seven, at least eight, at least nine, at least ten, at least 11, at least 12, at least 13, at least 14, at least 15, at least 16, at least 17, at least 18, at least 19, at least 20, at least 25, at least 30, at least 35, at least 40, at least 45, at least 50, at least 60, at least 70, at least 80, at least 90, at least 100, or more). In some embodiments, at least two amino acids are consecutive amino acids. In some embodiments, at least two amino acids are discontinuous amino acids.

[0269] In some embodiments, sequencing of a polypeptide molecule involves identifying less than 100% of all amino acids in the polypeptide molecule (e.g., less than 99%, less than 95%, less than 90%, less than 85%, less than 80%, less than 75%, less than 70%, less than 65%, less than 60%, less than 55%, less than 50%, less than 45%, less than 40%, less than 35%, less than 30%, less than 25%, less than 20%, less than 15%, less than 10%, less than 5%, less than 1%, or less). For example, in some embodiments, sequencing of a polypeptide molecule involves identifying less than 100% of one type of amino acid in the polypeptide molecule (e.g., identifying a portion of all amino acids of one type in the polypeptide molecule). In some embodiments, sequencing of a polypeptide molecule involves identifying less than 100% of each type of amino acid in the polypeptide molecule.

[0270] In some embodiments, sequencing of a polypeptide molecule involves identifying at least one, at least five, at least 10, at least 15, at least 20, at least 25, at least 30, at least 35, at least 40, at least 45, at least 50, at least 55, at least 60, at least 65, at least 70, at least 75, at least 80, at least 85, at least 90, at least 95, at least 100, or more types of amino acids in the polypeptide.

[0271] In some embodiments, the present application provides compositions and methods for sequencing a polypeptide by identifying a set of amino acids present at the terminals of the polypeptide over time (for example, by repeated detection and cleavage of terminal amino acids). In yet other embodiments, the present application provides compositions and methods for sequencing a polypeptide by identifying the labeled amino acid content of the polypeptide and comparing it with a reference sequence database.

[0272] In some embodiments, the present application provides compositions and methods for sequencing a polypeptide by sequencing multiple fragments of the polypeptide. In some embodiments, sequencing a polypeptide involves the step of combining sequence information of multiple polypeptide fragments to identify and / or determine the sequence of the polypeptide. In some embodiments, combining sequence information may be carried out by computer hardware and software. The methods described herein may enable sequencing of a set of relevant polypeptides, e.g., the entire proteome of an organism. In some embodiments, multiple single-molecule sequencing reactions are carried out in parallel according to embodiments of this application (e.g., on a single chip). For example, in some embodiments, multiple single-molecule sequencing reactions are each carried out in separate sample wells on a single chip.

[0273] In some embodiments, the methods provided herein may be used for sequencing and identifying individual proteins in a sample containing a complex mixture of proteins. In some embodiments, the present application provides a method for uniquely identifying individual proteins in a complex mixture of proteins. In some embodiments, individual proteins are detected in a mixed sample by determining the partial amino acid sequence of the protein. In some embodiments, the partial amino acid sequence of the protein is within a continuous stretch of approximately 5 to 50 amino acids.

[0274] While we do not wish to be bound by any particular theory, it is believed that most human proteins can be identified using incomplete sequence information by referring to proteomic databases. For example, simple modeling of the human proteome has shown that approximately 98% of proteins can be uniquely identified by detecting just four types of amino acids within a 6-40 amino acid stretch (see, for example, Swaminathan et al., PLoS Comput Biol., 2015, Vol. 11, No. 2, p.e1004080, and Yao et al., 2015, Vol. 12, No. 5, p.055003). Therefore, the protein complex mixture can be broken down (e.g., chemically or enzymatically) into short polypeptide fragments of approximately 6–40 amino acids, and sequencing of this polypeptide library will reveal the identity and abundance of each protein present in the original complex mixture. Compositions and methods for identifying polypeptides by selective amino acid labeling and determination of partial sequence information are described in detail in U.S. Patent Application No. 15 / 510,962, filed September 15, 2015, entitled "Single Molecular Peptide Sequencing" (which is incorporated in its entirety by reference).

[0275] The embodiments enable sequencing of single polypeptide molecules with high accuracy, such as at least about 50%, 60%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, 99.9%, 99.99%, 99.999%, or 99.9999%. In some embodiments, the target molecule used for single-molecule sequencing is a polypeptide immobilized on a surface of a solid support, such as the bottom surface or sidewall surface of a sample well. The sample well may also contain any other reagents required for the sequencing reaction according to the present invention, such as one or more suitable buffers, cofactors, labeled recognition molecules, and enzymes (e.g., catalytically active or inactive exopeptidase enzymes that may be luminescently labeled or unlabeled).

[0276] As described above, in some embodiments, the sequencing according to the present invention includes identifying amino acids by determining the probability that an amino acid is of a particular type. Conventional protein identification systems require the identification of each amino acid in a polypeptide in order to identify the polypeptide. However, accurately identifying each amino acid in a polypeptide is difficult. For example, data collected from interactions in which a first recognition molecule associates with a first amino acid may not differ to a sufficient degree to distinguish the two amino acids from data collected from interactions in which a second recognition molecule associates with a second amino acid. In some embodiments, the sequencing according to the present invention avoids this problem by using a protein identification system that, unlike conventional protein identification systems, does not require (but does not exclude) the identification of each amino acid in a protein.

[0277] Therefore, in some embodiments, the sequencing according to the present invention may be performed using a protein identification system that employs machine learning techniques to identify proteins. In some embodiments, the system operates by (1) collecting data on the polypeptide of a protein using a real-time protein sequencing device, (2) using a machine learning model and the collected data to identify the probability that a particular amino acid is part of the polypeptide at its respective position, and (3) using the identification probability as a “probabilistic fingerprint”. In some embodiments, the data on the polypeptide of a protein may be obtained using a reagent that selectively binds to amino acids. For example, the reagent and / or amino acids may be labeled with an luminescent label that emits light in response to the application of excitation energy. In this example, the protein sequencing device may apply excitation energy to a sample of protein (e.g., polypeptide) at the time of the binding interaction between the reagent and the amino acids in the sample. In some embodiments, one or more sensors of the sequencing device (e.g., photodetectors, electrical sensors, and / or any other preferred type of sensor) may detect the binding interaction. Meanwhile, data collected and / or obtained from the detection light emission may be provided to a machine learning model. Machine learning models, related systems, and methods are described in detail in U.S. Provisional Patent Application No. 62 / 860,750, filed June 12, 2019, entitled "Machine Learning Enabled Protein Identification" (which is incorporated in its entirety by reference).

[0278] In some embodiments, sequencing according to the present application may involve immobilizing polypeptides onto the surface of a substrate (e.g., a solid carrier, e.g., a chip, e.g., an integrated device as described herein). In some embodiments, polypeptides may be immobilized on the surface of a sample well on the substrate (e.g., on the bottom surface of a sample well). In some embodiments, the N-terminal amino acid of the polypeptide is immobilized (e.g., attached to the surface). In some embodiments, the C-terminal amino acid of the polypeptide is immobilized (e.g., attached to the surface). In some embodiments, one or more non-terminal amino acids are immobilized (e.g., attached to the surface). The immobilized amino acids can be attached using any preferred covalent or non-covalent linkage, for example, as described herein. In some embodiments, multiple polypeptides are attached to multiple sample wells, for example, in an array of sample wells on a substrate (e.g., one polypeptide is attached to the surface of each sample well, e.g., the bottom surface).

[0279] The sequencing according to the present invention may, in some embodiments, be carried out using a system that enables single-molecule analysis. The system may include an integrate device and an instrument configured to interface with the integrate device. The integrate device may include an array of pixels, and each individual pixel may include a sample well and at least one photodetector. The sample wells of the integrate device may be formed on or over the entire surface of the integrate device and may be configured to receive a sample placed on the surface of the integrate device. As a whole, the sample wells may be considered an array of sample wells. Multiple sample wells may have a size and shape suitable for at least a portion of the sample wells to receive a single sample (e.g., a single molecule such as a polypeptide). In some embodiments, the number of samples in the sample wells may be distributed within the sample wells of the integrate device such that some sample wells contain one sample, while others contain zero, two, or more samples.

[0280] Excitation light is supplied to the integrate device from one or more light sources outside the integrate device. The optical components of the integrate device can receive the excitation light from the light source, direct the light to an array of sample wells in the integrate device, and illuminate the illumination area within the sample wells. In some embodiments, the sample wells may have a configuration that allows the sample to be held close to the surface of the sample well and facilitates the delivery of excitation light to the sample and the detection of emitted light from the sample. A sample positioned within the illumination area may emit emitted light in response to illumination by the excitation light. For example, a sample may be labeled with a fluorescent marker that emits light in response to achieving an excited state via illumination by the excitation light. The emitted light emitted by the sample can then be detected by one or more photodetectors in pixels corresponding to the sample wells containing the sample to be analyzed. According to some embodiments, multiple samples can be analyzed in parallel when carried out across an entire array of sample wells, the number of which may range from approximately 10,000 to 1,000,000 pixels.

[0281] The integrated device may include an optical system for receiving excitation light and directing it between sample well arrays. The optical system may include one or more grating couplers configured to couple the excitation light to the integrated device and direct it to other optical components. The optical system may include optical components that direct the excitation light from the grating couplers to the sample well arrays. Such optical components may include optical splitters, optical combiners, and waveguides. In some embodiments, one or more optical splitters may couple the excitation light from the grating couplers and deliver it to at least one of the waveguides. According to some embodiments, the optical splitter may have a configuration that allows for substantially uniform delivery of excitation light across all waveguides so that each waveguide receives substantially similar amounts of excitation light. Such embodiments can improve the performance of the integrated device by improving the uniformity of the excitation light received by the sample wells of the integrated device. For example, suitable components to be included in an integrated device for coupling excitation light to a sample well and / or directing emitted light to a photodetector are described in U.S. Patent Application No. 14 / 821,688, filed August 7, 2015, entitled "Integrated Device for Probing, Detecting and Analyzing Molecules," and "Integrated Device with External Light Source for Probing, Detecting and Analyzing Molecules." The following is described in U.S. Patent Application No. 14 / 543,865, filed November 17, 2014, entitled “Probing, Detecting, and Analyzing Molecules” (both incorporated in their entirety by reference). Suitable examples of grating couplers and waveguides that may be implemented in an integrated device are “Optical Coupler and Waveguide Systems” It is described in U.S. Patent Application No. 15 / 844,403, filed on December 15, 2017, titled "WAVEGUIDE SYSTEM" (which is incorporated in its entirety by reference).

[0282] Additional photonic structures may be positioned between the sample well and the photodetector and may be configured to reduce or prevent excitation light reaching the photodetector that may contribute to signal noise when the emitted light is detected. In some embodiments, a metal layer that may act as a circuit in the integrated device may also act as a spatial filter. Examples of suitable photonic structures may include spectral filters, polarizing filters and spatial filters, and are described in U.S. Patent Application No. 16 / 042,968, filed July 23, 2018, titled "Optical Rejection Photonic Structures" (which is incorporated in its entirety by reference).

[0283] Components positioned away from the integrate device may be used to position and align the excitation source to the integrate device. Such components may include optical components such as lenses, mirrors, prisms, windows, apertures, attenuators, and / or optical fibers. Additional mechanical components may be included in the instrument to enable control of one or more alignment components. Such mechanical components may include actuators, stepper motors, and / or knobs. An example of a suitable excitation source and alignment mechanism is described in U.S. Patent Application No. 15 / 161,088, filed May 20, 2016, entitled "Pulsed Laser and System" (which is incorporated in its entirety by reference). Another example of a beam steering module is described in U.S. Patent Application No. 15 / 842,720, filed December 14, 2017, entitled "Compact Beam Shaping and Steering Assembly" (incorporated herein by reference). An example of additional suitable excitation sources is described in "Integrated Device for Probing, Detecting, and Analyzing Molecules." This is described in U.S. Patent Application No. 14 / 821,688, filed on August 7, 2015, titled “FOR BING, DETECTING AND ANALYZING MOLECULES” (which is incorporated in its entirety by reference).

[0284] A photodetector positioned with the individual pixels of the integrated device may be configured and positioned to detect light emitted from the corresponding sample well of the pixel. A suitable example of a photodetector is described in U.S. Patent Application No. 14 / 821,656, filed August 7, 2015, entitled "Integrated Device for Temporal Binning of Received Photons" (which is incorporated in its entirety by reference). In some embodiments, the sample wells and their respective photodetectors may be aligned along a common axis. In this way, the photodetectors may overlap the sample wells within the pixels.

[0285] The characteristics of the detected emitted light can provide an index for identifying a marker associated with the emitted light. Such characteristics may include any preferred type of characteristics, including the arrival time of photons detected by the photodetector, the amount of photons accumulated over time by the photodetector, and / or the distribution of photons across two or more photodetectors. In some embodiments, the photodetector may have a configuration that allows detection of one or more timing characteristics (e.g., luminescence lifetime) associated with the emitted light of the sample. The photodetector may detect the distribution of photon arrival times after the excitation light pulse has propagated through the integrated device, and the distribution of arrival times may provide an index of the timing characteristics of the emitted light of the sample (e.g., a proxy for luminescence lifetime). In some embodiments, one or more photodetectors provide an index of the probability of emitted light emitted by a marker (e.g., emission intensity). In some embodiments, multiple photodetectors may be sized and arranged to capture the spatial distribution of emitted light. Output signals from one or more photodetectors may then be used to distinguish markers among multiple markers, and multiple markers may be used to identify a sample within a sample. In some embodiments, the sample can be excited by multiple excitation energies, and the emitted light and / or timing characteristics of the emitted light emitted by the sample in response to multiple excitation energies can distinguish one marker from several other markers.

[0286] During operation, parallel analysis of the sample in the sample well is performed by exciting part or all of the sample in the well using excitation light and detecting the signal from the sample emission with a photodetector. The light emitted from the sample can be detected by the corresponding photodetector and converted into at least one electrical signal. The electrical signal can be transmitted along a conduction line in the circuit of the integrated device and can be connected to an instrument interfaced to the integrated device. The electrical signal can be subsequently processed and / or analyzed. Processing or analysis of the electrical signal can be performed on a suitable computing device located either on or outside the instrument.

[0287] The device may include a user interface for controlling the operation of the device and / or integrated device. The user interface may be configured to allow the user to input information to the device, such as commands and / or settings used to control the functions of the device. In some embodiments, the user interface may include buttons, switches, dials, and a microphone for voice commands. The user interface may allow the user to receive feedback regarding the performance of the device and / or integrated device, such as information obtained from proper alignment and / or readout signals from photodetectors on the integrated device. In some embodiments, the user interface may provide feedback using a speaker to provide audible feedback. In some embodiments, the user interface may include indicator lights and / or a display screen to provide visual feedback to the user.

[0288] In some embodiments, the device may include a computer interface configured to connect to a computing device. The computer interface may be a USB interface, a FireWire® interface, or any other suitable computer interface. The computing device may be any general-purpose computer, such as a laptop computer or a desktop computer. In some embodiments, the computing device may be a server (e.g., a cloud-based server) accessible via a wireless network through a suitable computer interface. The computer interface may facilitate the communication of information between the device and the computing device. Input information for controlling and / or configuring the device to the computing device may be provided and transmitted to the device via the computer interface. Output information generated by the device may be received by the computing device via the computer interface. Output information may include feedback on the performance of the device, the performance of an integrated device, and / or data generated from readout signals of a photodetector.

[0289] In some embodiments, the instrument may include a processing device configured to analyze data received from one or more photodetectors of the integrated device and / or to transmit control signals to an excitation source. In some embodiments, the processing device may include a general-purpose processor, a specially-adapted processor (e.g., a central processing unit (CPU), e.g., one or more microprocessor cores or microcontroller cores, a field-programmable gate array (FPGA), an application-specific integrated circuit (ASIC), a custom integrated circuit, a digital signal processor (DSP), or a combination thereof). In some embodiments, processing of data from one or more photodetectors may be performed by both the instrument's processing device and an external computing device. In other embodiments, the external computing device may be omitted, and processing of data from one or more photodetectors may be performed solely by the processing device of the integrated device.

[0290] According to some embodiments, instruments configured to analyze a sample based on luminescence characteristics can detect differences in luminescence lifetime and / or intensity between different luminescent molecules, and / or differences in lifetime and / or intensity of the same luminescent molecule in different environments. The inventors recognize and appreciate that differences in luminescence lifetime can be used to distinguish the presence or absence of different luminescent molecules and / or to distinguish different environments or conditions to which luminescent molecules are exposed. In some cases, distinguishing luminescent molecules based on lifetime (e.g., rather than emission wavelength) can simplify the configuration of the system. For example, wavelength-discriminating optical elements (e.g., wavelength filters, dedicated detectors for each wavelength, dedicated pulsed light sources for different wavelengths, and / or diffractive optical elements) can be reduced in number or eliminated when distinguishing luminescent molecules based on lifetime. In some cases, a single-pulse light source operating at a single characteristic wavelength can be used to excite different luminescent molecules that emit within the same wavelength region of the optical spectrum but have measurably different lifetimes. An analysis system that uses a single pulsed light source, rather than multiple light sources operating at different wavelengths, to excite and distinguish different light-emitting molecules emitting in the same wavelength region can reduce operational and maintenance complexity, be more compact, and be manufactured at a lower cost.

[0291] While analytical systems based on emission lifetime analysis may have certain advantages, the amount of information and / or detection accuracy obtained by such systems can be increased by allowing additional detection techniques. For example, some embodiments of the system may be additionally configured to distinguish one or more properties of a sample based on emission wavelength and / or emission intensity. In some realizations, emission intensity may be used additionally or alternatively to distinguish different emission labels. For example, some emission labels may emit at significantly different intensities or have significantly different excitation probabilities (e.g., a difference of at least about 35%), even if their decay rates are similar. Different emission labels may be distinguishable based on intensity levels by referring to the bin signal against the measured excitation light.

[0292] According to some embodiments, different emission lifetimes can be distinguished using a photodetector configured to time-bin the emission event following the excitation of an emission marker. Time binning can occur during a single charge accumulation cycle for the photodetector. A charge accumulation cycle is the interval between readout events in which photogenerating carriers are accumulated in the bins of the time-binning photodetector. An example of a time-binning photodetector is described in U.S. Patent Application No. 14 / 821,656, filed August 7, 2015, entitled "INTEGRATED DEVICE FOR TEMPORAL BINNING OF RECEIVED PHOTONS" (incorporated herein by reference). In some embodiments, the time-binning photodetector may generate charge carriers in a photon absorption / carrier generation region and directly transfer the charge carriers to a charge carrier storage bin in a charge carrier storage region. In such embodiments, the time-binning photodetector may not include a carrier transfer / capture region. Such time-binning photodetectors may be referred to as “direct binning pixels.” An example of a time-binning photodetector including direct binning pixels is described in U.S. Patent Application No. 15 / 852,571, filed December 22, 2017, entitled “Integrated Photodetector with Direct Binning Pixel” (incorporated herein by reference).

[0293] In some embodiments, an anequal number of fluorophores of the same type may be linked to different reagents in the sample so that each reagent can be identified based on its emission intensity. For example, two fluorophores may be linked to a first labeled recognition molecule, while four or more fluorophores may be linked to a second labeled recognition molecule. Because there are anequal numbers of fluorophores, different excitation probabilities and fluorophore release probabilities may exist associated with different recognition molecules. For example, with the second labeled recognition molecule, more release events may occur during the signal accumulation interval, resulting in a significantly higher apparent intensity in the bottle than with the first labeled recognition molecule.

[0294] The inventors recognized and appreciated that distinguishing nucleotides or any other biological or chemical samples based on fluorophore attenuation rate and / or fluorophore intensity could allow for simplification of optical excitation and detection systems. For example, optical excitation may be performed with a single-wavelength light source (e.g., a light source that produces one characteristic wavelength rather than multiple light sources, or a light source that operates at multiple different characteristic wavelengths). In addition, wavelength-discriminating optical elements and filters may not be required in the detection system. Furthermore, a single photodetector may be used for each sample well to detect emissions from different fluorophores. The terms “characteristic wavelength” or “wavelength” are used to mean the central or dominant wavelength within a limited bandwidth of radiation (e.g., the central or peak wavelength within a 20 nm bandwidth output from a pulsed light source). In some cases, “characteristic wavelength” or “wavelength” may be used to mean the peak wavelength within the total bandwidth of the radiation output from the light source.

[0295] calculation technology Aspects of this application relate to computational techniques for analyzing data generated by polypeptide sequencing techniques described herein. As discussed above, for example, in relation to Figures 1A and 1B, the data generated using these sequencing techniques may include a series of signal pulses that indicate instances in which amino acid recognition molecules associate with amino acids exposed at the terminal ends of the sequenced polypeptide. The series of signal pulses may have one or more different features (e.g., changes in pulse duration, inter-pulse duration, and magnitude) over time, depending on the type of terminal amino acid at that point as the degradation process proceeds by sequentially removing amino acids. The resulting signal trace may include characteristic patterns arising from one or more different features associated with each amino acid. The computational techniques described herein may be implemented as a part for analyzing such data obtained using these sequencing techniques to identify amino acid sequences.

[0296] Some embodiments may include obtaining data during the polypeptide degradation process, analyzing the data to determine portions of the data corresponding to amino acids sequentially exposed at the ends of the polypeptide during the degradation process, and outputting the amino acid sequence corresponding to the polypeptide. Figure 11 is a diagram of an exemplary processing pipeline 1100 for identifying an amino acid sequence by analyzing data obtained using the polypeptide sequencing techniques described herein. As shown in Figure 11, analyzing the sequencing data 1102 may include outputting an amino acid sequence 1108 using association event identification techniques 1104 and amino acid identification techniques 1106.

[0297] As discussed herein, sequencing data 1102 can be obtained during the polypeptide degradation process. In some embodiments, sequencing data 1102 serves as an indicator of the terminal amino acid identity of the polypeptide during the degradation process. In some embodiments, sequencing data 1102 serves as an indicator of the signal generated by one or more amino acid recognition molecules that bind to different types of terminal amino acids during the degradation process. Exemplary sequencing data are shown in Figures 1A and 1B discussed above.

[0298] Depending on how the signal is generated during the degradation process, the sequencing data 1102 can be an indicator of one or more different types of signals. In some embodiments, the sequencing data 1102 can be an indicator of a luminescence signal generated during the degradation process. For example, a luminescent label may be used to label an amino acid recognition molecule, and the luminescence emitted by the luminescent label may be detected when the amino acid recognition molecule associates with a specific amino acid, resulting in a luminescence signal. In some embodiments, the sequencing data 1102 can be an indicator of an electrical signal generated during the degradation process. For example, the polypeptide molecule to be sequenced may be immobilized on a nanopore, and the electrical signal (e.g., a change in conductance) may be detected when the amino acid recognition molecule associates with a specific amino acid.

[0299] Some embodiments include analyzing the sequencing data 1102 to determine portions of the sequencing data 1102 corresponding to amino acids sequentially exposed at the ends of a polypeptide during the degradation process. As shown in Figure 11, the association event identification technique 1104 may access the sequencing data 1102 and analyze the sequencing data to identify portions of the sequencing data 1102 corresponding to association events. Association events may correspond to characteristic patterns in the data, such as CP1 and CP2, shown in Figure 1B. In some embodiments, the association event identification technique 1104 may include detecting a series of cleavage events and determining portions of the sequencing data 1102 between sequential cleavage events. For example, the cleavage event between CP1 and CP2, shown in Figure 1B, may be detected such that a first portion of the data corresponding to CP1 can be identified as a first association event, and a second portion of the data corresponding to CP2 can be identified as a second association event.

[0300] Some embodiments include identifying the type of amino acid for one or more determined portions of sequencing data 1102. As shown in Figure 11, the amino acid identification technique 1106 may be used to determine the type of amino acid for one or more association events identified by the association event identification technique 1104. In some embodiments, the individual portions of data identified by the association event identification technique 1104 may include pulse patterns, and the amino acid identification technique 1106 may determine the type of amino acid for one or more portions based on their respective pulse patterns. Referring to Figure 1B, the amino acid identification technique 1106 may identify a first type of amino acid for CP1 and a second type of amino acid for CP2. In some embodiments, determining the type of amino acid may include identifying the time quantity within a portion of the data, such as a portion identified using the association event identification technique 1104 when the data is above a threshold, and comparing the time quantity with the duration for the portion of the data. For example, identifying the type of amino acid for CP1 may include the time quantity within CP1 when the signal is above a threshold, e.g., when the signal is M L This may include determining the time interval pd when it is greater than 1, and comparing it to the total duration for CP1. In some embodiments, determining the amino acid type may include identifying one or more pulse durations for one or more portions of the data identified by the association event identification technique 1102. For example, identifying the amino acid type for CP1 may include determining a pulse duration such as time interval pd for CP1. In some embodiments, determining the amino acid type may include identifying one or more inter-pulse durations for one or more portions of the data identified using the association event identification technique 1104. For example, identifying the amino acid type for CP1 may include identifying an inter-pulse duration such as ipd.

[0301] By identifying the amino acid types for sequential portions of sequencing data 1102, the amino acid identification technique 1106 can output an amino acid sequence 1108 corresponding to a polypeptide. In some embodiments, the amino acid sequence includes a set of amino acids corresponding to portions of data identified using the association event identification technique 1104.

[0302] Figure 12 is a flowchart of an exemplary process 1200 for determining the amino acid sequence of a polypeptide molecule according to some embodiments of the techniques described herein. Process 1200 can be performed on any suitable computing device (e.g., a single computing device, multiple computing devices located at multiple physical locations co-located or at multiple physical locations separated from each other, one or more computing device portions of a cloud computing system, etc.), since the embodiments of the techniques described herein are not limited to these embodiments. In some embodiments, association event identification techniques 1104 and amino acid identification techniques 1106 may perform part or all of process 1200 to determine the amino acid sequence. Process 1200 begins with operation 1202, which involves contacting a single polypeptide molecule with one or more terminal amino acid recognition molecules. Process 1200 then proceeds to operation 1104, which involves detecting a series of signal pulses indicative of the association between one or more terminal amino acid recognition molecules and sequential amino acids exposed at the ends of the single polypeptide while the single polypeptide is being degraded. The series of pulses may enable sequencing of the single polypeptide molecule, for example, by using association event identification techniques 1104 and amino acid identification techniques 1106.

[0303] In some embodiments, process 1200 may include operation 1206, which includes identifying a first type of amino acid of a single polypeptide molecule based on a first characteristic pattern of a series of signal pulses, for example, by using amino acid identification technique 1106.

[0304] Figure 13 is a flowchart of an exemplary process 1300 for determining the amino acid sequence of a polypeptide corresponding to some embodiment of the technology described herein. Process 1300 can be performed on any suitable computing device (e.g., a single computing device, multiple computing devices located at multiple physical locations co-located or at multiple physical locations separated from each other, one or more computing device portions of a cloud computing system, etc.), since the embodiments of the technology described herein are not limited to these embodiments. In some embodiments, association event identification techniques 1104 and amino acid identification techniques 1106 may perform part or all of process 1300 to determine the amino acid sequence.

[0305] Process 1300 begins with operation 1302, which provides data during the polypeptide degradation process. In some embodiments, the data serves as an indicator of the identity of the terminal amino acids of the polypeptide during the degradation process. In some embodiments, the data serves as an indicator of the signal generated by one or more amino acid recognition molecules that bind to different types of terminal amino acids during the degradation process. In some embodiments, the data serves as an indicator of the luminescence signal generated during t...

Claims

1. A recombinant amino acid-binding protein having an amino acid sequence that is at least 80% identical to a sequence selected from Table 1 or Table 2, and containing one or more labels.

2. The recombinant amino acid-binding protein according to claim 1, wherein the one or more labels include luminescence labels or conductivity labels.

3. The recombinant amino acid-binding protein according to claim 2, wherein the luminescence label comprises at least one fluorophore dye molecule.

4. The recombinant amino acid-binding protein according to claim 2 or 3, wherein the luminescence label comprises 20 or fewer fluorophore dye molecules.

5. The recombinant amino acid-binding protein according to any one of claims 2 to 4, wherein the luminescent label comprises at least one FRET pair including a donor label and an acceptor label.

6. The recombinant amino acid-binding protein according to claim 5, wherein the ratio of the donor label to the acceptor label is 1:1, 2:1, 3:1, 4:1, or 5:

1.

7. The recombinant amino acid-binding protein according to claim 5, wherein the ratio of the acceptor label to the donor label is 1:1, 2:1, 3:1, 4:1, or 5:

1.

8. The conductivity label comprises a charged polymer, wherein the recombinant amino acid-binding protein is according to any one of claims 1 to 7.

9. The recombinant amino acid-binding protein according to any one of claims 1 to 8, wherein the one or more labels include a tag sequence.

10. The recombinant amino acid-binding protein according to claim 9, wherein the tag sequence comprises one or more of a purified tag, a cleavage site, and a biotinylated sequence.

11. The recombinant amino acid-binding protein according to claim 10, wherein the biotinylated sequence includes at least one biotin ligase recognition sequence.

12. The recombinant amino acid-binding protein according to claim 10 or 11, wherein the biotinylated sequence comprises two biotin ligase recognition sequences oriented in series.

13. The recombinant amino acid-binding protein according to any one of claims 1 to 12, wherein the one or more labels include a biotin moiety.

14. The recombinant amino acid-binding protein according to claim 13, wherein the biotin portion comprises at least one biotin molecule.

15. The recombinant amino acid-binding protein according to claim 13 or 14, wherein the biotin portion is a bisbiotin portion.

16. The recombinant amino acid-binding protein according to claim 14 or 15, wherein the label comprises at least one biotin ligase recognition sequence to which at least one biotin molecule is attached.

17. The recombinant amino acid-binding protein according to any one of claims 1 to 16, wherein the one or more labels include one or more polyol moieties.

18. The recombinant amino acid-binding protein according to claim 17, wherein the one or more polyol moieties include dextran, polyvinylpyrrolidone, polyethylene glycol, polypropylene glycol, polyoxyethylene glycol, polyvinyl alcohol, or a combination or variation thereof.

19. The recombinant amino acid-binding protein according to any one of claims 1 to 18, wherein the recombinant amino acid-binding protein comprises one or more non-natural amino acids to which the one or more labels are attached.

20. The recombinant amino acid-binding protein according to any one of claims 1 to 19, wherein the amino acid sequence is 80-90%, 90-95%, or at least 95% identical to a sequence selected from Table 1.

21. A composition comprising a recombinant amino acid-binding protein according to any one of claims 1 to 20.

22. A polypeptide sequencing reaction composition comprising two or more amino acid recognition molecules, wherein at least one of the two or more amino acid recognition molecules is a recombinant amino acid-binding protein according to any one of claims 1 to 20.

23. The polypeptide sequencing reaction composition according to claim 22, comprising at least one type of cleavage reagent.

24. A polypeptide sequencing method, Contacting a polypeptide with the polypeptide sequencing reaction composition according to claim 22 or 23, A method comprising sequencing a polypeptide by detecting a series of interactions between the polypeptide and at least one amino acid recognition molecule while the polypeptide is being degraded.

25. A polypeptide sequencing reaction mixture comprising an amino acid-binding protein and a peptidase, wherein the amino acid-binding protein comprises one or more labels, and the molar ratio of the labeled amino acid-binding protein to the peptidase is about 1:1,000 to about 1:1, or about 1:1 to about 100:

1.

26. The polypeptide sequencing reaction mixture according to claim 25, wherein the molar ratio is approximately 1:100 to approximately 1:1, or approximately 1:1 to approximately 10:

1.

27. The polypeptide sequencing reaction mixture according to claim 25 or 26, wherein the molar ratios are approximately 1:1,000, approximately 1:500, approximately 1:200, approximately 1:100, approximately 1:10, approximately 1:5, approximately 1:2, approximately 1:1, approximately 5:1, approximately 10:1, approximately 50:1, and approximately 100:

1.

28. The polypeptide sequencing reaction mixture according to any one of claims 25 to 27, wherein the amino acid-binding protein is a synthetic protein or a recombinant protein.

29. The polypeptide sequencing reaction mixture according to any one of claims 25 to 28, wherein the amino acid-binding protein is a degradation pathway protein, an inactivated peptidase, an antibody, an aminotransferase, a tRNA synthetase, or an SH2 domain-containing protein or a fragment thereof.

30. The polypeptide sequencing reaction mixture according to any one of claims 25 to 29, wherein the amino acid-binding protein is a ClpS protein, a Gid protein, a UBR box protein or a UBR box domain-containing fragment thereof, or a p62 protein or a ZZ domain-containing fragment thereof.

31. The polypeptide sequencing reaction mixture according to any one of claims 25 to 30, wherein the amino acid-binding protein is a ClpS protein.

32. The polypeptide sequencing reaction mixture according to claim 31, wherein the ClpS protein is ClpS1 or ClpS2 derived from A. tumifaciens, C. crescentus, E. coli, S. elongatus, P. falciparum, T. elongatus, or homologs thereof.

33. The polypeptide sequencing reaction mixture according to any one of claims 25 to 32, wherein the amino acid-binding protein is a protein having at least 80%, 80-90%, 90-95%, or at least 95% identical amino acid sequences to a sequence selected from Table 1 or Table 2.

34. The polypeptide sequencing reaction mixture according to any one of claims 25 to 33, wherein the amino acid-binding protein is the recombinant amino acid-binding protein according to any one of claims A1 to A19.

35. The polypeptide sequencing reaction mixture according to any one of claims 25 to 34, wherein the peptidase is an exopeptidase.

36. The polypeptide sequencing reaction mixture according to any one of claims 25 to 35, wherein the peptidase is an aminopeptidase or a carboxypeptidase.

37. The polypeptide sequencing reaction mixture according to any one of claims 25 to 36, wherein the peptidase is a nonspecific exopeptidase that cleaves more than one type of amino acid from the end of the peptidase.

38. The polypeptide sequencing reaction mixture according to any one of claims 25 to 37, wherein the peptidase is proline aminopeptidase, proline iminopeptidase, glutamic acid / aspartic acid specific aminopeptidase, methionine specific aminopeptidase, or zinc metalloproteinase.

39. The polypeptide sequencing reaction mixture according to any one of claims 25 to 38, wherein the peptidase is an enzyme having at least 80%, 80-90%, 90-95%, or at least 95% identical amino acid sequences to a sequence selected from Table 4 or Table 5.

40. The polypeptide sequencing reaction mixture according to any one of claims 25 to 39, wherein the reaction mixture comprises one or more amino acid-binding proteins.

41. The polypeptide sequencing reaction mixture according to any one of claims 25 to 40, wherein the reaction mixture comprises one or more peptidases.

42. The polypeptide sequencing reaction mixture according to any one of claims 25 to 41, wherein the reaction mixture comprises peptide molecules immobilized on the surface.

43. A polypeptide sequencing reaction mixture comprising a single peptide molecule, at least one peptidase molecule, and at least three amino acid recognition molecules.

44. A polypeptide sequencing reaction mixture according to claim 43, comprising at least one and up to 10 peptidase molecules.

45. A polypeptide sequencing reaction mixture according to claim 43 or 44, comprising at least one and up to five peptidase molecules.

46. A polypeptide sequencing reaction mixture according to any one of claims 43 to 45, comprising at least one and up to three peptidase molecules.

47. A polypeptide sequencing reaction mixture according to any one of claims 43 to 46, comprising at least three and up to 30 amino acid recognition molecules.

48. The polypeptide sequencing reaction mixture according to claim 47, comprising up to 20, up to 10, or up to 5 amino acid recognition molecules.

49. A substrate comprising an array of sample wells, wherein at least one sample well of the array comprises a polypeptide sequencing reaction mixture according to any one of claims 43 to 48.

50. The substrate according to claim 49, wherein the at least one sample well includes a bottom surface, and the single polypeptide molecule is immobilized on the bottom surface.

51. A substrate comprising an array of sample wells, wherein at least one sample well of the array comprises a single polypeptide molecule, a cleaving means, and a binding means, wherein the binding means and the cleaving means are configured to achieve at least 10 association events between the binding means and a terminal amino acid on the polypeptide before the terminal amino acid is removed from the polypeptide by the cleaving means.

52. An amino acid recognition molecule comprising a polypeptide having at least a first amino acid-binding protein and a second amino acid-binding protein linked at their terminal ends, wherein the first amino acid-binding protein and the second amino acid-binding protein are separated by a linker containing at least two amino acids.

53. The amino acid recognition molecule according to claim 52, wherein the first amino acid-binding protein and the second amino acid-binding protein are the same.

54. The amino acid recognition molecule according to claim 52, wherein the first amino acid-binding protein and the second amino acid-binding protein are different from each other.

55. The amino acid recognition molecule according to any one of claims 52 to 54, wherein the first amino acid-binding protein and the second amino acid-binding protein each independently have at least 80% amino acid sequence identity with respect to an amino acid sequence selected from Table 1 or Table 2.

56. The linker is an amino acid recognition molecule according to any one of claims 52 to 55, comprising up to 100 amino acids.

57. The linker comprises about 5 to about 50 amino acids, according to any one of claims 52 to 56, the amino acid recognition molecule.

58. Equation (I): (Z 1 -X 1 ) n -Z 2 (I) (In the formula, Z 1 and Z 2 It is an independent amino acid-binding protein, X 1 (where n is a linker containing at least two amino acids, the amino acid-binding proteins are terminally linked by the linker, and n is an integer between 1 and 5 (including the values ​​at both ends)). An amino acid recognition molecule containing polypeptides.

59. Z 1 and Z 2 is the amino acid recognition molecule according to claim 58, comprising amino acid binding proteins of the same type.

60. Z 1 and Z 2 The amino acid recognition molecule according to claim 58, comprising amino acid-binding proteins of different types from each other.

61. Z 1 and Z 2 The amino acid recognition molecule according to any one of claims 58 to 60, wherein the amino acid recognition molecule relates independently and optionally to a labeling component comprising at least one detection label.

62. The amino acid recognition molecule according to any one of claims 58 to 61, wherein the polypeptide further comprises a tag sequence.

63. The amino acid recognition molecule according to claim 62, wherein the tag sequence is at or near the end of the polypeptide.

64. The amino acid recognition molecule according to claim 62 or 63, wherein the tag sequence includes at least one biotin ligase recognition sequence.

65. The amino acid recognition molecule according to any one of claims 62 to 64, wherein the tag sequence includes two biotin ligase recognition sequences oriented in series.

66. Z 1 and Z 2 The amino acid recognition molecule according to any one of claims 58 to 65, wherein each of the amino acids independently has an amino acid sequence that is at least 80% identical to an amino acid sequence selected from Table 1 or Table 2.

67. X 1 The amino acid recognition molecule according to any one of claims 58 to 66, comprising up to 100 amino acids.

68. X 1 The amino acid recognition molecule according to any one of claims 58 to 67, comprising approximately 5 to approximately 50 amino acids.

69. An amino acid recognition molecule comprising a polypeptide having an amino acid-binding protein and a labeled protein linked at their terminal ends, wherein the amino acid-binding protein and the labeled protein are separated by a linker containing at least two amino acids.

70. The linker is an amino acid recognition molecule according to claim 69, comprising up to 100 amino acids.

71. The linker comprises about 5 to about 50 amino acids, according to the amino acid recognition molecule of claim 69 or 70.

72. The labeled protein has a molecular weight of at least 10 kDa, and is an amino acid recognition molecule according to any one of claims 69 to 71.

73. The labeled protein has a molecular weight of about 10 kDa to about 150 kDa, and is an amino acid recognition molecule according to any one of claims 69 to 72.

74. The labeled protein has a molecular weight of about 15 kDa to about 100 kDa, and is an amino acid recognition molecule according to any one of claims 69 to 73.

75. The labeled protein comprises at least 50 amino acids, according to any one of claims 69 to 74, an amino acid recognition molecule.

76. The labeled protein comprises about 50 to about 1,000 amino acids, according to any one of claims 69 to 75, an amino acid recognition molecule.

77. The labeled protein comprises about 100 to about 750 amino acids, the amino acid recognition molecule according to any one of claims 69 to 76.

78. The amino acid recognition molecule according to any one of claims 69 to 77, wherein the labeled protein comprises a protein selected from the group consisting of DNA polymerase, maltose-binding protein, glutathione S-transferase, green fluorescent protein, and SNAP tag.

79. The labeled protein is an amino acid recognition molecule according to any one of claims 69 to 78, comprising luminescence labeling.

80. The amino acid recognition molecule according to claim 79, wherein the luminescent label comprises at least one fluorophore dye molecule.

81. The amino acid-recognizing molecule according to any one of claims 69 to 80, wherein the amino acid-binding protein is a Gid protein, a UBR box protein or a fragment containing its UBR box domain, a p62 protein or a fragment containing its ZZ domain, or a ClpS protein.

82. The amino acid-recognizing molecule according to any one of claims 69 to 81, wherein the amino acid-binding protein has an amino acid sequence that is at least 80% identical to an amino acid sequence selected from Table 1 or Table 2.

83. A polypeptide sequencing method, The process involves contacting a single polypeptide molecule in a reaction mixture with a composition containing binding and cleaving means, A method wherein the binding means and the cleaving means are configured to achieve at least 10 association events between the binding means and the terminal amino acids on the polypeptide before the terminal amino acids are removed from the polypeptide by the cleaving means.

84. The method according to claim 83, wherein the binding means and the cleaving means are configured to achieve at least 10 and up to 1,000 association events before the removal of the terminal amino acid.

85. The method according to claim 83 or 84, wherein the binding means and the cleaving means are configured to achieve up to 500, up to 200, or up to 100 association events before the removal of the terminal amino acid.

86. The method according to any one of claims 83 to 85, wherein the terminal amino acids are exposed at the end of the polypeptide during a cleavage event prior to the at least 10 association events.

87. The method according to claim 86, wherein the at least 10 meeting events occur after the cutting event.

88. The method according to any one of claims 83 to 87, wherein the coupling means and the disconnecting means are configured to achieve a time interval of at least one minute between disconnection events.

89. The method according to claim 88, wherein the time interval is approximately 1 minute to approximately 20 minutes, approximately 5 minutes to approximately 15 minutes, or approximately 1 minute to approximately 10 minutes.

90. The method according to any one of claims 83 to 89, wherein the binding means comprises one or more amino acid recognition molecules, and the cleavage means comprises one or more peptidase molecules.

91. The method according to claim 90, wherein the molar ratio of the amino acid recognition molecule to the peptidase molecule is configured to achieve at least 10 association events before the removal of the terminal amino acid.

92. The method according to claim 91, wherein the molar ratio of the amino acid recognition molecule to the peptidase molecule is about 1:1,000 to about 1:1 or about 1:1 to about 100:

1.

93. The method according to claim 91, wherein the molar ratio of the amino acid recognition molecule to the peptidase molecule is about 1:100 to about 1:1 or about 1:1 to about 10:

1.

94. The method according to claim 91, wherein the molar ratio of the amino acid recognition molecule to the peptidase molecule is about 1:1,000, about 1:500, about 1:200, about 1:100, about 1:10, about 1:5, about 1:2, about 1:1, about 5:1, about 10:1, about 50:1, and about 100:

1.

95. To obtain data during the polypeptide degradation process, Analyzing the data to determine a portion of the data corresponding to an amino acid sequentially exposed at the end of the polypeptide during the degradation process, wherein each independent portion includes a pulse pattern having at least one pulse duration, and the pulse pattern includes an average pulse duration of approximately 1 millisecond to approximately 10 seconds. A method comprising outputting an amino acid sequence corresponding to the polypeptide.

96. The method according to claim 95, wherein the average pulse duration is approximately 50 milliseconds to approximately 2 seconds.

97. The method according to claim 95 or 96, wherein the average pulse duration is approximately 50 milliseconds to approximately 500 milliseconds or approximately 500 milliseconds to approximately 2 seconds.

98. The method according to any one of claims 95 to 97, wherein the pulse pattern of the first type of amino acid differs from the pulse pattern of the second type of amino acid by an average pulse duration of at least 10 milliseconds.

99. The method according to any one of claims 95 to 97, wherein the pulse pattern of the first type of amino acid differs from the pulse pattern of the second type of amino acid by an average pulse duration of about 10 milliseconds to about 100 milliseconds or about 100 milliseconds to about 10 seconds.

100. A polypeptide sequencing method, Contacting a single polypeptide molecule with one or more amino acid recognition molecules, The process includes sequencing a single polypeptide molecule by detecting a series of signal pulses that serve as an indicator of the association between one or more amino acid recognition molecules and sequential amino acids exposed at the termini of the single polypeptide molecule while the single polypeptide is being degraded. A method wherein the association of one or more amino acid recognition molecules with each type of amino acid exposed at the terminal generates a characteristic pattern of the series of signal pulses that is different from other types of amino acids exposed at the terminal, and the signal pulses of the characteristic pattern have an average pulse duration of about 1 millisecond to about 10 seconds.

101. A method for sequencing polypeptides, Contacting a single polypeptide molecule in a reaction mixture with a composition containing one or more amino acid recognition molecules and a cleavage reagent, A series of signal pulses that serve as an indicator of the association of one or more amino acid recognition molecules with the terminal of the single polypeptide molecule in the presence of the cleavage reagent, wherein the series of signal pulses includes detecting a series of signal pulses that are indicators of a series of amino acids exposed at the terminal over time as a result of terminal amino acid cleavage by the cleavage reagent. A method wherein the association of one or more amino acid recognition molecules with each type of amino acid exposed at the terminal generates a characteristic pattern of the series of signal pulses that is different from other types of amino acids exposed at the terminal, and the signal pulses of the characteristic pattern have an average pulse duration of about 1 millisecond to about 10 seconds.

102. A polypeptide sequencing method, a) Identifying the first amino acid at the end of a single polypeptide molecule, b) Removing the first amino acid to expose the second amino acid at the terminal end of the single polypeptide molecule, c) Identifying the second amino acid at the terminal end of the single polypeptide molecule, (a) to (c) are carried out in a single reaction mixture containing one or more amino acid recognition molecules. The first amino acid and the second amino acid are identified by detecting a series of signal pulses that indicate the association between one or more amino acid recognition molecules and the terminal end of the single polypeptide molecule. A method comprising the association of one or more amino acid recognition molecules with the first amino acid to generate a characteristic pattern of a series of signal pulses different from that of the second amino acid, wherein the signal pulses of the characteristic pattern include an average pulse duration of about 1 millisecond to about 10 seconds.

103. A method for identifying the amino acids of a polypeptide, The process involves bringing a single polypeptide molecule into contact with one or more amino acid recognition molecules bound to the single polypeptide molecule, To detect a series of signal pulses that serve as an indicator of association between one or more amino acid recognition molecules and the single polypeptide molecule under polypeptide degradation conditions, A method comprising identifying a first type of amino acid of a single polypeptide molecule based on a characteristic pattern in a series of signal pulses, wherein the signal pulses of the characteristic pattern include an average pulse duration of about 1 millisecond to about 10 seconds.

104. The method according to any one of claims 100 to 103, wherein the average pulse duration is approximately 50 milliseconds to approximately 2 seconds.

105. The method according to any one of claims 100 to 104, wherein the average pulse duration is approximately 50 milliseconds to approximately 500 milliseconds or approximately 500 milliseconds to approximately 2 seconds.

106. The method according to any one of claims 100 to 105, wherein the characteristic pattern includes at least 10 signal pulses.

107. The method according to any one of claims 100 to 106, wherein the characteristic pattern of at least one type of amino acid comprises about 50 to about 200 signal pulses.

108. The method according to any one of claims 100 to 107, wherein the characteristic pattern of at least one type of amino acid comprises about 25 to about 100 signal pulses.

109. The method according to any one of claims 100 to 108, wherein the characteristic pattern of the first type of amino acid differs from the characteristic pattern of the second type of amino acid by an average pulse duration of at least 10 milliseconds.

110. The method according to any one of claims 100 to 109, wherein the characteristic pattern of the first type of amino acid differs from the characteristic pattern of the second type of amino acid by an average pulse duration of about 10 milliseconds to about 100 milliseconds or about 100 milliseconds to about 1 second.

111. The method according to any one of claims 100 to 110, wherein the method is carried out in a reaction mixture comprising a composition containing one or more amino acid recognition molecules and one or more cleavage reagents.

112. The method according to claim 111, wherein the molar ratio of amino acid recognition molecules to cleavage reagent in the reaction mixture is about 1:1,000 to about 1:1 or about 1:1 to about 100:

1.

113. The method according to claim 111, wherein the molar ratio of amino acid recognition molecules to cleavage reagent in the reaction mixture is about 1:100 to about 1:1 or about 1:1 to about 10:

1.

114. The method according to claim 111, wherein the molar ratio of amino acid recognition molecules to cleavage reagents in the reaction mixture is about 1:1,000, about 1:500, about 1:200, about 1:100, about 1:10, about 1:5, about 1:2, about 1:1, about 5:1, about 10:1, about 50:1, and about 100:

1.

115. The method according to any one of claims 100 to 114, wherein the one or more amino acid recognition molecules include one or more terminal amino acid recognition molecules.

116. The method according to any one of claims 100 to 115, wherein each of the one or more amino acid recognition molecules comprises an amino acid-binding protein having an amino acid sequence that is at least 80% identical to an amino acid sequence selected from Table 1 or Table 2.

117. At least one hardware processor, A system comprising: at least one non-temporary computer-readable storage medium storing processor-executable instructions that cause the at least one hardware processor to perform the method according to any one of claims 100 to 116 when executed by the at least one hardware processor.

118. At least one non-temporary computer-readable storage medium storing processor-executable instructions that cause the at least one hardware processor to perform the method according to any one of claims 100 to 116 when executed by the at least one hardware processor.

119. Formula (II): A-(Y) n -D (II) (In the formula, A is an amino acid binding component comprising at least one amino acid binding protein having at least 80% amino acid sequence identity with respect to an amino acid sequence selected from Table 1 or Table 2, Each of Y is a polymer that forms covalent or non-covalent linking groups. n is an integer between 0 and 10 (including the values ​​at both ends). Furthermore, D is a labeling component comprising at least one detection label, wherein D has a diameter of less than 20 nm (200 Å). A composition containing a soluble amino acid recognition molecule.

120. The soluble amino acid recognition molecule is composed of A and Y linked at their terminal ends. 1 The polypeptide comprises A and Y 1 The composition according to claim 119, wherein the elements are separated by a linker containing at least two amino acids.

121. The composition according to claim 120, wherein the linker comprises up to 100 amino acids.

122. The composition according to claim 120 or 121, wherein the linker comprises about 5 to about 50 amino acids.

123. Y 1 The composition according to any one of claims 119 to 122, wherein is a protein having a molecular weight of at least 10 kDa.

124. Y 1 The composition according to any one of claims 119 to 123, wherein is a protein having a molecular weight of about 10 kDa to about 150 kDa.

125. Y 1 The composition according to any one of claims 119 to 124, wherein the protein has a molecular weight of about 15 kDa to about 100 kDa.

126. Y 1 The composition according to any one of claims 119 to 125, wherein the protein comprises at least 50 amino acids.

127. Y 1 The composition according to any one of claims 119 to 126, wherein the protein is a protein containing about 50 to about 1,000 amino acids.

128. Y 1 The composition according to any one of claims 119 to 127, wherein the protein comprises about 100 to about 750 amino acids.

129. Y 1 The composition according to any one of claims 119 to 128, wherein is a protein selected from the group consisting of DNA polymerase, maltose-binding protein, glutathione S-transferase, green fluorescent protein, and SNAP tag.

130. The composition according to any one of claims 119 to 129, wherein A comprises a polypeptide having at least a first amino acid-binding protein and a second amino acid-binding protein linked at their terminal ends, the first amino acid-binding protein and the second amino acid-binding protein being separated by a linker containing at least two amino acids.

131. The composition according to claim 130, wherein the first amino acid-binding protein and the second amino acid-binding protein are the same.

132. The composition according to claim 130 or 131, wherein the first amino acid-binding protein and the second amino acid-binding protein are different from each other.

133. The composition according to any one of claims 130 to 132, wherein the first amino acid-binding protein and the second amino acid-binding protein each independently have at least 80% identical amino acid sequences to an amino acid sequence selected from Table 1 or Table 2.

134. The linker comprises up to 100 amino acids, according to any one of claims 130 to 133.

135. The composition according to any one of claims 130 to 134, wherein the linker comprises about 5 to about 50 amino acids.

136. Formula (III): A-Y 1 -D (III) (In the formula, A is an amino acid binding component comprising at least one amino acid binding protein having at least 80% amino acid sequence identity with respect to an amino acid sequence selected from Table 1 or Table 2, Y1 is a nucleic acid or polypeptide. Furthermore, D is a labeling component that includes at least one detection label, Y 1 When is a nucleic acid, the nucleic acid forms a covalent or non-covalent linking group, Y 1 When is a polypeptide, the polypeptide is 50 × 10 -9 Dissociation constant less than M (K D (Forms non-covalent linking groups characterized by) An amino acid recognition molecule.

137. The amino acid recognition molecule according to claim 136, wherein A comprises a polypeptide having at least a first amino acid-binding protein and a second amino acid-binding protein linked at their terminal ends, and the first amino acid-binding protein and the second amino acid-binding protein are separated by a linker containing at least two amino acids.

138. The amino acid recognition molecule according to claim 137, wherein the first amino acid-binding protein and the second amino acid-binding protein are the same.

139. The amino acid recognition molecule according to claim 137 or 138, wherein the first amino acid-binding protein and the second amino acid-binding protein are different from each other.

140. The amino acid recognition molecule according to any one of claims 137 to 139, wherein the first amino acid-binding protein and the second amino acid-binding protein each independently have an amino acid sequence that is at least 80% identical to an amino acid sequence selected from Table 1 or Table 2.

141. The linker is an amino acid recognition molecule according to any one of claims 137 to 140, comprising up to 100 amino acids.

142. The linker comprises about 5 to about 50 amino acids, the amino acid recognition molecule according to any one of claims 137 to 141.

143. nucleic acids and, At least one amino acid recognition molecule attached to a first attachment site on the nucleic acid, comprising an amino acid-binding protein having an amino acid sequence that is at least 80% identical to the amino acid sequence selected from Table 1 or Table 2, An amino acid recognition molecule comprising at least one detectable label attached to a second attachment site on the nucleic acid, The nucleic acid is an amino acid recognition molecule that forms a covalent or non-covalent linking group between the at least one amino acid recognition molecule and the at least one detectable label.

144. A polyvalent protein containing at least two ligand-binding sites, At least one amino acid recognition molecule attached to the protein via a first ligand moiety bound to a first ligand binding site on the protein, comprising an amino acid-binding protein having an amino acid sequence that is at least 80% identical to an amino acid sequence selected from Table 1 or Table 2, An amino acid recognition molecule comprising: at least one detectable label attached to the protein via a second ligand moiety bound to a second ligand binding site on the protein;