Means and methods for determining the sequence of single-molecule peptides

The use of aminopeptidases and chemical cleavage inducers for measuring residence time allows for efficient single-molecule peptide sequencing, addressing the complexity and cost issues of existing protein sequencing methods and enabling digital protein quantification.

JP7834427B2Active Publication Date: 2026-03-24VLAAMS INTERUNIVERSITAIR INST VOOR BIOTECHNOLOGIE VZW +1
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2018-09-28
Publication Date
2026-03-24

AI Technical Summary

Technical Problem

Current protein sequencing technologies, particularly those based on liquid chromatography-mass spectrometry (LC-MS), are complex, expensive, and require significant time for detection, making single-molecule sequencing impractical and hindering applications like single-cell proteomics. Additionally, proteins cannot be amplified like DNA, and existing methods face challenges in distinguishing the 20 amino acids for precise identification.

Method used

A method involving the use of catalytically active aminopeptidases or chemical cleavage inducers, such as isothiocyanates, to sequentially identify N-terminal amino acids by measuring the residence time of the cleavage inducer on the peptide, allowing for single-molecule peptide sequencing.

Benefits of technology

Enables efficient and simplified single-molecule peptide sequencing by identifying N-terminal amino acids based on the kinetics of the cleavage reaction, reducing complexity and cost, and facilitating digital protein quantification.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007834427000005
    Figure 0007834427000005
  • Figure 0007834427000006
    Figure 0007834427000006
  • Figure 0007834427000007
    Figure 0007834427000007
Patent Text Reader

Abstract

The present invention relates to the field of biochemistry, more particularly to proteomics, more particularly to protein sequencing, and even more particularly to single molecule peptide sequencing.The present invention discloses a means and method for single molecule protein sequencing and / or amino acid identification using a cleavage inducer.The cleavage inducer is not specific to one particular amino acid, but cleaves polypeptide step by step from the N-terminus, and provides information about the identity of the cleaved amino acid based on the kinetics of the reaction.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] Field of the present invention This invention relates to the field of biochemistry, more specifically to proteomics, more specifically to protein sequencing, and even more specifically to single-molecule peptide sequencing. This invention discloses means and methods for single-molecule protein sequencing and / or amino acid identification using cleavage inducers. The cleavage inducer is not specific to a single amino acid, but cleaves a polypeptide step by step from the N-terminus, and provides information regarding the identity of the cleaved amino acids based on the kinetics of the engagement between the cleavage inducer and the polypeptide, or information regarding the kinetics of the polypeptide cleavage reaction. [Background technology]

[0002] background Technologies for high-throughput sequencing of biomolecules (DNA, RNA, and proteins) are rapidly developing and are essential to modern science and medicine. DNA sequencing, in particular, has made a significant leap forward, now shifting from next-generation approaches to third-generation approaches operating at the single-molecule level (e.g., Pacific Biosciences, Oxford Nanopores; Ambardar et al. 2016 Indian J Microbiol 56:394-404). The uniform nature of DNA and the broad panel of available molecular tools are truly driving this field forward. In contrast, protein / peptide sequencing lags behind, largely relying on liquid chromatography-mass spectrometry (LC-MS) based techniques.

[0003] While LC-MS instrumentation has advanced significantly in terms of speed, sensitivity, and resolution, operating and maintaining the instruments remains highly complex and expensive. Furthermore, MS-based proteomics still requires approximately 10 minutes of detection time. 6 A copy of the protein / peptide is required. Single-molecule sequencing is currently impractical, and therefore applications such as single-cell proteomics are unmanageable. Moreover, since proteins cannot be amplified in the same way that DNA can be amplified, single-molecule sequencing is conceptually more suitable and would also enable digital protein quantification.

[0004] The next generation of protein sequencing concepts has emerged, but faces significant challenges. Firstly, there are 20 amino acids to distinguish (compared to only 4 nucleotides), and obtaining a single measurable parameter for precise identification of each amino acid seems unlikely. However, attributing a certain amino acid category to the position of each sequence, even probabilistically, is sometimes sufficient with profiling using constraint-based peptide identification from databases (as is done with LC-MS data) (Swaminathan et al. 2015 PLoS Comput Biol. 11:e1004080).

[0005] Secondly, proteins are extremely heterogeneous in terms of their physiological and chemical properties. This can be reduced to some extent by applying a bottom-up proteomics approach, i.e., protease-mediated digestion that converts the proteome into peptides. Parallel sequencing of such a complex mixture of peptides with a dynamic range of several orders of magnitude will inevitably face immeasurable technical difficulties. The complexity can be further reduced by purifying subsets of (proteotype) peptides from the peptide pool. Sequencer of peptides passed through nanopores (in solid state) is currently the most popular research platform. Large-scale studies have focused on modifying nanopores that can translocate peptides and distinguish amino acids from each other or from amino acid categories according to their sequence (Kennedy et al. 2016 Nat Nanotechnol 11:968-976; Wilson et al. 2016 Adv Funct Mater 26:4830-4838).

[0006] Another promising technique for sensitive and quantitative protein analysis has been developed by Mitra et al. (WO 2010 / 065531). This technique, called End sequencing or DAPES digital protein analysis, features a method for single-molecule protein analysis. To perform DAPES, a large amount of protein is denatured and cleaved into peptides. These peptides are immobilized on a nanogel surface applied to the surface of a microscope slide, and their amino acid sequences are determined in parallel using a method related to Edman degradation. Phenyl isothiocyanate (PITC) is added to the slide and reacts with the N-terminal amino acid of each peptide to form a stable phenylthiourea derivative. The identity of the N-terminal amino acid derivative is then determined by antibody conjugation, detection, and stripping with antibodies specific to each N-terminal amino acid derivatized with PITC, for example, 20 rounds.

[0007] The N-terminal amino acid is removed by increasing the temperature or decreasing the pH, and this cycle is repeated to sequence 12–20 amino acids from each peptide on a glass slide. The absolute concentration of each protein in the original sample can then be calculated based on the number of various peptide sequences observed. The PITC chemistry used in DAPES is the same as that used for Edman degradation and is efficient and robust (>99% efficiency). However, cleaving a single amino acid requires either a strong anhydride or, alternatively, an aqueous buffer at high temperature. Cycling between any of these harsh conditions is undesirable for performing multiple rounds of analysis on sensitive substrates used for single-molecule protein detection (SMD).

[0008] An alternative peptide sequencing method uses N-terminal amino acid-binding proteins (NAABs) instead of antibodies that bind to N-terminal amino acids derivatized with PITCs (WO20140273004). Such NAABs are developed, which may be modified with aminopeptidases or tRNA synthetases for each amino acid. The NAABs are labeled differently, and the N-terminal amino acids of the polypeptide are then identified by detecting the fluorescent label of the specific NAAB bound to the N-terminal amino acid during incubation and washing of such NAABs. Furthermore, instead of chemical / physical removal of the N-terminal amino acids, an enzyme called Edmanase (named after Edman degradation) may be used.

[0009] While edmanase partially addresses the shortcomings of Mitra et al.'s method, it relies on an arsenal of NAABs to derive amino acid identity information. The need for NAABs for all different amino acids, either coexisting or being added sequentially, adds complexity to the system. Furthermore, the ability to develop NAABs with sufficient affinity for single-molecule detection remains unproven. Consequently, developing simpler and more sophisticated protein sequencing techniques based on various physiological principles, rather than solely on reagent binding affinity, would be advantageous. [Overview of the project]

[0010] overview This invention describes alternative single-molecule peptide sequencing methods and the modified molecules involved. The N-terminal amino acids of the single peptide molecule are identified (or categorized) using the catalytic properties and kinetics of the enzymatic reaction of the aminopeptidase. The method described herein involves the turnover rate (k) of the modified aminopeptidase. cat This is based on the correlation between the aminopeptidase and the N-terminal amino acid it cleaves. Therefore, when a modified aminopeptidase is added, the N-terminal amino acid is identified by measuring the time it is present on the peptide substrate before it is cleaved.

[0011] The aminopeptidase can also be replaced by a chemical cleavage inducer. Similar to what has been observed with the aminopeptidase, the residence time of the chemical cleavage inducer is a read-out of the identity of the N-terminal amino acid to which it binds. More precisely, this application provides a method for sequencing a protein, comprising the following step cycle: N-terminal derivatization of a peptide immobilized through the C-terminal to the easily cleavable peptide moiety, measuring the time it takes for the cleavage inducer to cleave the N-terminal amino acid, leading to release of the N-terminal amino acid from the immobilized surface, and setting the system in preparation for the next cycle (Figure 1). The cleavage inducer may be a catalytically active aminopeptidase or an isothiocyanate-like chemical.

[0012] In a first aspect, a modified catalytically active aminopeptidase acting on a polypeptide is provided, wherein the polypeptide is immobilized on a surface via its C-terminus or via the peptide moiety of the C-terminus to the first peptide bond of the polypeptide, wherein the aminopeptidase cleaves the N-terminal amino acid of the polypeptide, and wherein the residence time of the aminopeptidase for cleaving the N-terminal amino acid identifies or categorizes the N-terminal amino acid. The N-terminal amino acid may be a derivatized N-terminal amino acid, if so, the aminopeptidase binds to and cleaves the derivatized N-terminal amino acid.

[0013] The N-terminal amino acid may be an N-terminal amino acid derivatized with an isothiocyanate or an isothiocyanate analog. More specifically, the aminopeptidase described above has at least 80% sequence identity with SEQ ID NO: 1 or SEQ ID NO: 2 and has a glycine residue at position 25, a serine residue at position 65, a cysteine ​​residue at position 138, and a histidine residue at position 208, and the aminopeptidase can bind to an N-terminal amino acid derivatized with CITC or SPITC. More specifically, the aminopeptidase contains the amino acid sequence as depicted in SEQ ID NO: 3 or SEQ ID NO: 4.

[0014] The aminopeptidase of the present application can also have at least 80% sequence identity to SEQ ID NO: 7, where a cysteine residue is inserted between a methionine residue at position 1 and an alanine residue at position 2. More specifically, the aminopeptidase comprises or consists of SEQ ID NO: 8. In a specific embodiment, the above aminopeptidase further comprises an optical label, an electrical label, or a plasmon label, and thus the aminopeptidase can be detected optically, electrically, or plasmonically. In other specific embodiments, the aminopeptidase is thermophilic and / or solvent resistant.

[0015] In a second aspect, the use of a cleavage inducer for obtaining sequence information of a polypeptide is provided, where the polypeptide is immobilized on a surface via its C-terminus or via a peptide portion of the C-terminal to the first peptide bond of the polypeptide, and where the residence time of the cleavage inducer on the N-terminal amino acid of the polypeptide identifies or categorizes the N-terminal amino acid. The cleavage inducer can be a catalytically active aminopeptidase, isothiocyanate, or an isothiocyanate analog. More specifically, the catalytically active aminopeptidase can be any of the aminopeptidases described in the original application. In a specific embodiment, the N-terminal amino acid is selected from the list consisting of Leu, Met, Tyr, Arg, Pro, Gly, Lys, Ala, and Val. The above use for obtaining sequence information at the single molecule level where the N-terminal amino acid is a derivatized N-terminal amino acid is also provided.

[0016] In a third aspect, there is provided a method for identifying or categorizing the N-terminal amino acid of a polypeptide, where the polypeptide is immobilized on a surface via its C-terminus, or via the peptide portion of the polypeptide from the C-terminus to the first peptide bond, the method comprising contacting the polypeptide immobilized on the surface with a cleavage inducer, where the cleavage inducer binds to the N-terminal amino acid and cleaves it from the polypeptide; measuring the residence time of the cleavage inducer on the N-terminal amino acid; and identifying or categorizing the N-terminal amino acid by comparing the measured residence time with a series of reference residence time values characteristic of the cleavage inducer and a series of N-terminal amino acids.

[0017] In the same line, there is provided a method for obtaining sequence information of a polypeptide immobilized on a surface via its C-terminus, the method comprising: a) contacting the polypeptide immobilized on the surface with a cleavage inducer, where the agent binds to the N-terminal amino acid and cleaves it from the polypeptide; b) measuring the residence time of the cleavage inducer on the N-terminal amino acid of the polypeptide immobilized on the surface; c) enabling the cleavage inducer to cleave off the N-terminal amino acid; d) identifying or categorizing the N-terminal amino acid by comparing the measured residence time with a series of reference residence time values characteristic of the cleavage inducer and a series of N-terminal amino acids; e) repeating steps a) to d) one or more times, or repeating steps b) to d) one or more times comprising

[0018] Also provided is the above method, where the cleavage inducer is an isothiocyanate or an isothiocyanate analog, where the residence time is the time taken until the N-terminal amino acid is removed, and where the N-terminal amino acid is identified by comparing the time taken with a series of reference values for various amino acids.

[0019] Another method provided is the one described above, where the cleavage inducer is an aminopeptidase, and where the residence time of the aminopeptidase is measured for each binding event of the aminopeptidase to the N-terminal amino acid.

[0020] The method of this application may also include the step of determining the cleavage of the N-terminal amino acid by measuring the optical, electrical, or plasmon signal of a polypeptide immobilized on a surface, where the difference in the optical, electrical, or plasmon signal represents the cleavage of the N-terminal amino acid. The above method is also provided, in which the polypeptide immobilized on the surface is further contacted with one or more N-terminal amino acid binding proteins, where the kinetics of the binding event of the one or more binding proteins to the N-terminal amino acid identifies the N-terminal amino acid or provides further information regarding the N-terminal amino acid. The above method may also include a first step of polypeptide denaturation, or a method is provided in which polypeptide denaturation conditions are present during one or more steps of the method, where the catalytically active aminopeptidase is a thermophilic and / or solvent-resistant aminopeptidase, and / or, where the cleavage inducer is an isothiocyanate or isothiocyanate analog. The above method is also provided, where the N-terminal amino acid is derivatized. The aminopeptidase obtained by the above method may be any of the aminopeptidases disclosed herein.

[0021] In specific embodiments, the methods of this application are intended to be used at the single-molecule level. In particular for these applications, it is foreseeable that the residence time of the cleavage inducer will be measured optically, electrically, or plasmonically. This can be done in large quantities when the polypeptide is immobilized on an active sensing surface. The active sensing surface may be a gold surface to which the polypeptide is chemically coupled, or an amide-, carboxyl-, thiol-, or azide-functionalized surface. [Brief explanation of the drawing]

[0022] Simple description of the diagram [Figure 1] Figure 1. Enzymatic degradation, more precisely, kinematic monitoring of N-terminal amino acid cleavage for single-molecule peptide sequencing. [Figure 2] Figure 2. Schematic diagram of optical and potentiometric readouts of enzyme residence time in a non-restrictive case. [Figure 3] Figure 3. Immobilization of Cy5-labeled test peptide (pepCy5) on a glass surface. Left: control; Center: pepCy5 at a concentration of 1 nM; Right: zoomed image showing a successful spatial distribution.

[0023] [Figure 4] Figure 4. Trypsin digestion (1 nM) of surface-immobilized peptide (1 nM pepCy5). A. Trypsin reaction to immobilized pepCy5 in the absence of passivation agent. B. Trypsin treatment of immobilized pepCy5 in the presence of the passivation agent dbco-peg8-amide (1 μM). A signal was detected at 639 nm (upper panel). The background was assessed in the lower λEm channel (Cy3 channel) at 561 nm (lower panel). [Figure 5] Figure 5. Computer-aided good docking of sulfophenyl-isothiocyanate-Ala-Phe(A) and 3-coumarinyl-isothiocyanate-Ala-Phe(B) on virtually modified edmanase. [Figure 6] Figure 6. Successful transformation of T. cruzi (cruzipain) (A) and T. aquaticus (aminopeptidase T) (B) modified in E. coli BL21. [Figure 7] Figure 7. SDS-PAGE analysis of purified T. aquaticus aminopeptidase T. Aquaticus; unprocessed soluble fraction; B. Ni-NTA purification; C. heat treatment; D. Ni-NTA purification + heat treatment.

[0024] [Figure 8] Figure 8. Peptidase assay with L-leucine-p-nitroaniline. E, enzyme; S, substrate. [Figure 9] Figure 9. Molecular map of the pET24b(+) plasmid. [Figure 10] Figure 10. Schematic presentation of the acylation and deacylation steps of aminopeptidase. For the peptide, this case retains XH=RNH2 (thus creating the presence of a peptide bond in the scheme). The N-terminus of the peptide is symbolized by the red portion, while the C-terminus of the peptide is symbolized by X. [Figure 11] Figure 11. Aminopeptidase assays with various amino acid p-nitroanilide substrates. The enzyme and substrate (1.5 mM) were incubated in PBS at 40°C or 80°C for 2 hours, and the released p-nitroanilide was quantified by measuring the absorbance at 405 nm.

[0025] [Figure 12] Figure 12. Organic solvent tolerance of aminopeptidase T from T. aquaticus. A. The activity of aminopeptidase T was measured with an L-leucine-p-nitroanilide substrate. The enzyme and substrate were incubated for 3 hours at 40°C in 50 mM TrisHCl (pH 8) containing various concentrations of organic solvents: acetonitrile (ACN), methanol (MeOH), and ethanol (EtOH). B. Circular dichroism analysis of the secondary structure of aminopeptidase T in 0% vs. 50% MeOH (in 10 mM dipotassium phosphate buffer (K2HPO4)). C. Enzyme activity in MeOH vs. deionized water (MilliQ) at varying concentrations in buffer (50 mM TrisHCl, pH 8). [Figure 13]Figure 13. Site-specific fluorescence labeling of aminopeptidase T. A. SDS-PAGE analysis of aminopeptidase T after labeling of the N-terminal cysteine ​​with maleimide-DyLight650 in isomer and 1×, 10×, 100×, and 1000× molar excess. Aminopeptidase T was visualized by fluorescence (DyLight650 labeling) and Coomassie (total protein). B. Check of aminopeptidase T activity with L-leucine-p-nitroanilide after labeling with maleimide-DyLight650.

[0026] [Figure 14] Figure 14. Monitoring of combinations of enzyme-substrate binding and substrate cleavage events. Single-molecule enzyme residence time is monitored by TIRF microscopy using fluorescently labeled immobilized peptide substrates and free aminopeptidase. [Figure 15] Figure 15. Schematic diagram of the Edman degradation mechanism. Edman degradation inevitably involves the coupling of phenyl isothiocyanate (PITC) to the N-terminus of a free protein / peptide (under alkaline conditions), followed by the release of the N-terminal amino acid as a phenylthiohydantoin (PTH) derivative (under acidic conditions). The released PTH-amino acid is then identified by chromatography. This procedure is then continuously repeated to obtain protein / peptide sequence information (Source: https: / / en.wikipedia.org / wiki / Edman_degradation). [Figure 16] Figure 16. Spontaneous Edman degradation of various amino acid p-nitroanilide substrates. A. Sulfophenyl isothiocyanate (SPITC, 15 mM) and substrate (1.5 mM) were incubated in 300 mM triethanolamine in 50% ACN at 40°C for 30 min, and the released p-nitroanilide was quantified by measuring the absorbance at 405 nm. B. Time-reaction rate (kinetic) measurement of SPITC-derivatized amino acid p-nitroanilide substrate cleavage. [Modes for carrying out the invention]

[0027] Detailed description definition The present invention may be described with reference to certain drawings in relation to specific embodiments, but the present invention is not limited thereto, but is limited only by the claims. No reference numerals in the claims shall be construed as limiting the scope. The drawings provided are schematic only and non-limiting. The sizes of some elements in the drawings may be exaggerated and may not be drawn to scale for illustrative purposes. Where the term “comprising” is used in this description and claims, it does not exclude other elements or steps. Where the indefinite or definite article is used with a singular noun, e.g., “a” or “an” or “the”, it includes the plural form of that noun unless otherwise specifically stated.

[0028] Furthermore, the terms first, second, third, etc., in this description and claims are used to distinguish between similar elements and are not necessarily used to indicate order or chronology. It should be understood that the terms used in this manner are interchangeable under appropriate circumstances, and that the aspects of the invention described herein can be operated in an order other than that described or explained herein. The following terms or definitions are provided solely to aid in the understanding of the invention. Unless specifically defined herein, all terms used herein have the same meaning as to a person skilled in the art of the invention. Practitioners should refer specifically to Michael R. Green and Joseph Sambrook, Molecular Cloning: A Laboratory Manual, 4 for definitions and terms in the art. thThis refers to *Cold Spring Harbor*, ed., *Laboratory Press*, *Plainsview*, New York (2012); and *Ausubel et al.*, *Current Protocols in Molecular Biology* (Supplement 47), John Wiley & Sons, New York (1999). The definitions provided herein should not be construed as being less extensive than those understood by those skilled in the art.

[0029] In the present application, the applicants describe a peptide sequencing method using a multi-step approach in which N-terminal amino acids are identified one by one. More precisely, a method is provided for sequencing a polypeptide, wherein the method comprises the following steps: a) contacting the polypeptide with a cleavage inducer, more specifically a catalytically active aminopeptidase, isothiocyanate, or isothiocyanate analog; b) measuring the residence time of the agent on the N-terminal amino acids of the polypeptide, or alternatively, the k of the enzymatic reaction. cat Measuring the value; c) the residence time or the k cat The method comprises identifying or categorizing the N-terminal amino acid by value; and repeating steps a) to c) one or more times. In one embodiment, the polypeptide is immobilized on a surface. When the agent cleaves the N-terminal amino acid from the polypeptide, it goes without saying that the polypeptide is immobilized on the surface by its C-terminus. Similarly, when the agent cleaves the C-terminal amino acid from the polypeptide, the polypeptide is immobilized on the surface by its N-terminus. In another embodiment, the method is a method for sequencing a polypeptide immobilized on a surface at the single-molecule level.

[0030] Given that the peptide sequencing method described relies on the sequential identification of N-terminal amino acids, the present application also discloses a method for identifying or categorizing the N-terminal amino acids of a polypeptide by determining the residence time of a cleavage inducer (more specifically, a catalytically active aminopeptidase, isothiocyanate, or isothiocyanate analog) on ​​the N-terminal amino acids, wherein the method involves contacting the polypeptide with the agent and determining the residence time of the agent or alternatively the k of the enzymatic reaction. cat This includes measuring values. In this case as well, the method can be used at the single-molecule level, with or without the use of surface-immobilized peptides. Therefore, in one embodiment, the method identifies or categorizes the N-terminal amino acids of a surface-immobilized polypeptide at the single-molecule level.

[0031] As used herein, the terms “peptide” and “polypeptide” are interchangeable and refer to amino acids of any length in polymeric form, which may include coding and non-coding amino acids, natural and unnatural amino acids, chemically or biochemically modified or derivatized amino acids, and polypeptides having a modified peptide backbone. As used herein, “peptide” or “polypeptide” is shorter than a full-length protein from which it is derived and which is formed by, for example, the protein digestion of trypsin or proteinase K (but not limited to these). In specific embodiments, the peptide or polypeptide has an amino acid length between 20 and 500, or between 25 and 200, or between 30 and 100, or has an amino acid length of less than 500, less than 250, less than 200, less than 150, less than 100, or less than 50. In any case, the "peptide" or "polypeptide" contains at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, or at least 20 amino acids.

[0032] "Single-molecule" refers to the investigation of the properties of individual molecules in the manner of a single molecule, at the single-molecule level, or as used in single-molecule experiments. Single-molecule studies are sometimes contrasted with measurements of molecular aggregates or bulk collections, where the individual behavior of molecules cannot be identified and only average characteristics can be measured.

[0033] Proteins are polymers of amino acids. Proteins are made by ribosomes, which "read" RNA encoded by codons in genes and assemble the essential amino acid combinations from the genetic instructions in a process known as translation. The newly formed protein strand then undergoes post-translational modification, in which additional atoms or molecules, such as copper, zinc, or iron, are added. Once this post-translational modification process is complete, the protein begins to fold (sometimes spontaneously, sometimes with the help of enzymes) so that the hydrophobic elements of the protein are buried deep within its internal structure and the hydrophilic elements end up on the outside, curling itself up. The final shape or structure of a protein is determined by how it interacts with its environment. Therefore, proteins have a primary structure (i.e., a sequence of amino acids held together by covalent peptide bonds), a secondary structure (i.e., regular repeating patterns such as alpha-helices and beta-pleated sheets), a tertiary structure (i.e., covalent interactions between amino acid side chains, such as disulfide bridges between cysteine ​​groups), and a quaternary structure (i.e., protein subunits that interact with each other).

[0034] However, for the methods disclosed in this application, the protein and its N-terminal amino acids should be easily accessible to the aminopeptidase of this application, and preferably, the protein is immobilized in a linear configuration. Therefore, in various embodiments, the protein to be sequenced should be denatured. Denaturation is the process by which a protein loses its quaternary, tertiary, and secondary structures in which it exists in its native state, but the peptide bonds of the primary structure between amino acids remain intact. Protein denaturation can be achieved by applying external stress or compounds such as strong acids or bases, concentrated inorganic salts, organic solvents (e.g., alcohol or chloroform), radiation, or heat. Therefore, the aminopeptidase to be used in various embodiments of this application is thermophilic and / or solvent-resistant (see below).

[0035] When used herein, “thermophilic” refers more precisely to an organism or enzyme that is highly active or maintains its activity at relatively high temperatures, particularly between 40°C and 122°C, rather than having “enhanced temperature tolerance.” An unspecified example of a thermophilic organism is Thermus aquaticus, whose enzymes, such as aminopeptidase T, function at high temperatures and are therefore thermophilic. In specific embodiments, the aminopeptidases for use and methods of the present application have optimal peptidase activity in temperature ranges of 40°C and 100°C, or 40°C and 80°C, or 50°C and 70°C, or 60°C and 80°C. In other specific embodiments, the aminopeptidases of this application maintain their enzymatic activity in the presence of solvents such as acetic acid, trichloroacetic acid, sulfosalicylic acid, sodium bicarbonate, ethanol, and alcohol; crosslinking agents such as formaldehyde and glutaraldehyde; chaotropic agents such as urea, guanidine chloride, or lithium perchlorate; and disulfide bond-breaking agents such as 2-mercaptoethanol, dithiothreitol, or tris(2-carboxyethyl)phosphine.

[0036] The use of thermophilic and / or solvent-tolerant aminopeptidases is particularly useful for fine-tuning the "on-time" value of the aminopeptidase when bound to various N-terminal amino acids. By changing the reaction conditions during the experiment (e.g., protein sequencing), temperature, pH, solvent, etc., can be adjusted to better distinguish between the "on-time" values ​​of amino acid X and amino acid Y.

[0037] In the method of the present application, the N-terminal amino acid is cleaved from the polypeptide substrate. This can be achieved enzymatically, for example chemically, or by a peptidase, more specifically an aminopeptidase. The cleavage inducer may be covalently or noncovalently bound to the N-terminal amino acid.

[0038] Chemical cleavage inducers Edman degradation is a chemical technique that enables the determination of the N-terminal sequence of proteins. It was first described by Pehr Edman in 1950, and the degradation reaction was fully automated in 1967. This method inevitably involves the coupling of phenyl isothiocyanate (PITC) to the N-terminus of the protein / peptide (under alkaline conditions), followed by the release of the N-terminal amino acid as a phenylthiohydantoin (PTH) derivative (under acidic conditions) (Figure 15). The released PTH-amino acid is then identified, for example, by chromatography. This procedure is then continuously repeated to obtain the protein / peptide sequence information.

[0039] In the present application, it was surprisingly found that PTH is released even under non-acidic conditions or without heating, and thus the N-terminal amino acid is cleaved from the polypeptide substrate. Even more surprisingly, the inventors found that the time from when the isothiocyanate (or analogue) binds to the N-terminal amino acid until the N-terminal amino acid is cleaved depends on the characteristics of the amino acid. Therefore, the N-terminal amino acid can be identified by measuring this time, which is essentially the residence time of the isothiocyanate (or analogue) on the N-terminal amino acid.

[0040] In a preferred embodiment of this application, the cleavage inducer referred to in the use and method of this application is a chemical agent, more specifically, a chemical agent selected from the list consisting of isothiocyanates (ITC), phenyl isothiocyanates (PITC), azido-PITC, coumarinyl isothiocyanates (CITC), and sulfophenyl isothiocyanates (SPITC). In this application, PITC, azido-PITC, CITC, and SPITC will be referred to as isothiocyanate analogs. Thus, the method of this application is provided, in which the cleavage inducer is an isothiocyanate or an isothiocyanate analog.

[0041] Also provided herein is the use of a cleavage inducer to obtain sequence information of a polypeptide immobilized on a surface via its C-terminus, where the residence time of the cleavage inducer on the N-terminal amino acid of the polypeptide identifies the N-terminal amino acid. In a specific embodiment, the cleavage inducer is an isothiocyanate or isothiocyanate analog, and the residence time is the time it takes from the binding of the ITC or ITC analog to the N-terminal amino acid until the N-terminal amino acid is removed. The N-terminal amino acid can then be identified by comparing the residence time with a series of reference values ​​obtained for ITCs or ITC analogs related to various amino acids.

[0042] Aminopeptidase As used herein, "aminopeptidase" refers to an enzyme that catalyzes the cleavage of amino acids from the amino-terminus (N-terminus) of a protein or peptide substrate. They are widely distributed across the plant and animal kingdoms and are found in many intracellular organelles, cytosols, and as membrane components. Aminopeptidases are classified by 1) the number of amino acids cleaved from the amino-terminus of the substrate (for example, aminodipeptidases remove intact dipeptides from the amino-terminus, and aminotripeptidases catalyze the hydrolysis of tripeptides from the amino-terminus), 2) the location of the aminopeptidase in the cell, 3) its sensitivity to inhibition by bestatin, 4) its metal ion content and / or the metal-binding residues of the enzyme, 5) the pH at which maximum activity is observed, and 6) the relative efficiency of residue removal, which is most relevant to this application (Taylor 1993 FASEB J 7:290-298). Aminopeptidases may have broad or narrow substrate specificity. While this application focuses on the development or use of aminopeptidases with broad substrate specificity, the use of multiple aminopeptidases with overlapping or complementary substrate specificity is also envisioned in this application.

[0043] Generally, the specificity of an enzyme for a specific substrate under specific environmental conditions is given by the specificity constant k cat / K Mcan be quantified by k cat is the metabolic turnover number, the number of substrate molecules that each enzyme site converts to product per unit time, or the number of substrates with the ability to produce reactants per catalytic center and per unit time. K M is defined as the substrate concentration required for an enzyme to reach half of its maximum velocity under the conditions required for the measurement of a proper steady-state enzyme kinetics, which is well-known in the art. When distinguishing between two enzyme substrates A and B, based on the conversion rates of these substrates to products, this type of relationship is as follows: [Number] is maintained, where v is the velocity and [A] is the concentration of A.

[0044] As a result, information regarding the identity of the various substrates of an enzyme can be obtained from measurements of the conversion rates of these substrates by the enzyme. The relative velocities under conditions of equal substrate concentrations are determined by k cat and K M When observing a single substrate molecule, once the enzyme is added, the time required to form a product molecule is governed by k cat Therefore, in the observations of single molecules, information regarding the identity of the substrate can be obtained from the "on-time" or residence time of the enzyme on the substrate. This information can be further complemented by modifying the substrate and / or the enzyme so that it can be distinguished from unproductive matches between the catalytically productive match of the enzyme and the substrate. Thus, "on-time", as used herein, refers to the residence time of the enzyme on the substrate, the contact time of the enzyme solution with the substrate, or more specifically, the reciprocal of k cat which is well-known in the art. Hereafter, "on-time" and residence time shall be used interchangeably and can refer to the time until one enzyme molecule acts on one peptide molecule to cause cleavage, or the time required until multiple enzyme molecules act successively on a peptide molecule to cause cleavage.

[0045] The finding that the "on-time" of an enzyme on a substrate can be used to identify that substrate is particularly true for aminopeptidases. Peptidases generally operate through a two-step mechanism (Figure 10). Firstly, during the acylation reaction, the N-terminal portion of the peptide (for aminopeptidases) or the C-terminal portion of the peptide (for carboxypeptidases) is cleaved and covalently linked to the peptidase. Secondly, in the deacylation reaction, the enzyme releases the cleaved amino acids.

[0046] Aminopeptidases gain their specificity for specific groups of amino acids through stereoelectronic fit with the transition state of the acylation reaction, which depends on the nature of the substrate's side chain (one or more) and the easily cleavable bond, particularly at the N-terminus. Typically, the aminopeptidase has a much weaker binding interaction with the C-terminus of the peptide portion and the easily cleavable bond, and will therefore rapidly dissociate from the peptide (or the surface to which the peptide was bound) during the acylation or hydrolysis step that determines the reaction rate. If the peptide is immobilized at the C-terminus from the easily cleavable peptide bond that is cleaved by the peptidase, then during the acylation reaction, the N-terminal amino acid or amino acid derivative of the peptide will be covalently linked to the enzyme in the case of serine or cysteine ​​peptidases, or acovalently linked to the enzyme in the case of peptidases that directly hydrolyze, while the C-terminus will remain conjugated to the surface to which the peptide was immobilized (Figure 10).

[0047] Consequently, for a selected aminopeptidase, the residence time or "on-time" on the surface-immobilized peptide substrate correlates with the rate of the acylation or hydrolysis step, and therefore with the nature of the N-terminal—easily cleavable—bond portion. The "on-time" of an aminopeptidase can be readily determined in this case by molecular labeling the aminopeptidase. Thus, molecular labeling acts as a substitute for the "on-time" of the aminopeptidase, and as a substitute for the identity of the N-terminal amino acid cleaved by the aminopeptidase. In specific embodiments of this application, the aminopeptidase may be labeled optically, fluorescently, electrically, or plasmonically (see below).

[0048] In an alternative embodiment, a solution of aminopeptidase molecules is brought into contact with a peptide substrate, and the residence time / on-time is measured until the N-terminal amino acid (or its derivative) is cleaved. In such embodiment, the total residence time of the enzyme in contact with the substrate is measured up to such cleavage event, and this value is the k of the enzyme relating to the specific N-terminal amino acid (derivative) on the peptide substrate under the conditions of use. cat It correlates with the reciprocal of [the given expression].

[0049] The situation is different for carboxypeptidases derived from a group of cysteine ​​and serine proteases. More precisely, in the case of these carboxypeptidases, the enzyme remains covalently bound to the immobilized peptide portion after cleaving the C-terminal amino acid. The carboxypeptidase will not dissociate from the peptide during the acylation step, and its "on-time" value on the peptide on the immobilized surface will be determined by the rate of the deacylation (hydrolysis) step. The latter hydrolysis step has little to no informational value to the properties of the C-terminal amino acid (which has already been released into the solvent during the acylation step). However, in embodiments in which a solution of carboxypeptidase molecules is brought into contact with a peptide substrate and the residence time / on-time is measured until the C-terminal amino acid (or its derivative) is cleaved, this value is the k of the enzyme relating to the specific C-terminal amino acid (derivative) on the peptide substrate under the conditions of use. cat Correlating with the reciprocal of such carboxypeptidase, such carboxypeptidase can be used within the scope of the present invention.

[0050] Interestingly, carboxypeptidases derived from a group of metalloproteinases neither constitute this covalent bond nor cleave the C-terminal amino acid by hydrolysis. Therefore, the "on-time" nature of the carboxy-metallopeptidase is informative to the C-terminal amino acid it binds to and cleaves, just as it is to the N-terminal amino acid with aminopeptidases. Thus, the use of carboxy-metallopeptidases is also envisioned in the methods of the present application, although there is a significant difference in that the polypeptide is subsequently immobilized on the surface through its N-terminus or through the peptide's side chains. In short, in addition to the practicality of isothiocyanates and / or ITC analogs, aminopeptidases or carboxy-metallopeptidases are also particularly useful in the methods disclosed in this application and herein.

[0051] Therefore, in various specific embodiments of the present application, the cleavage inducer is a peptidase, specifically an aminopeptidase or carboxy-metallopeptidase, more specifically an aminopeptidase, as described in the methods and uses of the present application.

[0052] In specific embodiments of this application, the use of an active peptidase is provided, the cleavage rate or kinetics of its peptidase activity being characterized by the amino acid substrate of the polypeptide, more specifically by the terminal amino acids, and thus identifying these amino acids. A desired strategy is to utilize an aminopeptidase, specifically a single unique aminopeptidase, more specifically a catalytically active aminopeptidase that recognizes each of 20 viable N-terminal amino acids. However, it is also conceivable that two, three, four, or more aminopeptidases may be used that subsequently identify different groups of amino acids (for example, non-aromatic amino acids to aromatic amino acids, or hydrophobic terminal amino acids, positively charged amino acids, negatively charged amino acids, and small amino acids, but not limited to these).

[0053] It is also assumed that the "on-time" value of an aminopeptidase for a particular N-terminal amino acid may change by altering the reaction conditions during an experiment (e.g., protein sequencing). This is particularly desirable when the aminopeptidases used have very similar "on-time" values ​​for a given N-terminal amino acid. Therefore, in one embodiment, reaction conditions, including temperature, pH, and solvent, among many others, are adjusted to increase the distinction between the "on-time" value for amino acid X and the "on-time" value for amino acid Y. In another embodiment, the aminopeptidase itself is modified to distinguish between various amino acids that would otherwise have similar residence times in a native aminopeptidase. When used herein, "modified" is synonymous with "synthetic," "recombinant," "man-made," or "non-natural."

[0054] As a non-limiting example, aminopeptidase T derived from T. aquaticus may be used in the method of the present application (see below). Residues that are viable for modifying the aminopeptidase are those within a radius of 8 angstroms from the divalent metal ion in the catalytic site, more precisely, E250, F252, G315, E316, V317, A318, T336, E340, H345, I346, A347, F348, Q350, Y352, N355, H376, V377, D378, and / or W379. The positions of these residues refer to their positions in wild-type aminopeptidase T as depicted in Sequence ID No. 7.

[0055] A desired alternative strategy is that, when the aminopeptidase described in the present application binds, the enzyme "on-time" value can be used to identify the N-terminal amino acids of the immobilized polypeptide. Also envisioned in this application are aminopeptidases, more specifically catalytically active aminopeptidases, their uses and methods, where the aminopeptidase is used, and its enzyme "on-time" value is beneficial or informative for a group or subgroup of amino acids. Thus, in some embodiments, the enzyme "on-time" value will classify or categorize the N-terminal amino acids in a group or subgroup of amino acids with a certain probability.

[0056] In one embodiment, the aminopeptidase used in the method disclosed in this application is an aminodipeptidase, more specifically a catalytically active aminodipeptidase. Aminodipeptidase is synonymous with diaminopeptidase and refers to an enzyme that cleaves the two N-terminal amino acids of a polypeptide.

[0057] In a preferred embodiment, the aminopeptidase or aminopeptidase used in the method disclosed in the present application is catalytically active. "Catalytically active" means that the aminopeptidase is a fully functional catalytic enzyme. This is in contrast to a catalytically non-functional (dead) aminopeptidase, such as in WO20140273004, which is modified to bind to the N-terminal amino acid but not cleave it.

[0058] Thermophilic aminopeptidases are particularly intended for use and methods described in the present application. Non-limiting examples of such aminopeptidases that may be used in the methods described in the present application include aminopeptidase T derived from Thermus aquaticus (AMPT_THEAQ), aminopeptidase T derived from Thermus thermophilus (AMPT_THET8), PepC derived from Streptococcus thermophiles (PEPC_STRTR), aminopeptidase S derived from Streptomyces griseus (APX_STRGG), aminopeptidase TH-2 derived from Streptomyces septatus TH-2 (Q75V72_9ACTN), and aminopeptidase 2 derived from Bacillus stearothermophilus (AMP2_GEOSE).

[0059] Non-limiting examples of catalytically active aminopeptidases envisioned in the present application and disclosed herein include modified Trypanosoma cruzi (cruzain) and Thermus aquaticus aminopeptidase T.

[0060] In the present application, a modified catalytically active aminopeptidase is provided, comprising a binding domain to any N-terminal amino acid of a polypeptide or to a different set of N-terminal amino acids of a polypeptide, wherein the polypeptide is immobilized on a surface, wherein the aminopeptidase cleaves the N-terminal amino acid upon binding, and wherein the "on-time" of the aminopeptidase enzyme identifies or categorizes the N-terminal amino acid. Also provided herein is a modified catalytically active aminopeptidase that binds to a polypeptide immobilized on a surface, wherein the aminopeptidase cleaves the N-terminal amino acid of the polypeptide, and wherein the residence time of the aminopeptidase on the N-terminal amino acid identifies or categorizes the N-terminal amino acid.

[0061] In a specific embodiment, the polypeptide is immobilized on the surface through a peptide C-terminus—a readily cleavable bond. In another specific embodiment, the N-terminal amino acid is a derivatized N-terminal amino acid, and the aminopeptidase binds to and cleaves the derivatized N-terminal amino acid. In a more specific embodiment, the derivatized N-terminal amino acid is an N-terminal amino acid derivatized with ITC, CITC, SPITC, PITC, azide-PITC, or a click chemistry-modified product of azide-PITC (hereinafter collectively referred to as "azide-PITC"), and the aminopeptidase binds to and cleaves each of the N-terminal amino acids derivatized with ITC, CITC, SPITC, PITC, or azide-PITC.

[0062] The term "derivative" originates from the chemistry technique or biochemical mechanism of transforming a chemical compound into a derivative, which is a product of a similar chemical structure (a derivative of a reaction). Generally, specific functional groups of a compound are added to a derivatization reaction, transforming the product into a derivative that deviates from its reactivity, solubility, boiling point, melting point, aggregation state, chemical composition, interactions, or optical, electrical, or plasmonic characteristics. In alternative embodiments, "derivative" means labeled.

[0063] In another embodiment, the N-terminal amino acid is selected from a list consisting of Leu, Met, Tyr, Arg, Pro, Gly, Lys, Ala, and Val, and the binding domain to a different set of N-terminal amino acids, or to any N-terminal amino acid, is a binding domain to an N-terminal acid selected from a list consisting of Leu, Met, Tyr, Arg, Pro, Gly, Lys, Ala, and Val. In yet another embodiment, the N-terminal amino acid is Leu, Met, Tyr, Arg, Pro, Gly, Lys, Ala, or Val, and the binding domain to any N-terminal amino acid is a binding domain to Leu, Met, Tyr, Arg, Pro, Gly, Lys, Ala, or Val.

[0064] In another specific embodiment, a modified synthetic or recombinant aminopeptidase is provided comprising a binding domain to the N-terminal amino acid of one or more different derivatized or labeled polypeptides, wherein the aminopeptidase cleaves the derivatized or labeled N-terminal amino acid upon binding, and the cleavage rate of the aminopeptidase or kinetics of the aminopeptidase activity identifies the N-terminal amino acid. In one embodiment, the derivatized or labeled N-terminal amino acid is an N-terminal amino acid derivatized or labeled with ITC, CITC, SPITC, PITC, or azide-PITC. In another embodiment, the derivatized or labeled N-terminal amino acid is a derivatized or labeled Leu, Met, Tyr, Arg, Pro, Gly, Lys, Ala, or Val. Thus, the aminopeptidase is catalytically active and not a non-catalyst aminopeptidase.

[0065] In one aspect, this application provides a modified synthetic or recombinant aminopeptidase comprising the amino acid sequence of wild-type Trypanosoma cruzi, or cruzain, having a glycine residue at position 25, a serine residue at position 65, a cysteine ​​residue at position 138, and a histidine residue at position 208, wherein the remaining amino acid sequence of the aminopeptidase is that of the wild-type T. as depicted in Sequence ID No. 1. The sequence contains at least 75%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity with respect to the amino acid sequence of cruzipain or cruzain as depicted in Sequence ID No. 2.

[0066] This is equivalent to providing a modified aminopeptidase having at least 75%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity with respect to SEQ ID NO: 1 or SEQ ID NO: 2, as well as a modified aminopeptidase having a glycine residue at position 25, a serine residue at position 65, a cysteine ​​residue at position 138, and a histidine residue at position 208. In one embodiment, the cysteine ​​residue is inserted after the first methionine residue of the aminopeptidase. The cysteine ​​is used to label the modified cruzipain or cruzain aminopeptidase. In a specific embodiment, the aminopeptidase comprises the amino acid sequence as depicted in SEQ ID NO: 3 or SEQ ID NO: 4. In a more specific embodiment, the aminopeptidase consists of the amino acid sequence as depicted in SEQ ID NO: 5 or SEQ ID NO: 6. In a specific embodiment, the aminopeptidase in the above aspect and its embodiments is a catalytically active aminopeptidase.

[0067] The specific mutations described above in T. cruzi, cruzipain and cruzain, direct their activity (binding and cleavage) towards a derivatized N-terminal amino acid, more specifically, an N-terminal amino acid derivatized with CITC or SPITC. Therefore, this method or present application is also useful for sequencing peptides containing a derivatized amino acid or for identifying a derivatized N-terminal amino acid. In specific embodiments, the derivatized amino acid is an amino acid derivatized with ITC, PITC, azide-PITC, CITC, and / or SPITC, more specifically, an amino acid derivatized with CITC and / or SPITC.

[0068] In addition to the utility of the modified T. cruzi crujipain and cruzin described above for cleaving N-terminal amino acids derived at CITC or SPITC in the method of the present application, the modified T. cruzi crujipain and cruzin can also be used in the method of WO20140273004. In the latter document, the identification of the derived N-terminal amino acids is made by a series of N-terminal amino acid binding proteins, and as soon as the identified N-terminal amino acids are removed by edmanase. In cases where the N-terminal amino acids are derived at CITC or SPITC, the modified T. cruzi crujipain and cruzin described in the present application can be used as edmanase.

[0069] In the following aspect, a modified synthetic or recombinant aminopeptidase is provided which has at least 75%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity with respect to the amino acid sequence of wild-type Thermus aquaticus aminopeptidase T as depicted in Sequence ID No. 7, wherein the cysteine ​​residue is inserted after the first methionine residue of the wild-type aminopeptidase. This provides a modified aminopeptidase having at least 75%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity with respect to SEQ ID NO: 7, where the cysteine ​​residue is inserted between the methionine residue at position 1 and the alanine residue at position 2. In a specific embodiment, the aminopeptidase T comprises the amino acid sequence as depicted in SEQ ID NO: 8. In a more specific embodiment, the aminopeptidase consists of the amino acid sequence as depicted in SEQ ID NO: 8. In a specific embodiment, the aminopeptidase of the above aspect and its embodiments is a catalytically active aminopeptidase.

[0070] As used herein, the terms “identical,” “similarity,” or “percent “identity,” “percent “similarity,” or “percent “homology” refer to two or more polypeptide sequences or subsequences that, when compared and sequenced to be as identical as possible across a comparison window or specified region, as measured by a sequence comparison algorithm or by manual alignment and visual inspection, have the same (e.g., 75% identity across a specified region) amino acid residues or a specified percentage. Preferably, identity exists over a region of at least about 25 amino acids, more preferably over a region of 50 to 100 amino acids, and even more preferably over a region of 100 to 500 amino acids or even longer.

[0071] The terms “sequence identity” or “sequence homology,” as used herein, refer to the degree to which sequences are identical on an amino acid-by-amino acid basis across a comparison window. Thus, the “percentage of sequence homology” is calculated by comparing two optimally aligned sequences across a comparison window, determining the number of positions where identical amino acids occur in both sequences to obtain the number of matched positions, dividing the number of matched positions by the total number of positions in the comparison window (i.e., the window size), and multiplying the result by 100 to obtain the percentage of sequence identity. Gaps, i.e., positions in alignment where a residue is present in one sequence but not in the other, are considered positions with non-identical residues. Determining the percentage of sequence homology is done manually or by utilizing computer programs available in the art. Examples of useful algorithms include PILEUP (Higgins & Sharp, CABIOS 5:151 (1989)), BLAST, and BLAST 2.0 (Altschul et al. J. Mol. Biol. 215: 403 (1990)). Software for performing BLAST analysis is publicly available through the National Center for Biotechnology Information (http: / / www.ncbi.nlm.nih.gov / ). In a specific embodiment, the comparison window for determining the sequence identity of two or more polypeptides (such as aminopeptidases) is the full-length protein sequence.

[0072] Immobilization and labeling When used herein, "immobilization on a surface" refers to the attachment of one or more polypeptides to an inert, insoluble material, such as a glass surface, resulting in the loss of the polypeptide's mobility. In the methods disclosed in this application, immobilization ensures that the polypeptide(s) remain in place throughout the process, through sequencing of the polypeptide or identification or categorization of its N-terminal amino acids. Therefore, the N-terminus should be freely accessible, and thus the polypeptide should be immobilized through its C-terminus. Furthermore, high-density immobilization of proteins on a surface allows for the use of small amounts of sample solution. Many immobilization techniques have been developed in recent years, primarily based on three mechanisms: physical, covalent, and biocompatible immobilization (Rusmini et al 2007 Biomacromolecules 8: 1775-1789; U.S. Patent No. 6,475,809; WO2001040310; US7358096; US20100015635; WO1996030409). In this application, polypeptides are immobilized on a glass surface using an azido-dibenzocyclooctyl (DBCO) click reaction (see Examples 1 and 2) according to protocols available in the art, e.g., Eeftens et al. (2015 BMC Biophys 8:9).

[0073] In various embodiments of the present application, polypeptides may be immobilized on a surface prior to contact with an aminopeptidase. The peptides may be immobilized on any preferred surface (see below). Essential to the methods disclosed in the present application is that the polypeptide to be sequenced or the polypeptide whose N-terminal amino acid is to be identified or categorized is immobilized through the most C-terminal portion of the polypeptide or through the C-terminal portion of a cleavable bond. Thus, the polypeptide is attached to the surface of the present application at its C-terminus or at the C-terminal to cleavable bond portion along the peptide structure (through methods well known in the art, e.g., cysteine ​​thiol function, e.g., maleimide chemistry or gold-thiol bonding).

[0074] When used herein, "easily cleavable bond" refers to a covalent chemical bond that is intended to be cleaved by any one of the aminopeptidases of this application.

[0075] As used herein, “surface” is synonymous with a carrier or layer. Surfaces or layers of the present application are suitable for use in the detection of molecular labels, electrochemical signals, electromagnetic signals, and plasmon-related events. The molecular labels may be optical labels (including, but not limited to, luminescent and fluorescent labels) or electrical labels (including, but not limited to, potentiometric, voltammetric, and coulometric labels).

[0076] The layer may also be multilayer, i.e., a layer comprising several layers. In the multilayer case, at least one layer should enable suitable detection of the molecular label or the electrochemical, electromagnetic, or plasmon-related event. Thus, according to a specific embodiment, the surface is an active sensing surface. Hence, the polypeptide immobilized on the surface of the method for arranging polypeptides immobilized on a surface at the single-molecule level is a polypeptide immobilized on an active sensing surface. In a more specific embodiment, the active sensing surface is either a gold surface to which the polypeptide of the method is chemically coupled, or an amide-, carboxyl-, thiol-, or azide-functionalized surface. In other specific embodiments, the support is a nanoparticle, a nanodisk, a nanostructure, or a chip. In the most specific embodiment, the surface is a self-assembled monolayer (SAM).

[0077] To detect the "on-time" value or residence time, two labeling options may be selected. First, the polypeptides to be sequenced may be labeled, for example, through their N-terminal amino acids. Alternatively, or in addition, the inner amino acids may be labeled, for example, as shown in Figure 14.

[0078] Polypeptide labeling can be performed using fluorescent probes such as fluorescein, o-phthalaldehyde, dansyl chloride, and coumarinyl isothiocyanate (CITC), but is not limited to these. In specific embodiments, the N-terminal amino acids of the immobilized polypeptides of this application are derivatized with CITC, or, to put it another way, labeled with CITC. Polypeptides can also be electrically labeled. Electroanalytic methods are a type of technique in which the presence of an analyte, peptide, enzyme, etc., can be determined by measuring the potential (volts) and / or current (amperes) of an electrical label on the analyte, peptide, enzyme, etc. These methods can be broken down into several categories depending on the labeling. The three main categories are potentiometrics (where the electrode potential difference is measured), coulometrics (where the current is measured over time), and voltammetry (where the current is measured while the potential is actively changed).

[0079] There are two basic categories of coulometric techniques. Potential coulometry involves maintaining a constant electric potential during a reaction using a potentiostat. On the other hand, so-called coulometric titration or amperostatic coulometry involves maintaining a constant current (measured in amperes) using an amperostat. A non-limiting example of electrical labeling is sulfophenyl isothiocyanate (SPITC). SPITC is a negatively charged variant of the phenyl isothiocyanate (PITC) probe used in MS de novo peptide sequencing to neutralize ions of the N-terminal fragment (Samyn et al. 2004 J Am Soc Mass Spectrom 15:1838-1852). In specific embodiments, electrical labeling may be potentiometric, amperometric, or voltammetric labeling.

[0080] To detect and measure the “on-time” value or residence time of the aminopeptidase of this application on the N-terminal amino acid of an immobilized polypeptide or until the immobilized polypeptide is cleaved (see above), the aminopeptidase must be detectable. The aminopeptidase may interact with the substrate within the measured residence time until the N-terminal amino acid is cleaved, either cleavage-productively or cleavage-non-productively. Of both interaction types, their length, sum of lengths, and average length may be part of the measured residence time relevant to this invention, as all these parameters are part of the measurement that provide information about how long it takes for the aminopeptidase to cleave the N-terminal amino acid. Any type of detection is not essential to this invention, as long as the enzyme “on-time” or residence time of the aminopeptidase can be detected. In some embodiments of this application, the “on-time” of the aminopeptidase is detected optically, electrically, or plasmonically. One method for detecting aminopeptidase in this application involves fusing it with a molecular label and subsequently detecting the molecular label. As described above, aminopeptidase can be labeled optically, electrically, or plasmonically.

[0081] Optical detection requires optical labeling and includes, but is not limited to, luminescence detection and fluorescence detection. Therefore, the label may be a fluorophore. A broad catalog of commercially available optical labels exists, but is not limited to, Cy3, Cy5, coumarin, Alexa fluor labeling, GFP, YFP, RFP, etc. In one embodiment, the aminopeptidase labeling interacts with the labeling of an immobilized polypeptide, or the N-terminal amino acids of the polypeptide, or the immobilized surface. The label may, for example, be a fluorophore. In another embodiment, at least one molecule shares a common base between the first and second groups of the labeled molecule. In one embodiment, the detection step produces an image, for example, a fluorescence image (e.g., acquired using fluorescence resonance energy transfer (FRET), total internal reflection fluorescence (TIRF), or zero-mode waveguide (ZMW)). In another embodiment, image editing constitutes a digital profile, for example, a digital profile identifying the immobilized polypeptide or its N-terminal amino acids. In a specific embodiment, optical labeling is fluorescent labeling. In an even more specific embodiment, the fluorescent labeling is measured or detected by TIRF microscopy.

[0082] The binding of aminopeptidases to the N-terminal amino acids of immobilized proteins, and consequently the residence time or "on-time" measurement of aminopeptidases, can also be detected even without molecular labeling. An unrestricted example of label-free electrical detection is the use of field-effect transistor-based biosensors or BioFETs. A BioFET is a gated field-effect transistor that opens and closes by a change in surface potential induced by molecular binding (i.e., aminopeptidase binding to the N-terminal amino acids of an immobilized polypeptide). When charged molecules, such as SPITC-labeled aminopeptidases, which are usually dielectric materials, bind to the FET gate, they can alter the charge distribution of the underlying semiconductor material, resulting in a change in the conductance of the FET channel. A BioFET consists of two main compartments: one is the biological recognition element, and the other is the field-effect transistor (FET).

[0083] BioFETs can be constructed from ion-sensitive field-effect transistors (ISFETs), silicon nanowires (SiNWs), EIS capacitive sensors, and light-addressable potentiometers (LAPS) simply by modifying the gate or coupling it with various biologically recognizing elements (receptors) (Poghossian and Schoening 2014 Electroanalysis 26: 1197-1213). These encompass either a complex array of biomolecular species (e.g., enzymes, antibodies, antigens, proteins, peptides, or DNA) or living biological systems (e.g., cells, tissue sections, intact organs, or entire organisms). In this application, the entire biologically recognizing system selectively recognizes aminopeptidase-polypeptide bonds, ITC-polypeptide bonds, or ITC analog-polypeptide bonds to be detected, translating (bio-)chemical information into chemical or physical signals. The most crucial aspect of information transfer from biological recognition to the transducer is the interface between these two domains. In BioFED, the potential effect (or charge effect) is used to transform these recognition phenomena.

[0084] Generally, BioFEDs are highly sensitive to any kind of charge or potential change generated by intermolecular interactions at or near the gate insulator / electrolyte interface. Binding of a charged species to the gate insulator is analogous to applying an additional voltage to the gate. Therefore, it can be expected that the adsorption or binding of charged biomolecules on the gate surface will modulate the silicon space charge region at the insulator / semiconductor interface. This results in modulation of the drain current of an ISFET, the conductance or current of a SiNWFET, the capacitance of an EIS sensor, or the photocurrent of a LAPS. Consequently, by measuring changes in the FET drain current, the SiNW conductance, the EIS sensor capacitance, or the photocurrent of a LAPS, the aminopeptidase, ITC, or ITC analogue "on-time" values ​​can be quantitatively determined. In various embodiments, detection of aminopeptidase binding to immobilized polypeptides is carried out using techniques relating to BioFETs or field-effect transistors.

[0085] Furthermore, plasmonic readout may also be used to detect the "on-time" of the aminopeptidase or ITC or ITC analogues of this application. In physics, plasmons can be defined as quanta relating to collective oscillations of free electrons, usually at the interface between a (noble) metal and a dielectric. The term plasmon refers to the plasma-like behavior of free electrons in a metal under the influence of electromagnetic radiation. Surface plasmons are coherent delocalized electron oscillations that exist at the interface between any two materials where the real part of the dielectric function changes sign across the interface (e.g., a metal-dielectric interface such as a metal sheet in air). Excitation of surface plasmons, which can be done very efficiently with light in the visible range of the electromagnetic spectrum, is frequently used in an experimental technique known as surface plasmon resonance (SPR).

[0086] In SPR, the maximum excitation of surface plasmons is detected by monitoring the reflected power from the prism coupler as a function of the angle of incidence or wavelength. This technique can be used to observe changes in thickness at the nanometer level, density fluctuations, or molecular absorption, and to screen and quantify protein-binding events. Commercial instruments operating according to these principles are available. Thus, in specific embodiments, the "on-time" of the aminopeptidase or ITC or ITC analogue of this application is determined by surface plasmon resonance.

[0087] Another example where the "on-time" of aminopeptidases or ITCs or ITC analogs can be measured is a plasmon-enhanced whispering gallery microcavity sensor, as demonstrated for polymerase-DNA interactions (Kim et al 2017 Sci Adv 3:e1603044).

[0088] However, even with the above description, it is clear that labeling and the resulting detection are not essential to the present invention, as long as the "on-time" or residence time of the cutting inducer can be detected.

[0089] Use of cutting-inducing agents In another aspect of the present application, the use of a cleavage inducer to obtain sequence information of a polypeptide immobilized on a surface is provided, where the residence time of the cleavage inducer on the terminal amino acids of the polypeptide identifies the terminal amino acids. In one embodiment, the terminal amino acids are derivatized. In another embodiment, the cleavage inducer is an isothiocyanate or an isothiocyanate analog, or a peptidase. More specifically, the isothiocyanate analog is selected from the list consisting of ITC, CITC, PITC, CITC, and azido-PITC. The peptidase is a catalytically active peptidase, more specifically a catalytically active aminopeptidase.

[0090] In another embodiment, the use of any of the aminopeptidases described in this application for cleaving the N-terminal amino acids of a polypeptide is provided. In one embodiment, the polypeptide is immobilized on a surface. In another embodiment, the cleavage of the N-terminal amino acids of the polypeptide is carried out at the single-molecule level. In a specific embodiment, a modified synthetic or recombinant aminopeptidase (where the remaining amino acid sequence of the aminopeptidase is that of wild-type Trypanosoma cruzi or cruzain, having a glycine residue at position 25, a serine residue at position 65, a cysteine ​​residue at position 138, and a histidine residue at position 208) is provided. The use of a sequence having at least 75%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity to the amino acid sequence of cruzipain or cruzain as depicted in SEQ ID NO: 2 is provided for cleaving the N-terminal amino acids of a polypeptide. More specifically, the aminopeptidase provided for cleaving the N-terminal amino acids of a polypeptide has a cysteine ​​residue inserted after the first methionine residue of wild-type T. cruzipain or cruzain. More specifically, the aminopeptidase provided for cleaving the N-terminal amino acid of a polypeptide comprises the amino acid sequence as depicted in SEQ ID NO: 3 or SEQ ID NO: 4, or consists of the amino acid sequence as depicted in SEQ ID NO: 5 or SEQ ID NO: 6. Most specifically, the cleavage is carried out at the single-molecule level, and the polypeptide is immobilized on a surface, more specifically through the C-terminus of the polypeptide.

[0091] In another specific embodiment, the use of a modified synthetic or recombinant aminopeptidase (where a cysteine ​​residue is inserted after the first methionine residue of the wild-type aminopeptidase) having at least 75%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity with respect to the amino acid sequence of wild-type Thermus aquaticus aminopeptidase T as depicted in Sequence ID No. 7 is provided for cleaving the N-terminal amino acids of a polypeptide. In a more specific embodiment, the aminopeptidase T provided for cleaving the N-terminal amino acids of a polypeptide comprises or comprises the amino acid sequence as depicted in Sequence ID No. 8. Most specifically, the cleavage is carried out at the single-molecule level, and the polypeptide is immobilized on a surface. In other embodiments, the aminopeptidase is a catalytically active aminopeptidase. Also provided herein is the use of T. aquaticus aminopeptidase T for cleaving or identifying the N-terminal amino acid from a surface-immobilized polypeptide, where the N-terminal amino acid is selected from the list of Leu, Met, Tyr, Arg, Pro, Gly, Lys, Ala, and Val.

[0092] In another aspect, the use of an aminopeptidase, more specifically a catalytically active aminopeptidase, is provided for identifying or categorizing the N-terminal amino acids of a polypeptide, or for sequencing the polypeptide. In a specific embodiment, such identification, categorization, or sequencing is performed at the single-molecule level. More specifically, the polypeptide is immobilized on a surface through its C-terminus. In one embodiment, the aminopeptidase is labeled, specifically by optical, electrical, or plasmonic labeling, or the aminopeptidase is detected optically, electrically, or plasmonically. In a specific embodiment, a modified synthetic or recombinant aminopeptidase comprising an amino acid sequence having a glycine residue at position 25, a serine residue at position 65, a cysteine ​​residue at position 138, and a histidine residue at position 208 of wild-type Trypanosoma cruzi (where the remaining amino acid sequence of the aminopeptidase is the wild-type T. as depicted in SEQ ID NO: 1). The use of sequences having at least 75%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity to the amino acid sequence of cruzipain or cruzain as depicted in Sequence ID No. 2 is provided for identifying the N-terminal amino acids of a polypeptide or for sequencing a polypeptide.

[0093] More specifically, the aminopeptidase provided for identifying the N-terminal amino acids of a polypeptide or for sequencing the polypeptide has a cysteine ​​residue inserted after the first methionine residue of wild-type T. cruzi or cruzain. More specifically, the aminopeptidase provided for identifying the N-terminal amino acids of a polypeptide or for sequencing the polypeptide contains the amino acid sequence as depicted in SEQ ID NO: 3 or SEQ ID NO: 4, or consists of the amino acid sequence as depicted in SEQ ID NO: 5 or SEQ ID NO: 6. In specific embodiments, the identification, categorization, or sequencing is performed at the single-molecule level. More specifically, the polypeptide is immobilized on a surface through its C-terminus.

[0094] In another specific embodiment, the use of a modified synthetic or recombinant aminopeptidase having at least 75%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity with respect to the amino acid sequence of wild-type Thermus aquaticus aminopeptidase T as depicted in Sequence ID No. 7 (where a cysteine ​​residue is inserted after the first methionine residue of the wild-type aminopeptidase) is provided for identifying the N-terminal amino acids of a polypeptide or for sequencing a polypeptide.

[0095] In a more specific embodiment, the aminopeptidase T provided for identifying or categorizing the N-terminal amino acids of a polypeptide, or for sequencing a polypeptide, comprises or consists of the amino acid sequence depicted in SEQ ID NO: 8. In a specific embodiment, the identification, categorization, or sequencing is performed at the single-molecule level. Even more specifically, the polypeptide is immobilized on a surface through its C-terminus. In another embodiment, the aminopeptidase is a catalytically active aminopeptidase. In the most specific embodiment, the N-terminal amino acid identified using the T. aquaticus aminopeptidase T is selected from the list consisting of Leu, Met, Tyr, Arg, Pro, Gly, Lys, Ala, and Val.

[0096] Method of this application The present invention as described in this application is based on several methods disclosed below. In one aspect, a method is provided for identifying or categorizing terminal amino acids of a surface-immobilized polypeptide, the method for identifying or categorizing said terminal amino acids as follows: a) Contacting the polypeptide immobilized on the surface with a cleavage inducer, the cleavage inducer binding to terminal amino acids and cleaving them from the polypeptide; b) Measuring the residence time of the cleavage inducer on the terminal amino acid; c) Comparing the measured residence time with a set of residence time reference values ​​characteristic of the cleavage inducer and a set of terminal amino acids. Includes.

[0097] Furthermore, a method for obtaining sequence information of a polypeptide immobilized on a surface is also provided, and this method is as follows: a) Contacting the polypeptide immobilized on the surface with a cleavage inducer, the inducer which binds to terminal amino acids and cleaves them from the polypeptide; b) Measuring the residence time of the cleavage inducer on the terminal amino acids of the polypeptide immobilized on the surface; c) Identifying or categorizing the terminal amino acids by comparing the measured residence time with a set of residence time reference values ​​characteristic of the cleavage inducer and a set of terminal amino acids; d) The cleavage inducer is made capable of cleaving the terminal amino acid; e) Repeat steps a) to d) at least once. Includes.

[0098] In this method, the residence time is measured optically, electrically, or by plasmon. In a specific embodiment, the series of terminal amino acids includes or consists of Leu, Met, Tyr, Arg, Pro, Gly, Lys, Ala, and Val.

[0099] In various embodiments, the cleavage inducer in the method is an isothiocyanate or ITC analog selected from a list consisting of PITC, CITC, SPITC, and azide-PITC. In such cases, the residence time is the time it takes for the terminal amino acid to be removed, and the terminal amino acid is identified by comparing the residence time with a set of reference values ​​for various amino acids.

[0100] In other embodiments, the cleavage inducer in the method is a peptidase, more specifically a catalytically active peptidase, and even more specifically a catalytically active aminopeptidase. When the aminopeptidase is used in the method of the present application, the polypeptide is immobilized on a surface through its C-terminus. In a particular embodiment, the step of measuring the residence time of the cleavage inducer on the terminal amino acid in the above method is to measure the residence time of the cleavage inducer on the terminal amino acid until the terminal amino acid of the polypeptide immobilized on the surface is cleaved.

[0101] Throughout the present application, the cleavage inducer may be an aminopeptidase, an ITC, or an ITC analog. As already discussed herein, the polypeptide immobilized on a surface should be denatured so that its N-terminus is easily accessible for chemical or enzymatic cleavage, but also to avoid steric hindrance or interference of said cleavage (in the case where the polypeptide is immobilized through its C-terminus). Thus, the present application also provides a method that includes a first step of polypeptide denaturation. Under such denaturation conditions, the catalytically active aminopeptidase to be used should withstand the denaturation conditions. Therefore, in these cases, the aminopeptidase is preferably a thermophilic and / or solvent-resistant aminopeptidase.

[0102] In various embodiments, the method described herein is provided, wherein the aminopeptidase is any of the aminopeptidases disclosed in this application, more precisely, either the cruzain and cruzipain peptidases derived from T. cruzi, or any of the T. aquaticus aminopeptidase T described herein or as described herein. In a specific embodiment, the method is provided, wherein the N-terminal amino acid is selected from the list consisting of Leu, Met, Tyr, Arg, Pro, Gly, Lys, Ala, and Val.

[0103] In other embodiments, the method described herein is provided, wherein the N-terminal amino acid is derivatized, and wherein the cleavage inducer is an aminopeptidase capable of cleaving the derivatized N-terminal amino acid. In more specific embodiments, the N-terminal amino acid is derivatized with CITC or SPITC, and the cleavage inducer is either cruzaine or cruzipain, both derived from T. cruzi, as disclosed herein in their modified form.

[0104] In various embodiments of this application, the methods herein for identifying or categorizing N-terminal amino acids from a C-terminally immobilized polypeptide or for obtaining sequence information from said polypeptide are methods performed at the single-molecule level.

[0105] For single-molecule measurement, the polypeptide is assumed to be immobilized on an active detection surface according to the method of the present application. In a specific embodiment, the active detection surface is either a gold surface to which the polypeptide is chemically coupled, or an amide-, carboxyl-, thiol-, or azide-functionalized surface.

[0106] Multiple measurements of residence time and use with non-cleaving binders In an alternative embodiment, the aminopeptidase used in the method of the application may be an aminopeptidase that cleaves the N-terminal amino acid only after several rounds of binding and dissociation of the N-terminal amino acid. Each residence time of the aminopeptidase may be informative for determining the residence time until the N-terminal amino acid is cleaved and may also be useful for identifying the N-terminal amino acid. To detect the point at which the identity of the N-terminal amino acid changes due to the aminopeptidase and to more accurately predict the N-terminal amino acid in a single-molecule mechanism (set-up), it is recommended to have multiple measurements for each N-terminal amino acid. This can be achieved by using an aminopeptidase that will dock (associate) and undock (dissociate) the N-terminal amino acid several times before the actual cleavage will occur.

[0107] Therefore, it is also assumed that the step in the present invention to measure the residence time of a catalytically active aminopeptidase implies measuring multiple residence times of the aminopeptidase before the aminopeptidase cleaves the N-terminal amino acid. Alternatively, the present invention provides a method in which the residence time of the catalytically active aminopeptidase is measured for each binding event of the aminopeptidase to the N-terminal amino acid. This is demonstrated in Example 13 and Figure 14.

[0108] In a specific embodiment, the method disclosed in the present application is provided, wherein the aminopeptidase used for the enzymatic cleavage of the N-terminal amino acid has, on average, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 20, or at least 50 association / dissociation cycles during the time window required for the aminopeptidase to cleave the N-terminal amino acid. This means that at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 20, or at least 50 association / dissociation cycles without cleavage-production capacity occur between cleavage-production cycles.

[0109] Also provided is the method of the present application, in which the polypeptide immobilized on the surface is further contacted with a protein that binds to one or more terminal amino acids, and the kinetics of the binding event of the one or more binding proteins to the terminal amino acids identifies the terminal amino acids. The feasibility of using the binding specificity of proteins that bind to N-terminal amino acids to collect substrate information has been theoretically demonstrated by Rodriques et al. (2018, bioRxiv, doi: http: / / dx.doi.org / 10.1101 / 352310). The additional use of the non-cleavable binder (apart from catalytically active aminopeptidases) in the method of the present application may provide additional information to predict or identify N-terminal amino acids with greater accuracy in single-molecule experiments. In a specific embodiment, the non-cleavable binder has at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 20, or at least 50 binding / dissociation cycles with the N-terminal amino acid during the time window required for the cleavage inducer to cleave the N-terminal amino acid. The cleavage inducer is an aminopeptidase or an ITC or an ITC analog, more specifically, the aminopeptidase is one of the aminopeptidases described in the present application.

[0110] Cut detection One additional aspect of the method of the present application is that the cleavage of a terminal amino acid is to be detected or confirmed. Therefore, also provided herein are methods of the present application that include the additional step of determining the cleavage of the terminal amino acid by measuring the optical, electrical, or plasmonic signal of a surface-immobilized polypeptide, where the difference in the optical, electrical, or plasmonic signal represents the cleavage of the terminal amino acid. Indeed, immobilized peptides with a free N-terminus have several properties that can be used to determine when the N-terminal amino acid is cleaved by the cleavage inducer of the present invention.

[0111] In the first example, the free N-terminal amine group is positively charged over a wide pH range. The distance between this positive charge of the peptide and the anchor point (through which the peptide is immobilized) can be measured, for example, by measuring random telegraph noise (Sorgenfrei et al 2011 Nano Lett 11:3739-3743) in potentiometric detection when the peptide is immobilized on a suitably designed detector element (such as carbon nanotubes, nanometer-scale transistors including field-effect transistors, particularly fin-shaped field-effect transistors, gate all-around field-effect transistors, nanoribbon field-effect transistors, etc.).

[0112] When the N-terminal amino acid is cleaved by the cleavage-inducing reagent of the present invention, the positively charged N-terminal amino group moves closer to the peptide's anchor point and, consequently, to the detector surface. In a fully stretched-out peptide, the shortened distance between this charge and the anchor point is approximately 3.8 angstroms (contour length), constrained by the geometry of the covalent bonds in the peptide backbone. Therefore, under environmental conditions that disrupt the peptide's secondary structure (high temperature, exposure to organic (co)solvents, etc.), the maximum value of the length measurement distribution between the amino end charge and the peptide anchor point has an upper limit constrained by the geometry of the covalent bonds in the peptide backbone. Measuring this change in maximum length by repeatedly observing the amino end charge of the peptide while the cleavage-inducing agent is present reveals the exact moment when the cleavage-inducing agent cleaves the N-terminal amino acid.

[0113] In the second example, the amino acid at the amino end can be reacted with a reagent such that the positive charge on the terminal amino group is eliminated, converted to an amino acid derivative carrying one or more negative charges, or an amino acid derivative in which a single positive charge is increased to a compound positively charged amino acid derivative. This can be achieved, for example, by contacting the immobilized peptide with a suitably selected N-hydroxysuccinimidyl reagent that is uncharged or carries one or more positive charges or one or more negative charges.

[0114] Alternatively, the charge-modulating reagent may be the cleavage-inducing reagent itself, as in the case where the terminal amino group of the immobilized peptide is reacted with a suitably selected isothiocyanate reagent (such as PITC, CITC, SPITC (4-sulfophenyl isothiocyanate), or azidophenyl isothiocyanate). The latter in the isothiocyanate reagent may be further modified by click chemistry on the azido group either prior to, during, or after contact of the agent with the immobilized peptide. Thus, the charge difference between the peptide carrying the amino acid derivative and this peptide after the N-terminal amino acid derivative has been cleaved is rendered binary (neutral to positive conversion, negative to positive conversion, or multiple positive charges to single positive charge conversion), or is enhanced, or both. Using detection techniques similar to those in Example 1, the point at which the cleavage-inducing agent effectively cleaves the amino acid can be measured using the detection of this charge change.

[0115] In the third example, the amino group or side chain of the N-terminal amino acid may be reacted with an agent that confers spectroscopically distinguishable properties to produce an N-terminal amino acid derivative that can be detected using spectroscopic methods such as fluorescence quantification, Raman spectroscopy, or plasmon resonance. In particular, single-molecule detection using total internal reflection fluorescence (TIRF) microscopy is a preferred method because it is designed to detect fluorescence from a thin layer parallel to a reflective surface (e.g., glass) to which the peptide can be immobilized.

[0116] When a cleavage inducer (which may, for example, be a molecule containing aminopeptidase, edmanase, or isothiocyanate) is brought into contact with the immobilized peptide, the point at which cleavage occurs can be detected by changes in the spectroscopic properties of the immobilized peptide over an observational time series (e.g., loss of fluorescence signal due to cleavage of a fluorescently labeled N-terminal amino acid derivative).

[0117] Alternatively, loss of Förster resonance energy transfer (FRET) signaling may be observed when the immobilized peptide contains a suitable FRET donor or acceptor, or when the N-terminal amino acid derivative contains a matching FRET acceptor or donor. In another embodiment, the N-terminal amino acid is derivatized (e.g., using a biotinylated isothiocyanate, for example, with biotin) so that a binder (e.g., an avidin such as streptavidin or neutraavidin) carrying a spectroscopically identifiable label (e.g., a fluorophore) can bind to the derivatized N-terminal amino acid. The time until the N-terminal amino acid is cleaved can then be measured as the point at which a change occurs in the observational time series of the immobilized peptide's ability to bind to the binder.

[0118] Binding ability is the ability of a peptide to bind to or not bind to a binder, or binding affinity, k on , k off These are characteristics of such binding. Detection can be performed, for example, using TIRF. In another embodiment, the N-terminal amino acid is converted (for example, by reaction with a molecule containing isothiocyanate) to a derivative to which a binder (for example, a catalytically active or inactive aminopeptidase) carrying a spectroscopically identifiable label (for example, a fluorophore) cannot bind. Then, upon cleavage with a cleavage inducer, such a binder can bind to the immobilized peptide, and the time until the N-terminal amino acid is cleaved again can then be measured as the point in time at which a change occurs in the observational time series of the immobilized peptide's ability to bind to the binder.

[0119] In another embodiment, the time until cleavage by a cleavage inducer can be detected by detecting changes in the binding affinity or binding kinetics of the peptide binding agent (e.g., a catalytically active or inactive aminopeptidase or edmanase) to the immobilized peptide. For example, the residence time of the peptide binding agent (the time between binding and dissociation) can be measured using any of the techniques described above. The number of cycles of binding and dissociation of the peptide binding agent can be measured, and changes in the average residence time can be detected during cleavage of the N-terminal amino acid.

[0120] In a specific example, the peptide binding agent used for detection has a binding / dissociation kinetics faster than the time required for a cleavage inducer to induce cleavage of the N-terminal amino acid, such that multiple measurement points of the binding / dissociation of the peptide binding agent to the immobilized peptide are typically observable between two cleavage events of the N-terminal amino acid. In a specific example, the peptide binding agent is the same as the cleavage inducer. For example, catalytically active aminopeptidases have both binding / dissociation cycles with cleavage-production capability and binding / dissociation cycles without cleavage-production capability. During a peptide binding event without cleavage-production capability, changes in these properties can be observed as the N-terminal amino acid is cleaved and therefore a new N-terminal amino acid is exposed to interact with the aminopeptidase, by measuring the affinity, particularly the kinetics, of the aminopeptidase binding / dissociation over time.

[0121] In a specific embodiment, the aminopeptidase is used under conditions far from the optimal conditions for the catalytic reaction rate of the enzyme, such that most of the binding / dissociation events of the aminopeptidase to the immobilized peptide are non-cleavage-productive, and the kinetic changes of these binding events in the time series are used to inform about the point in time when the aminopeptidase cleaves the N-terminal amino acid.

[0122] In another aspect, a method is provided for identifying or categorizing the N-terminal amino acids of a polypeptide immobilized on a surface through its C-terminus at the single-molecule level by determining the “on-time” value or residence time of the aminopeptidase on the N-terminal amino acids, the method comprising contacting the surface-immobilized polypeptide with the aminopeptidase and measuring the “on-time” value or residence time of the aminopeptidase.

[0123] In a specific embodiment, the aminopeptidase is any of the aminopeptidases described in this application. Thus, the method is provided, wherein the aminopeptidase is a modified synthetic or recombinant aminopeptidase comprising the amino acid sequence of wild-type Trypanosoma cruzi (cruzipain or cruzain), having a glycine residue at position 25, a serine residue at position 65, a cysteine ​​residue at position 138, and a histidine residue at position 208, wherein the remaining amino acid sequence of the aminopeptidase is the wild-type T. as depicted in Sequence ID No. 1. The sequence contains at least 75%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity with respect to the amino acid sequence of cruzipain or cruzain as depicted in Sequence ID No. 2.

[0124] More specifically, the aminopeptidase provided for the method has a cysteine ​​residue inserted after the first methionine residue of wild-type T. cruzi or cruzain. More specifically, the aminopeptidase comprises the amino acid sequence as depicted in SEQ ID NO: 3 or SEQ ID NO: 4, or consists of the amino acid sequence as depicted in SEQ ID NO: 5 or SEQ ID NO: 6.

[0125] In another specific embodiment, a sixth aspect of the method is provided, where the aminopeptidase has at least 75%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity with respect to the amino acid sequence of wild-type Thermus aquaticus aminopeptidase T as depicted in SEQ ID NO: 7, where the cysteine ​​residue is inserted after the first methionine residue of the wild-type aminopeptidase. In a more specific embodiment, the aminopeptidase T comprises or consists of the amino acid sequence depicted in SEQ ID NO: 8. In another embodiment, the aminopeptidase is a catalytically active aminopeptidase. In its most specific embodiment, the N-terminal amino acid to be identified or categorized is selected from the list consisting of Leu, Met, Tyr, Arg, Pro, Gly, Lys, Ala, and Val.

[0126] In a more specific embodiment, the N-terminal amino acid is derivatized, and the aminopeptidase is an aminopeptidase that includes a binding domain for the derivatized N-terminal amino acid and cleaves the derivatized N-terminal amino acid. In an even more specific embodiment, the derivatized N-terminal amino acid is a CITC- or SPITC-derivative amino acid, and the aminopeptidase can bind to and cleave the CITC- or SPITC-derivative amino acid. For the latter purpose, the aminopeptidase is specifically modified. A non-limiting example of such a modified aminopeptidase that binds to and cleaves a CITC- or SPITC-derivative N-terminal amino acid is T. cruzi cruzipain or cruzain of this application, modified to have a glycine residue at position 25, a serine residue at position 65, a cysteine ​​residue at position 138, and a histidine residue at position 208.

[0127] In another aspect, a method is provided for sequencing a surface-immobilized polypeptide at the single-molecule level, the method comprising: a) contacting the surface-immobilized polypeptide with an aminopeptidase, more specifically a catalytically active aminopeptidase; b) measuring the enzyme “on-time” value of the aminopeptidase; c) identifying or categorizing the N-terminal amino acid by the “on-time” value; and repeating steps a) to c) one or more times. In a specific embodiment, the aminopeptidase is any of the aminopeptidases described in this application.

[0128] Therefore, the present invention provides a modified synthetic or recombinant aminopeptidase comprising the amino acid sequence of wild-type Trypanosoma cruzi (cruzipain or cruzain), having a glycine residue at position 25, a serine residue at position 65, a cysteine ​​residue at position 138, and a histidine residue at position 208, wherein the remaining amino acid sequence of the aminopeptidase is the wild-type T. as depicted in Sequence ID No. 1. The aminopeptidase contains a sequence having at least 75%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity with respect to the amino acid sequence of cruzipain or cruzain as depicted in SEQ ID NO: 2. More specifically, the aminopeptidase has a cysteine ​​residue inserted after the first methionine residue of wild-type T. cruzipain or cruzain. More specifically, the aminopeptidase contains the amino acid sequence as depicted in SEQ ID NO: 3 or SEQ ID NO: 4, or consists of the amino acid sequence as depicted in SEQ ID NO: 5 or SEQ ID NO: 6.

[0129] In another specific embodiment, a seventh aspect of the method is provided, where the aminopeptidase is a modified synthetic or recombinant aminopeptidase having at least 75%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity with respect to the amino acid sequence of wild-type Thermus aquaticus aminopeptidase T as depicted in SEQ ID NO: 7, wherein the cysteine ​​residue is inserted after the first methionine residue of the wild-type aminopeptidase. In a more specific embodiment, the aminopeptidase T comprises or consists of the amino acid sequence depicted in SEQ ID NO: 8. In its most specific embodiment, the N-terminal amino acid from the method is selected from a list consisting of Leu, Met, Tyr, Arg, Pro, Gly, Lys, Ala, and Val.

[0130] In a more specific embodiment, the N-terminal amino acid is derivatized, and the aminopeptidase includes a binding domain for the derivatized N-terminal amino acid and is an aminopeptidase that cleaves the derivatized N-terminal amino acid. In an even more specific embodiment, the derivatized N-terminal amino acid is a CITC- or SPITC-derivatized amino acid, and the aminopeptidase can bind to and cleave the CITC- or SPITC-derivatized amino acid. In another embodiment, the aminopeptidase is a catalytically active aminopeptidase.

[0131] In another embodiment, the "on-time" value in a method for sequencing a surface-immobilized polypeptide at the single-molecule level is measured optically, electrically, or by plasmons.

[0132] In the most specific embodiments of the present application, the method described herein is carried out under protein denaturation conditions. These protein denaturation conditions are obtained by high temperature and the presence of a solvent. In specific embodiments, the high temperature is between 40°C and 120°C, or between 50°C and 110°C, or between 60°C and 100°C, or between 70°C and 90°C. In specific embodiments, the solvent is selected from a list consisting of crosslinking agents such as acetic acid, trichloroacetic acid, sulfosalicylic acid, sodium bicarbonate, ethanol, alcohol, formaldehyde and glutaraldehyde, chaotropic agents such as urea, guanidine chloride and lithium perchlorate, and disulfide bond-breaking agents such as 2-mercaptoethanol, dithiothreitol, or tris(2-carboxyethyl)phosphine. Most specifically, the solvent is acetonitrile, ethanol, or methanol.

[0133] In another most specific embodiment of the present application, the cleavage inducer is a covalent bond cleavage inducer. In another most specific embodiment of the present application, the non-cleavage binder is a covalent bond non-cleavage binder.

[0134] The following examples are intended to facilitate a further understanding of the invention. While the invention is described herein with reference to the embodiments described, it should be understood that the invention is not limited thereto. Those skilled in the art and those with access to the teachings herein will recognize additional modifications and embodiments within that scope. Accordingly, the invention is limited only by the claims appended herein.

[0135] example Havranek and Borgo's peptide sequencing relies on the use of NAABs, which are aminopeptidases or tRNA synthetases that do not function catalytically. As an alternative to existing Edman degradation, which cleaves the N-terminal amino acid, an Edmanase enzyme was generated (WO2014273004). This cruzain cysteine ​​protease from Trypanosoma cruzi was modified at four positions to be able to bind to (and cleave) the N-terminal amino acid of PITC-derivative peptides. A first introduction-point mutation (C25G) replacing the catalytic cysteine ​​at position 25 with glycine rendered the enzyme catalytically incompetent unless rescued by a sulfur atom of the PITC-derivative peptide substrate. Three additional mutations, namely G65S, A138C, and L160Y, were needed to improve adaptation to PITC-derivative substrates.

[0136] While Havranek and Borgo used non-catalyzed NAABs for amino acid identification and edmanase for cleaving the N-terminal amino acid after identification, we have surprisingly developed a method that does not require non-catalyzed NAABs. This method relies entirely on a modified aminopeptidase, the residence time from the aminopeptidase is informative for the N-terminal amino acid it binds to and cleaves. The enzymes disclosed herein have affinity for the (derivativeized) amino acid at the N-terminus, regardless of the N-terminal amino acid identity. However, to obtain peptide sequence information, the enzymes exhibit variability in catalytic efficiency and residence time depending on the N-terminal amino acid identity (Figures 1-2). Even more surprisingly, we demonstrate a method for identifying an N-terminal amino acid that correlates exclusively with the residence time of an ITC or ITC analog on the N-terminal amino acid.

[0137] Example 1: TIRF microscopy for single peptide detection In the first step, a system was developed to immobilize peptides that would be sequenced on the surface. An azide-functionalized, oven-cleaned glass plate was used as the surface, and the peptide NNGGNNGGRGNK, with a DBCO-PEG8 group attached to the N-terminus and a sulfo-Cy5 fluorescent probe attached to the C-terminus, was used as the test peptide. The test peptide was immobilized via an azide-DBCO click reaction. 1 ml of 1 nM test peptide was placed on the azide-functionalized glass plate and incubated in the dark for 24 hours. For functionalization, 11-azido-undecyl(trimethoxy)silane was used to make the glass surface hydrophobic, allowing the liquid to float on the glass. After 24 hours, the glass plate was washed three times with 1 ml of MS-grade water (30 min washing each time). For the control sample, the glass plate was incubated in water. Microscopy was performed using a Zeiss TIRF microscope, and photographs were taken using λ. Em Images were taken at 639 nm. The peptide was well immobilized on the glass surface, and an appropriate spatial distribution was obtained at a concentration of 1 nM (Figure 3).

[0138] Example 2: Trypsin digestion of surface-immobilized peptides Next, the conditions were optimized for enzymatic cleavage of the surface-immobilized peptide. The test peptide DBCO-PEG8-NNGGNNGGRGNK-Cy5 was used again, but this time with trypsin. Good enzymatic surface reaction was detected after cleavage with arginine to remove the fluorescent probe. 1 ml of 1 nM test peptide was placed on an azide-functionalized and oven-cleaned glass plate and incubated in the dark for 24 hours. After 3 × wash with MS-grade water, the glass plate was incubated with 100 nM trypsin (sequencing grade, Promega) at room temperature for 1 hour. The control was incubated with water. After trypsin treatment, the plate was washed again with water for 3 × wash. The experiment was also repeated in the presence of 1 μM DBCO-PEG8-amide passivator (added to the test peptide during the 24-hour azide-DBCO click reaction) to evaluate its effect on clustering and specific trypsin surface interaction.

[0139] The trypsin reaction in the absence of a passivator was not favorable due to high background signaling from nonspecific binding of trypsin to the surface (Figure 4A). The background was reduced to a lower λ of 561 nm. Em The assessment was performed using the Cy3 channel. When the passivation agent was present, the background disappeared, and a significant decrease was observed at the spot (Figure 4B), indicating that a considerable amount of immobilized peptide was cleaved.

[0140] Example 3: Peptide N-terminal reactive probe Depending on the aminopeptidase used (i.e., whether it binds to the derivatized N-terminal amino acid or to an unlabeled N-terminal amino acid, see Example 4 and further), it is determined whether the immobilized peptide should be labeled. The choice of probe will depend on the readout strategy (Figure 2). Nevertheless, the probe must be carefully selected: the probe must be highly reactive to the peptide N-terminus, and the derivatized peptide substrate must be compatible with the catalytic site of the enzyme (see Example 4 for enzyme modification).

[0141] Potential fluorescent probe candidates include fluoresamine, o-phthalaldehyde, dansyl chloride, and coumarinyl isothiocyanate (CITC) derivatives. For charged probes, an interesting candidate would be sulfophenyl isothiocyanate (SPITC), a negatively charged variant of the PITC probe used in MS de novo peptide sequencing to neutralize the N-terminal fragment ion (Samyn et al. 2004 J Am Soc Mass Spectrom 15:1838-1852). Specificity to N-terminal primary amines can be increased by carefully controlling the pH (unlike with lysine primary ε-amines), but some degree of specificity should be taken into consideration, especially when using tryptic peptides, where the peptides are derivatized with either CITC or SPITC.

[0142] Example 4: Re-modification of edmanase In the quest to find an aminopeptidase capable of binding to and cleaving any N-terminal amino acid of any peptide, two research lines were explored in parallel: one using a modified cruzipain / cruzain cysteine ​​protease derived from Trypanosoma cruzi, and the other using Thermus aquaticus aminopeptidase T.

[0143] We first remodeled the cruzain and cruzipain cysteine ​​proteases derived from Trypanosoma cruzi so that they could bind to the derivatized N-terminal amino acids, more precisely to the CITC-derivative and SPITC-derivative N-terminal amino acids. Several potential remodeling options were provided by computer docking (AutoDock-Vina; Trott and Olsen 2010 J Comput Chem 31:455-461) of fluorescent or charged isothiocyanate probes onto the cruzain or cruzipain enzymes and hypothetical mutants. SPITC-Ala-Phe docked well onto cruzain and cruzipain when Tyr160 was renatived to Leu (Figure 5A). The nearby Glu208 could have resulted in a charge collision with the sulfate of SPITC, but this could be avoided by effectively mutating it to His. By doing so, we developed modified T. cruzi cruzain and cruzipain containing four point mutations (i.e., C25G, G65S, A138C, E208H), and the fluorescent 3-CITC-Ala-Phe on the modified enzyme, as well as its 5-, 6-, 7-, and 8-coumarinyl analogs, resulted in the overall correct docking configuration within the groove of the active site (Figure 5B).

[0144] Example 5. Transformation of T. cruzi_pET24b(+) by E. coli BL21(DE3) In the next step, modified crujipain (C25G, G65S, A138C, E208H) were recombinantly produced. The codon-optimized DNA sequences encoding the re-modified crujipain were cloned into the pET24b(+) plasmid (NdeI-BamHI). Cys was added to the N-terminus of the protein, and 6xHis was added to the C-terminus. The full-length protein has a molecular weight of ±37.5 kDa and an estimated pI of 5.7. Its sequence is depicted in Sequence ID No. 5: [ka]

[0145] Successful transformation of the remodeled crujipain was obtained (Figure 6A). Next, the kinetics of the remodeled crujipain were assayed with synthetic SPITC- and CITC-derivativeated 7-amino-4-methylcoumarin (AMC) amino acid analogs. Fluorescent AMC was released and detected upon removal of the SPITC- or CITC-derivativeated N-terminal amino acid. Computer simulations were then used to assess the feasibility of using enzyme "on-time" values ​​for peptide sequencing.

[0146] Example 6: Transformation of E. coli BL21(DE3) with T. aquaticus aminopeptidase T_pET24b(+) A codon-optimized DNA sequence encoding aminopeptidase T (Taq-APT or TaqAPT) from Thermus aquaticus was synthesized. The Taq-APT gene was then cloned into the pET24b(+) plasmid (NdeI-BamHI) (Figure 9). Cys was added to the N-terminus and 6xHis to the C-terminus. The full-length protein has a molecular weight of ±45.7 kDa and an estimated pI of 5.6. Its sequence is depicted in Sequence ID No. 8: [ka] Successful transformation of the modified Taq-APT was obtained (Figure 6).

[0147] Example 7: Taq-APT expression and purification Recombinant expression of Taq-APT in E. coli BL21(DE3) transformants was validated and optimized. Purification of the (thermophilic) protein was tested using Ni-NTA spin column chromatography, heat treatment, or a combination of both. Aliquots of BL21(DE3) pET-24b(+)-Taq-APT transformed cells were added to 5 ml of LB + 50 μg / ml kanamycin and grown overnight at 37°C. The cells were then diluted 100× in 5 ml of fresh LB + 50 μg / ml kanamycin. After incubation at 37°C for 2 hours, induction with 1 mM IPTG was performed, followed by overnight growth at 28°C. Cells were collected (4,000×g for 10 min), resuspended in 1 ml PBS + 10 mM imidazole (pH 7.5-8), sonicated (1 s on, 1 s off, 90 sec, 30% amplitude), and centrifuged at full speed for 5 min. The supernatant was collected and divided into 4 × 200 μl portions to test four different conditions: (A) unprocessed soluble fraction, (B) Ni-NTA purification, (C) heat treatment, and (D) Ni-NTA purification + heat treatment. After treatment, the samples were concentrated on a 10 kDa spin column until the volume was approximately 50 μl. 10 μl of this concentrated portion was mixed with 10 μl of SDS sample buffer, and the samples were analyzed by SDS-PAGE (Figure 8). For details on the Ni-NTA purification protocol, please refer to the experimental procedure. Purification by heat was performed at 80°C for 30 min.

[0148] The combination of Ni-NTA and heating resulted in an acceptablely pure Taq-APT extract (Figure 7). Peptidase assays using L-leucine-p-nitroaniline (140 μl PBS buffer, 24 mM L-leucine-p-nitroaniline in 10 μl MeOH, and 10 μl purified Tag-APT) confirmed the presence of an active (amino)peptidase with thermophilic properties (activity at 70°C, optimal for Taq) (Figure 8).

[0149] Example 8. Determination of kinetic parameters of T. aquaticus aminopeptidase T. The kinetic parameters of T. aquaticus aminopeptidase T for cleaving various amino acid substrates were determined by a p-nitroanilide assay. The substrates consisted of N-terminal amino acids with a p-nitroanilide attached to their C-terminus. During amino acid cleavage, the release of nitroanilide could be monitored by measuring the absorbance at 405 nm. T. aquaticus aminopeptidase T, as depicted in Sequence ID No. 8, was added at a concentration of 2.0625 μM (in PBS) to various concentrations of amino acid p-nitroanilide substrates (0.0625, 0.125, 0.25, 0.5, 1, and 2 mM in PBS). Subsequently, p-nitroanilide release was continuously measured at 40°C using a FLUOstar Omega microplate reader (MBG LabTech). The initial reaction rate at each substrate concentration was derived from this (v0). Lineweaver-Burke plots were generated for each amino acid, and the reaction V was derived from these plots. max and enzyme-substrate K M The following was determined: Metabolic turnover rate k cat V max and (k) calculated from enzyme concentration cat =V max / [E]). Next, the enzyme on-time value k cat The values ​​were calculated by taking the reciprocal of the values. The kinetic parameters for the nine amino acids are listed in Table 1. The "on-time" value is 1 / k as shown in Table 1. cat This is calculated as the total time required for the catalytic reaction to occur on the peptide in the enzyme solution.

[0150] Table 1. Kinetic parameters of the p-nitroanilide assay using nine different amino acids, including the on-time of T. aquaticus aminopeptidase T related to nine different amino acids. [Table 1]

[0151] The assay was performed at 40°C, which is below the enzyme's optimal temperature (temperature optimum). When the enzyme was operated at a temperature closer to the optimal temperature of 70°C, the reaction rate (k) was lower. cat The enzyme (on-time) is sped up. In short, we have found that, surprisingly, T. aquaticus aminopeptidase T, as described herein, can bind to and cleave nine different amino acids with a differential kinetic. Moreover, the reaction kinetics are linked to the identity of the amino acids, and even more surprisingly, the k of various amino acids cat The spread of the values ​​distinguishes various amino acids from each other, and thus these amino acids can be identified. The results presented herein not only demonstrate the practicality of the aminopeptidase disclosed herein, but also support and embody the methods and uses disclosed in this application.

[0152] Example 9. Activity of T. aquaticus aminopeptidase T against various amino acid p-nitroanilide substrates at 40°C and 80°C. The TaqAPT enzyme exhibits activity for all amino acid p-nitroanilide substrates in the current test panel. As described herein and in light of the present invention as described in this application, the activity differs among various amino acids. At 80°C, the amino acid substrate panel can be broadly divided into fast-cleaved substrates (L, M, Y, R, F) and slow-cleaved substrates (D, P). However, at 40°C, the panel appears to be divided into fast-cleaved (Y, R) substrates, slow-cleaved (D, P) substrates, and substrates in between (L, M, F). At 40°C, the activity shows a reduction of roughly 10× to 3× depending on the N-terminal amino acid.

[0153] Ultimately, TaqAPT is not only active in p-nitroanilide assays. More importantly, we have demonstrated that TaqAPT also cleaves peptide substrates. For example, Figure 11 shows that TaqAPT cleaves dipeptides even when proline is in the second position. While peptide bonds adjacent to the amino acid proline are resistant to cleavage by most peptidases (Iyver et al. 2015 FEBS Open Bio. 2015 Apr 2;5:292-302), the N-terminal amino acid from peptides with proline in the second position is not easily cleaved. In contrast to what is stated by Minagawa et al. (1988 Agricultural and Biological Chemistry 52:1755-1763), we here show that T. aquaticus aminopeptidase T can cleave the N-terminal amino acid even when proline is in the second position. This remarkable finding greatly encourages the comprehensive use of Taq-APT in the methods and uses disclosed in the present application, as well as in single-molecule peptide sequencing in general.

[0154] Example 10. Tolerance of aminopeptidase T from T. aquaticus to organic solvents. For the detection of a single peptide of a polypeptide immobilized on a surface, it is extremely important that the polypeptide is completely denatured and no longer possesses any secondary structure. It is well known in the art that this can be achieved by solvents (such as methanol) or high temperatures. When these stringent conditions are required, the aminopeptidase used should be resistant to solvents and / or high temperatures. Next, to demonstrate that TaqAPT is active at a temperature of 80°C, we also investigated whether the T. aquaticus-derived aminopeptidase described herein is resistant to organic solvents. As shown here, T. aquaticus-derived aminopeptidase T remained completely active up to 50% methanol, 33% acetonitrile, and 33% ethanol, demonstrating that the enzyme is quite resistant to organic solvents (Figure 12A). Despite higher organic solvent concentrations, activity can still be detected, even if it is lower. Analysis of the enzyme by circular dichroism (CD) in 0% methanol versus 50% methanol revealed no structural differences (Figure 12B). Furthermore, the enzyme appears to be fully active in deionized (MS-grade) water. This may be advantageous when the enzyme is to be used in ultra-sensitive chip technologies (e.g., electrical biosensors such as field-effect transistors) (Figure 12C).

[0155] Example 11. Site-specific N-terminal labeling of aminopeptidase T from T. aquaticus. Recombinant Taq aminopeptidase T containing extra N-terminal cysteine ​​was incubated overnight with a fluorescent maleimide-DyLight650 probe in the presence of the reducing agent TCEP (10 mM). Maleimide-DyLight650 was added at equimolar concentrations as well as at excess molar concentrations of 10×, 100×, and 1000×. After separation of the aminopeptidase by SDS-PAGE, it was visualized using Coomassie staining, and fluorescence was used to evaluate the DyLight650-labeled protein. Figure 13A shows that the aminopeptidase is labeled with DyLight650. Furthermore, an L-leucine-p-nitroanilide assay demonstrates that the labeled aminopeptidase does not lose its function (Figure 13B).

[0156] Example 12. Labeling of aminopeptidases Two sensor options are used for reading out the sequencing step and, consequently, detecting the enzyme "on-time" value: optical measurement and potentiometric measurement (Figure 2). Optical labeling of aminopeptidase is shown in Example 11. Alternative optical readout strategies with fluorescent probes may be on aminopeptidase and peptide substrates to measure on-time with fluorescence resonance energy transfer (FRET) and to detect good cleavage events.

[0157] Given that the concept is based on a single molecule, predictable and specific labeling of the enzyme is required. To have reasonable coverage of the human proteome, up to 10 9 This requires readings (for example, 10,000 copies of a protein need to be read). 4With a dynamic range of , and with least 10 readouts required for least abundant proteins (Geiger et al. 2012 Mol Cell Proteomics 11:M111.014050). Single-molecule detection can be achieved with zero-mode waveguides (ZMWs) (such as those used in single-molecule DNA sequencing following Rhoads and Au (2015 Genomics Proteomics Bioinformatics 13:278-289)). On the other hand, readouts from potentiometric measurements require a charged probe as it affects the potential of the field-effect transistor (FET).

[0158] For site-specific labeling of aminopeptidases, a one-step chemical modification is performed at the N-terminus of the enzyme. The enzyme is specifically modified with a pyridinecarboxaldehyde derivative with a primary amine at the N-terminus, as described by MacDonald et al. (2015 Nat Chem Biol 11:326-331). The cysteine ​​added to both aminopeptidases (see Examples 5 and 6) is specifically modified with an aldehyde probe or a maleimide probe when no other (surface-exposed) cysteine ​​is present (Gunnoo and Madder 2016 Chembiochem 17:529-553). For on-time potentiometric measurements, the need to add charge to the enzyme depends on its isoelectric point (pI). If necessary, the net charge of the enzyme can be altered by introducing a charged probe through (site-specific) modification or by unilateral neutralization of positively or negatively charged residues (e.g., formylation of lysine).

[0159] Example 13. Identification or categorization of N-terminal amino acids of single-molecule peptides by measuring aminopeptidase residence time. The use of aminopeptidase T derived from T. aquaticus was further validated in an independent experimental setup. A series of synthetic peptide substrates having identical primary structures except for the N-terminal amino acid were immobilized, and the residence time of labeled aminopeptidase T from T. aquaticus was then measured using TIRF microscopy. The peptide substrates have the following overall structure: X-DGGNNGGK(fluo)GGK(dbco / mal / nhs), where the C-terminal lysine in the substrate has a DBCO group, maleimide group, or N-hydroxysuccinimide group attached to its side chain to immobilize the substrate on the surface. The second lysine has a fluorescent group attached to its side chain to precisely target the single molecule substrate on its surface. The N-terminus will have a variable amino acid (X) or a varying sequence thereof, followed by an aspartic acid residue that acts as a brake on the reaction (proceeded by). Aminopeptidase T derived from T. aquaticus is thought to be extremely low in activity against aspartic acid residues.

[0160] After immobilizing the peptide substrate and determining the location of the single-molecule substrate in the field of view, fluorescently labeled aminopeptidase is added, and the enzyme residence time at the substrate location is continuously measured. Enzyme-substrate binding kinetics (K M ) and substrate cleavage kinetics (k cat Considering that both enzyme-substrate "on-off" events and substrate cleavage depend on the identity of the N-terminal amino acid, the identity of the N-terminal amino acid can be derived or categorized by measuring the number of enzyme-substrate "on-off" events and the total time until the substrate is cleaved (Figure 14). Verification of substrate cleavage is derived from the measurable change in the frequency of enzyme-substrate "on-off" events before and after cleavage (bottom of Figure 14). When thermophilic aminopeptidases (e.g., T. aquaticus aminopeptidase T) are used at temperatures far below the optimal temperature, the cleavage kinetics are significantly reduced, resulting in an increase in the number of enzyme-substrate "on-off" events.

[0161] As described in detail in the present application, a single aminopeptidase may be used to capture both "on-off" events on the N-terminal amino acid and to cleave the N-terminal amino acid. Alternatively, a combination of two different aminopeptidase enzymes may also be used. Or a combination of aminopeptidase and a chemical N-terminal amino acid binder / cleaver may be used (see description).

[0162] Example 14. Edman degradation reaction kinetics depend on the N-terminal amino acid residue. Next, we surprisingly found that Edman degradation chemistry can be employed for single-molecule sequencing of immobilized peptides. First, proteins are immobilized on a surface through their C-terminuses, and then amino acids are continuously cleaved via Edman degradation chemistry.

[0163] The N-terminal coupling reaction of Edman's reagent is independent of the identity of the N-terminal amino acid, although the speed of the cleavage reaction depends on it. Therefore, sequence information can be obtained by monitoring the cleavage reaction time on each subsequent N-terminal amino acid. To avoid problems related to N-terminal modification, the immobilized protein is first proteolyzed (e.g., with trypsin) to leave a C-terminal polypeptide with a freely accessible N-terminus. The chemical reaction time can be monitored using a traceable ITC agent. For example, sulfophenyl isothiocyanate carries a negative charge useful for electrical measurements. Alternatively, azidophenyl isothiocyanate, which can be labeled via click chemistry, can be used (charged group, fluorescent probe). Finally, the reaction can be carried out in a high-organic solvent that allows for thorough structural denaturation and reaction control.

[0164] To check the spontaneous cleavage activity of Edman's reagent 4-sulfophenyl isothiocyanate (SPITC) on various amino acid p-nitroanilide substrates, 5 μl of 24 mM amino acid p-nitroanilide substrate (in methanol) and 240 mM SPITC (in water) were added to 70 μl of 300 mM triethanolamine (in 50% acetonitrile (pH 9)) and incubated at 40°C for 30 min. Endpoint activity measurements revealed that the tested amino acid substrates were divided into those that were cleaved quickly (L, M, Y, R, F) and those that were cleaved slowly (D, P) (Figure 16A).

[0165] Time-rate assays demonstrate the difference in reaction kinetics for a series of identical substrates (Figure 16B). Importantly, this spontaneous cleavage activity was also observed under conditions not typically used during the cleavage step of classical Edman degradation. In classical Edman degradation, ITC coupling is achieved under weakly alkaline conditions (pyridine, trimethylamine, N-methylpiperidine), and amino acid cleavage is achieved under acidic conditions (trifluoroacetic acid). Here, both ITC coupling and amino acid cleavage are achieved under weakly alkaline conditions (triethanolamine).

[0166] Experimental Procedure E. coli transformation Thaw chemocompetent E. coli BL21 (DE3) cells (NEB) on ice, add 100 ng plasmid DNA, and keep on ice for 30 minutes. Incubate in a warm water bath at 42°C for 1.5 minutes, then place on ice for 10 minutes. Add 1 ml of LB medium to a vial, secure it to a shaker with tape, and leave at 37°C for 1 hour (let it rest) (LAF). Seed on an LB-Kan agar plate (50 μg / ml kanamycin) and grow overnight at 37°C.

[0167] Collection of cultured material (picking) Collect colonies from the plate (store the plate at 4°C) and add the colonies to 10 ml of liquid TB medium + 50 μg / ml Kan. Grow overnight at 37°C, prepare 500 μl aliquots (+ 500 μl glycerol), and store at -80°C.

[0168] cloning For cloning, the pET-24b(+) plasmid was used (Figure 9).

[0169] Ni-NTA purification Load 200 μl of lysate (in PBS + 10 mM imidazole) onto a pre-equilibrated Qiagen Ni-NTA spin column and centrifuge at 100 × g for 5 min. Wash the spin column 3 × with 500 μl PBS + 20 mM imidazole and elute with 500 μl PBS + 250 mM imidazole.

[0170] heat treatment The sample is heated at 80°C for 30 minutes and then centrifuged at full speed for 10 minutes.

Claims

1. A method for identifying the N-terminal amino acid of a polypeptide, wherein the polypeptide is immobilized on a surface either via its C-terminus or via the peptide portion of the C-terminus to the first peptide bond of the polypeptide. The method is as follows: a) Contacting the polypeptide immobilized on the surface with a cleavage inducer, wherein the cleavage inducer is a modified catalytically active aminopeptidase that binds to and cleaves an N-terminal amino acid from the polypeptide; b) Measuring the residence time of the cleavage inducer on the N-terminal amino acid; c) Identifying the N-terminal amino acid by comparing the measured residence time with a series of residence time reference values ​​characteristic of the cleavage inducer and a series of N-terminal amino acids. The method comprising, wherein the N-terminal amino acid is selected from the list consisting of Leu, Met, Tyr, Arg, Pro, Gly, Lys, Ala, and Val.

2. A method for obtaining sequence information of a polypeptide, wherein the polypeptide is immobilized on a surface via its C-terminus, and the method is as follows: a) Contacting the polypeptide immobilized on the surface with a cleavage inducer, wherein the cleavage inducer is a modified catalytically active aminopeptidase that binds to and cleaves an N-terminal amino acid from the polypeptide; b) Measuring the residence time of the cleavage inducer on the N-terminal amino acids of the polypeptide immobilized on the surface; c) The cleavage inducer is made capable of cleaving the N-terminal amino acid; d) Identifying the N-terminal amino acid by comparing the measured residence time with a series of residence time reference values ​​characteristic of the cleavage inducer and a series of N-terminal amino acids; e) Repeat steps a) to d) at least once, or repeat steps b) to d) at least once. The method comprising, wherein the N-terminal amino acid is selected from the list consisting of Leu, Met, Tyr, Arg, Pro, Gly, Lys, Ala, and Val.

3. The residence time is the time it takes for the N-terminal amino acid to be removed. The method according to claim 1 or 2, wherein the N-terminal amino acid is identified by comparing the time taken with a series of reference values ​​for various amino acids.

4. The method further includes the step of determining the cleavage of the N-terminal amino acid by measuring the optical, electrical, or plasmon signal of a polypeptide immobilized on its surface, The method according to any one of claims 1 to 3, wherein the difference in the optical signal, electrical signal, or plasmon signal represents the cleavage of the N-terminal amino acid.

5. In addition, the polypeptide immobilized on the surface is brought into contact with one or more N-terminal amino acid-binding proteins. The method according to any one of claims 1 to 4, wherein the kinetics of the binding event of the one or more binding proteins to the N-terminal amino acid identifies the N-terminal amino acid or provides further information thereto.

6. The first step of polypeptide modification is included, or the polypeptide modification state is present during one or more steps in any one of claims 1 to 5. The method according to any one of claims 1 to 5, wherein the aminopeptidase is a thermophilic and / or solvent-resistant aminopeptidase.

7. The method according to any one of claims 1 to 6, wherein the N-terminal amino acid is derivatized.

8. The method according to claim 7, wherein the derivatized N-terminal amino acid is derivatized with CITC or SPITC.

9. The N-terminal amino acid is identified, or The method according to any one of claims 1 to 8, wherein the polypeptide immobilized on the surface is sequenced at the single-molecule level.

10. The method according to any one of claims 1 to 9, wherein the residence time is measured optically, electrically, or by plasmons.

11. The method according to any one of claims 1 to 10, wherein the polypeptide is immobilized on an active detection surface.

12. The method according to claim 11, wherein the active detection surface is a gold surface to which the polypeptide is chemically coupled, or an amid-, carboxyl-, thiol-, or azide-functionalized surface.

13. The method according to any one of claims 1 to 12, wherein the residence time corresponds to the binding event between the aminopeptidase and the N-terminal amino acid.

14. The method according to any one of claims 1 to 13, wherein the aminopeptidase binds to and dissociates from the N-terminal amino acid at least once before the aminopeptidase binds to and cleaves the N-terminal amino acid from the polypeptide.

15. The method according to claim 14, further comprising measuring multiple residence times corresponding to multiple binding events between an aminopeptidase and an N-terminal amino acid.

16. The method of claim 15, wherein the identity of the N-terminal amino acid is determined by comparing the average of multiple residence times with a series of residence time reference values.

17. A method for identifying the N-terminal amino acid of a polypeptide, A polypeptide is brought into contact with a composition comprising one or more terminal amino acid-binding proteins and a cleavage inducer, wherein the cleavage inducer is a modified catalytically active aminopeptidase; Measuring the residence time of one or more terminal amino acid-binding proteins on the N-terminal amino acid, where each residence time corresponds to a binding event between the terminal amino acid-binding protein and the N-terminal amino acid; and The N-terminal amino acid is identified by comparing the measured residence time with a series of residence time reference values ​​characteristic of one or more terminal amino acid-binding proteins and a series of N-terminal amino acids. The method comprising, wherein the N-terminal amino acid is selected from the list consisting of Leu, Met, Tyr, Arg, Pro, Gly, Lys, Ala, and Val.

18. The method according to claim 17, wherein the residence time corresponds to the binding / dissociation cycle between one or more terminal amino acid-binding proteins and an N-terminal amino acid.

19. The method according to claim 18, wherein the residence time corresponds to at least three binding / dissociation cycles.

20. The method according to claim 18, wherein the residence time corresponds to at least 10 binding / dissociation cycles.

21. The method according to claim 18, wherein the residence time corresponds to at least 20 binding / dissociation cycles.

22. The method according to any one of claims 17 to 21, wherein the terminal amino acid-binding protein comprises optical labeling, electrical labeling, or plasmon labeling.

23. The method according to any one of claims 17 to 22, wherein the residence time is measured optically, electrically, or by plasmons.

24. The method according to any one of claims 17 to 23, wherein the terminal amino acid-binding protein is a catalytically inactive aminopeptidase.

25. The method according to any one of claims 17 to 24, wherein the polypeptide is immobilized on a surface.

26. The method according to claim 25, wherein the polypeptide is immobilized on the surface via its C-terminus or via the peptide portion of the polypeptide between the C-terminus and the first peptide bond.

27. The method according to any one of claims 17 to 26, further comprising using a cleavage inducer in the composition to cleave an N-terminal amino acid from a polypeptide.

28. The method according to claim 27, wherein the cleavage inducer is a thermophilic and / or solvent-resistant aminopeptidase, and the polypeptide is subject to conditions for causing the cleavage inducer in the composition to cleave an N-terminal amino acid from the polypeptide.

29. The method according to claim 27, wherein the terminal amino acid-binding protein binds to the N-terminal amino acid at a rate faster than the time required for the cleavage-inducing agent in the composition to cleave the N-terminal amino acid from the polypeptide.

Citation Information

Patent Citations

  • Macromolecular analysis using nucleic acid encoding

    JP2019523635A

  • Molecules and methods for iterative polypeptide analysis and processing

    US20140273004A1

  • Protein sequencing methods and reagents

    WO2017063093A1