Nanopore peptide profiling and sequencing by hydrolysis
The nanopore-based method addresses the limitations of existing protein sequencing techniques by using a nanopore with a sensing moiety to measure ionic current changes for high-resolution peptide sequencing and modification detection, achieving efficient peptide sequencing and purity analysis.
Patent Information
- Application Number
- PCT/CN2025/116383
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-08-25
- Filing Date
- 2025-08-22
- Publication Date
- 2026-03-05
AI Technical Summary
Current protein sequencing methods, such as Edman degradation and mass spectrometry, face limitations in speed, read length, throughput, and the identification of short or long peptides, particularly for low-abundance proteins and peptides with post-translational modifications, lacking a robust single-molecule sequencing method for efficient peptide sequencing.
A nanopore-based method using a nanopore incorporating a sensing moiety, such as Ni2+, to interact with peptides, measuring ionic current changes during translocation for peptide characterization, including proteinogenic components and post-translational modifications, and employing machine learning for sequence reconstruction.
Enables high-resolution peptide sequencing, identifying amino acid sequences, modifications, and purity of polypeptides, overcoming limitations of existing methods by providing sensitivity to low-abundance proteins and direct sensing of PTMs.
Smart Images

Figure CN2025116383_05032026_PF_FP_ABST
Abstract
Description
NANOPORE PEPTIDE PROFILING AND SEQUENCING BY HYDROLYSISFIELD OF THE INVENTION
[0001] The present invention relates to a method for identifying a peptide using nanopore.BACKGROUND OF THE INVENTION
[0002] Proteins, which consist of amino acids joined together in sequence, are key building blocks of life and play pivotal roles in the functioning of cells in all living organisms (1) . The three-dimensional structures of proteins, which are determined by the genetically encoded protein sequence, result in the stunning diversity of protein functions (2) . Protein sequencing, a fundamental step in proteomic analysis, thus holds immense significance and is useful in cellular process monitoring, biomarker identification and disease diagnosis (3, 4) . Despite rapid advancements of techniques for nucleic acid sequencing, the development of protein sequencing methods significantly lags behind (5) . At present, protein sequencing is achieved by either Edman degradation (6) or a variety of mass spectrometry (MS) strategies (2, 7) . Despite their indispensable roles in proteomic analysis, clear technical drawbacks still exist. Edman degradation only reports sequence of the N-terminus of protein and is generally limited in the speed, read length and the throughput. On the other side, the mass spectrometers are generally disadvantageous for the identification of extremely short or long peptides, leading to coverage gaps in protein sequence (8, 9) . The presence of PTMs and incomplete enzymatic hydrolysis products may also introduce ambiguity during spectra interpretation (10-12) . For both methods, due to the lack of amplification techniques for protein, identification of low-abundance proteins in a complex sample is still challenging. It is widely anticipated that a robust single-molecule protein sequencing method may fully address these issues by providing an improved sensitivity for low-abundance proteins and a direct sensing capacity for PTMs. Existing approaches towards single molecule protein sequencing include single molecule fluorescence (13, 14) , tunnelling current (15, 16) , DNA nanotechnology (17, 18) and nanopore (19, 20) .
[0003] Inspired by the success of nanopore nucleic acid sequencing (21, 22) , it is widely anticipated that with a similar principle, protein or peptide may as well be sequenced by nanopore (23) . According to results of peptide translocation previously demonstrated by fragaceatoxin C (FraC) (24, 25) , aerolysin (AeL) (26, 27) and cytotoxin K (CytK) (28) nanopores, the sensing information gathered during peptide translocation is insufficient for sequencing. To achieve nanopore peptide sequencing, two main strategies, including peptide strand sequencing and exopeptidase sequencing (19, 20, 29-37) , have been proposed. Reported approaches towards peptide strand sequencing include peptide sequencing by nanopore induced phase shift sequencing (NIPSS) (19, 20) and retarded peptide translocation modulated by ClpX unfoldase (29) or electroosmotic flow (EOF) (30-32) . However, due to the limited spatial resolution of the pore constriction, it is hard to directly resolve individual amino acids without interferences of neighbouring amino acids, posing a great challenge for peptide sequence decoding. On the other side, approaches towards exopeptidase sequencing include the construction of a proteasome conjugated nanopore (33) and the engineering of a nanopore for direct identification of proteinogenic amino acids (34, 35) . However, sequential nanopore reading of exopeptidase-cleaved amino acid has not yet been demonstrated, due to the technical complexity required for this approach. Different from above two strategies, Oukhaled et al. and Wu et al. reported nanopore discrimination of peptides with sequence substitutions of amino acids (36) or chemically conjugated amino acids (37) . Though a high resolution of nanopore peptide discrimination is well demonstrated, a strategy for efficient peptide sequencing following these principles was not clearly demonstrated (36, 37) . To date, no nanopore techniques have demonstrated a satisfying peptide sequencing performance suitable for proteomics investigations.
[0004] BRIEF SUMMARY OF THE INVENTION
[0005] In one aspect, the present invention provides a method of characterizing one or more target analytes in a sample, comprising:
[0006] (i) providing a nanopore incorporating at least one sensing moiety which is capable of interacting with a peptide, and causing the target analyte to contact the sensing moiety in a region that allows the target analyte to be detected based on a change of an ionic current across the nanopore; and
[0007] (ii) measuring the ionic current as the target analyte migrates to provide a representative current pattern of the target analyte, wherein the representative current pattern indicates a characteristic of the target analyte;
[0008] wherein the one or more target analytes comprise a peptide; wherein the sensing moiety is a metal ion.
[0009] In some embodiments, the sensing moiety is Ni2+, Cu2+, Co2+, Zn2+, Cd2+, Ag+, Mn2+, Pb2+, Fe2+ or Fe3+; preferably, the sensing moiety is Ni2+.
[0010] In some embodiments, the sensing moiety is attached to the nanopore by a ligand; preferably, the ligand is a metal ion chelating agent; more preferably, the ligand is nitrilotriacetic acid (NTA) .
[0011] In some embodiments, the nanopore is a protein nanopore comprising a reactive amino acid to which the sensing moiety is attached.
[0012] In some embodiments, the reactive amino acid is selected from cysteine, methionine, histidine and lysine.
[0013] In some embodiments, the protein nanopore is selected from alpha-hemolysin (α-HL) , Mycobacterium smegmatis porin A (MspA) , Aerolysin (Ael) , curli production assembly / transport component (CsgG) , outer membrane porin F (OmpF) , Cytolysin A (ClyA) , ferric hydroxamate uptake component A (FhuA) , Fragaceatoxin C (FraC) , Cytotoxin K (CytK) , Pleurotolysin A (PlyA) / Pleurotolysin B (PlyB) , Curli production assembly / transport component CsgG (CsgG) , Phi29 connector protein, and any variant thereof; preferably, the protein nanopore is a variant of MspA.
[0014] In some embodiments, the protein nanopore is a heterogeneous protein nanopore in which one or more but not all monomers incorporate the sensing moiety and the other monomers do not incorporate the sensing moiety.
[0015] In some embodiments, the protein nanopore is a variant of MspA which has the reactive amino acid residue located at a position selected from 83-111, preferably 90, 91, 92 and 93.
[0016] In some embodiments, the protein nanopore has a mutation of N90C, N90M, N91M or N91C on one or more monomers compared to a wild-type MspA or a variant thereof; preferably, the variant is M2 MspA.
[0017] In some embodiments, the peptide comprises 2-50 amino acids in length.
[0018] In some embodiments, the peptide comprises a N atom in the N-terminal amino group and the N atom in the N-terminal amide bond moiety.
[0019] In some embodiments, the peptide comprises one or more amino acids selected from histidine, lysine, cysteine, methionine and proline.
[0020] In some embodiments, the one or more target analytes comprises one or more peptides and / or one or more amino acids, each of which optionally and independently and have a modification; preferably, the modification is a post-translational modification; preferably, the modification is acetylation, phosphorylation, methylation, or any combination thereof.
[0021] In some embodiments, the characteristic comprises the identity of the target analyte, the presence or absence of the target analyte, the structure of the target analyte, the conformation of the target analyte, the sequence of the target analyte, the modification of the target analyte, the confirmation of the target analyte, the fingerprint of the target analyte, or any combination thereof.
[0022] In some embodiments, the sample is a target polypeptide sample the one or more target analytes are comprised in the proteinogenic components in the target polypeptide sample, and the method comprises:
[0023] (i) providing a nanopore incorporating at least one sensing moiety which is capable of interacting with a peptide, and causing each proteinogenic component of the target polypeptide sample to contact the sensing moiety in a region that allows each proteinogenic component of the target polypeptide sample to be detected based on a change of an ionic current across the nanopore;
[0024] (ii) measuring the ionic current as each proteinogenic component of the target polypeptide sample migrates to provide a representative current pattern of the target polypeptide sample, and obtain a fingerprint of the target polypeptide sample; and
[0025] (iii) comparing the fingerprint of the target polypeptide sample with a reference fingerprint of a reference polypeptide and detecting the purity of the target polypeptide sample.
[0026] In some embodiments, if the fingerprint of the target polypeptide sample comprises one or more additional results that are not contained in the fingerprint of the reference polypeptide, it means that the target polypeptide sample contains impurities.
[0027] In some embodiments, the target polypeptide sample is a sample obtained by solid-phase peptide synthesis method.
[0028] In another aspect, the present invention provides a method of characterizing a target polypeptide, comprising:
[0029] (a) fragmenting the target polypeptide to obtain a peptide fragment mixture;
[0030] (b) characterizing the peptide fragment mixture using the method of any one of the preceding methods, to determine the characteristic of the peptide fragment mixture; and
[0031] (c) determining the characteristic of the target polypeptide based on the characteristic of the peptide fragment mixture.
[0032] In some embodiments, the target polypeptide has or hasn’ t a modification; preferably, the modification is a post-translational modification; preferably, the modification is acetylation, phosphorylation, methylation, or any combination thereof.
[0033] In some embodiments, the target polypeptide is fragmented by endopeptidase digestion.
[0034] In some embodiments, the step (b) comprises causing each component of the peptide fragment mixture to contact the sensing moiety in a region that allows each component of the peptide fragment mixture to be detected based on a change of an ionic current across the nanopore, measuring an ionic current as each component of the peptide fragment mixture migrates to provide a representative current pattern of the peptide fragment mixture, comparing the representative current pattern of the peptide fragment mixture to reference current patterns in a reference database, and determine the characteristic of the peptide fragment mixture based on the comparison.
[0035] In some embodiments, the method further comprises an assisting-assay on the target polypeptide comprising hydrolyzing the target polypeptide into an amino acid mixture using an exopeptidase, determining the identity of each amino acid in the mixture, selecting reference current patterns of amino acids and peptides containing the identified amino acid to form the reference database.
[0036] In some embodiments, the characteristic of the target polypeptide is the identity of the target polypeptide, the presence or absence of the target polypeptide, the amino acid sequence of the target polypeptide, the amino acid alteration in the target polypeptide, the modification in the target polypeptide, the fingerprint spectrum of the target polypeptide, the purity of the target polypeptide, or any combination thereof.
[0037] In some embodiments, the step (b) includes identifying the amino acid sequence of each peptide fragment in the peptide fragment mixture; and wherein the step (c) includes assembling the amino acid sequence of each peptide fragment in the peptide fragment mixture to the amino acid sequence of the target polypeptide.
[0038] In some embodiments, the target polypeptide is fragmented by multiple fragmentation operations to obtain multiple peptide fragment mixtures, and the multiple peptide fragment mixtures are separately subjected to step (b) to identify the amino acid sequences of each peptide fragment in each peptide fragment mixture, and these sequences are subsequently subjected to step (c) to be assembled to the amino acid sequence of the target polypeptide.
[0039] In some embodiments, the multiple fragmentation operations are multiple digestions by the same endopeptidase or by different endopeptidases respectively.
[0040] In some embodiments, the step (b) includes identifying the fingerprint of the target polypeptide; and wherein the step (c) includes comparing the fingerprint of the target polypeptide to a fingerprint of a reference polypeptide, and determining whether the target polypeptide is identical to or different from the reference polypeptide, or whether the target polypeptide has amino acid alteration and / or modifications relative to the reference polypeptide based on the comparison.
[0041] In some embodiments, the fingerprint spectra of the reference polypeptide and the target polypeptide were obtained by performing the same fragmentation operation and the same nanopore detection operation on the reference peptide and the target peptide.
[0042] In another aspect, the present invention provides a method of characterizing a target polypeptide, comprising:
[0043] (1) preparing a target polypeptide sample including the target polypeptide;
[0044] (2) preparing different shortened polypeptide samples, wherein each shortened polypeptide sample including a shortened target polypeptide by truncating the target polypeptide, and lengths of the shortened target polypeptides in different shortened polypeptide samples are different;
[0045] (3) characterizing the different shortened polypeptides using the method as mentioned above, to determine the characteristic of the shortened target polypeptides separately; and
[0046] (4) determining the characteristic of the target polypeptide based on the characteristic of the different shortened polypeptides.
[0047] In some embodiments, the different shortened polypeptide samples are prepared by stepwise shortening the target polypeptide.
[0048] In some embodiments, the different shortened polypeptide samples are prepared by stepwise shortening the target polypeptide from N-terminus or C-terminus of the target polypeptide.
[0049] In some embodiments, the different shortened polypeptide samples are prepared by Edman degradation.
[0050] In some embodiments, the characteristic of the target polypeptide is the identity of the target polypeptide, the presence or absence of the target polypeptide, the amino acid sequence of the target polypeptide, the amino acid alteration in the target polypeptide, the modification in the target polypeptide, the fingerprint spectrum of the target polypeptide, the purity of the target polypeptide, or any combination thereof.
[0051] In another aspect, the present application provides a use of the method as mentioned above for determining characteristic of a neoantigen.
[0052] In some embodiments, the characteristic of the neoantigen is the identity of the neoantigen, the presence or absence of the neoantigen, the amino acid sequence of the neoantigen, the amino acid alteration in the neoantigen, the modification in the neoantigen, the fingerprint spectrum of the neoantigen, the purity of the neoantigen, or any combination thereof.BRIEF DESCRIPTION OF THE DRAWINGS
[0053] Figure 1. Nanopore peptide sensing. (a) The workflow. A pool containing a variety of synthetic peptides with known sequences is first established. Each peptide in the pool is respectively measured by MspA-NTA-Ni and the resulting nanopore events are acquired to build the database, with which a custom machine learning algorithm is developed for further peptide classification. (b) Representative peptide events acquired with MspA-NTA-Ni. Here, results of thirty different types of peptides are demonstrated. The corresponding peptide sequence is placed above each representative event. Some peptides such as GP and GH can report multiple types of peptide events. (c) The event scatter plot of ΔI versus SD generated by results acquired with peptides in b. 100 events acquired with each peptide are used to generate the plot, according to which, all peptides are clearly distinguishable. (d) A zoomed-in view of the area in c, which is marked with a box. (e) Representative traces acquired during simultaneous sensing of DQ, IK, NKR, LRG and SLR. All five peptides were simultaneously added to the cis chamber. During the measurement, the final concentrations of DQ, LRG and SLR were 0.2 mM and the final concentrations of NKR and IK were 0.3 mM. The event identities were automatically predicted by the trained quadratic SVM model and marked above corresponding events. All measurements were performed by MspA-NTA-Ni (Methods) . I0 represents the open pore current of MspA-NTA-Ni.
[0054] Figure 2. Nanopore peptide profiling. (a) The workflow of nanopore peptide profiling. A polypeptide is first fragmented by endopeptidase hydrolysis. Afterwards, the generated fragments are sensed by MspA-NTA-Ni to produce corresponding raw data. The event features of all raw data are extracted, with which the peptide profiling result is produced. (b-e) Nanopore profiling of representative peptides. Four bioactive peptides, including angiotensin III (7 aa) , secretin (27 aa) , glucagon (29 aa) and ACTH (39 aa) , were used for the demonstration. Each peptide was separately treated with trypsin (~10000 BAEE U / mg) at an enzyme / peptide ratio of 1: 50 (w / w) at 37℃ for 8 h (Methods) . The hydrolysates were measured by MspA-NTA-Ni for peptide profiling. Each peptide reports a unique peptide profile in corresponding kernel density plot results.
[0055] Figure 3. Nanopore peptide sequencing by hydrolysis. (a) The event scatter plot of ΔI versus SD of events acquired with the exopeptidase hydrolysate of DQARNKRSLRGYIK (n=888) . Leucine aminopeptidase (LAP) was used to perform the exopeptidase treatment. All events were recognized by the previously trained machine learning classifier. Eleven populations of events respectively corresponding to R, K, N, S, A, I, Y, G, Q, L and D were identified. (b) Database simplification. The integrated database, which contains nanopore events of thirty peptides and twenty-four amino acids, is further simplified by the amino acid composition results demonstrated in a. To simplify later peptide identification, all unidentified amino acids and relevant peptides were excluded from the database. (c, f) Representative events acquired with endopeptidase hydrolysate I (c, trypsin) and endopeptidase hydrolysate II (f, thermolysin) are demonstrated. (d, g) The event scatter plots of ΔI versus SD of results acquired with endopeptidase hydrolysate I (d, n=807) and endopeptidase hydrolysate II (g, n=970) . All events are identified by the quadratic SVM model derived from the simplified peptide database. Clusters of events respectively corresponding to SLR, NKR, GYIK and DQAR are identified in d. Those corresponding to LRGY, LRG, DQARNKRS, IK and Y are identified in g. (e, h) The event appearance frequency of results acquired with endopeptidase hydrolysate I (e) and endopeptidase hydrolysate II (h) . All identified peptides are marked with pentacles. (i) Sequence reconstruction of peptide. According to above results, the sequence of the original peptide is determined to be DQARNKRSLRGYIK.
[0056] Figure 4. Analysis of peptides containing single-amino acid mutations. (a) The technique overview. Generally, peptides with a single amino acid difference can be discriminated by the demonstrated nanopore sequencing strategy. (b) Representative events acquired with the tryptic hydrolysate of DQARNKRSLRGYIK (reference peptide) . (c) The event scatter plot of ΔI versus SD of results acquired with the hydrolysate of the reference peptide. Four populations of events were identified by the previously trained quadratic SVM model. Its original sequence is also determined as described in Figure 3i. (d, f, h, j) Representative events acquired with the hydrolysates of D1F (d) , K6M (f) , I13L (h) and K14 deletion (j) . All peptides (b, d, f, h, j) were hydrolysed under the same condition. The lines on the peptide sequence (b, d, f, h, j) represent the trypsin cleavage sites. The mutation sites are highlighted. The events of mutated peptides are marked with dotted boxes. (e, g, i, k) The event scatter plots of ΔI versus SD of results acquired with the hydrolysate of D1F (e, n=1126) , K6M (g, n=693) , I13L (i, n=856) and K14 Deletion (k, n=537) . All events are identified by machine learning and labelled with color-coded dots. The event clusters corresponding to the mutated peptide fragment and the un-mutated counterpart are respectively marked with circles in each plot.
[0057] Figure 5. Analysis of peptide containing PTMs. (a-b) Representative events of (pS) LR, (pS) LRGYIK, NKR (pS) LR, GYI (acK) , SLRGYI (acK) and NKRSLRGYI (acK) (Table 5) . The event features of these peptides are added to the existing peptide database to create a larger peptide database containing post translational modifications (PTMs) . (c) Tryptic hydrolysis of pS8 peptide (DQARNKR (pS) LRGYIK) . The solid lines represent complete cleavage sites and the dashed line represents incomplete cleavage site. (d) Representative events acquired with the hydrolysate of pS8. (e) The event scatter plot of ΔI versus SD of results acquired with the hydrolysate of pS8 (n=701) . The event cluster corresponding to the unmodified peptide counterpart (SLR) is indicated by a grey circle. (f) Tryptic hydrolysis of DQARNKRSLRGYI (acK) (acK14) . (g) Representative events of results acquired with the hydrolysate of acK14. (h) The event scatter plot of ΔI versus SD of results acquired with the hydrolysate of acK14 (n=783) . The event cluster corresponding to the unmodified peptide counterpart (GYIK) is indicated by a grey circle. All events in d, e, g and h are identified by machine learning. The identified events are respectively marked with dots.
[0058] Figure 6. Database expansion. (a) Technique overview. MspA-NTA-Ni demonstrates a general compatibility with amino acids, peptides, amino acids containing PTMs, peptides containing PTMs and a variety of bioactive peptides, including cyclic peptides, aspartame and oxiglutatione (GSSG) . Their nanopore event features are collected and fed to the machine learning program for downstream sensing and sequencing applications. (b) The confusion matrix results generated by the previously trained quadratic SVM model. This model contains a total of seventy-three analyte classes including amino acids and peptides. A validation accuracy of 97.4%is achieved (Table 7) . Acknowledging the high resolution of MspA-NTA-Ni, the model can be further expanded or modularly established for specific applications while a satisfactory accuracy is still retained.
[0059] Figure 7. Peptide sensing by MspA-NTA-Ni. (a) Peptide sensing by MspA-NTA-Ni. Peptides with varying sequences, lengths, modifications and conformations such as linear and cyclic peptides can be sensed by MspA-NTA-Ni to report their characteristic event features. (b) The sensing mechanism. The sole nickel ion immobilized at the constriction of MspA-NTA-Ni serves to reversibly interact with the N-termini of peptide to report characteristic nanopore features.
[0060] Figure 8. Single molecule results of glycine peptides. (a-e) Representative traces respectively acquired with GG (a) , GGG (b) , GGGG (c) , GGGGG (d) and GGGGGG (e) using MspA-NTA-Ni. Representative events of peptides, which were extracted from the corresponding traces (marked by arrows) , were placed to the right of each source trace. I0 represents the open pore current of MspA-NTA-Ni. The event levels corresponding to peptide binding were marked with different bands. The measurements were carried out using MspA-NTA-Ni in a 1.5 M KCl buffer (1.5 M KCl, 10 mM CHES, pH 9) . A voltage of +100 mV was continually applied. Each analyte was separately added to the cis chamber with a final concentration of 1 mM. (f-j) The event scatter plots of ΔI versus SD respectively derived from events of GG (f) , GGG (g) , GGGG (h) , GGGGG (i) and GGGGGG (j) . 200 events of each peptide were used to generate the scatter plot.
[0061] Figure 9. A representative trace acquired with MspA-NTA-Ni. I0 represents the open pore current of MspA-NTA-Ni. Even without the presence of any analyte, the pore spontaneously reports short-residing spiky events (marked with asterisks without circle) and events with a lower amplitude (marked with asterisks in circle) . These events can be easily differentiated from events produced by amino acids and peptides. The measurements were carried out using MspA-NTA-Ni in a 1.5 M KCl buffer (1.5 M KCl, 10 mM CHES, pH 9) . A voltage of +100 mV was continually applied.
[0062] Figure 10. Summary of the nanopore events of glycine peptides. I0 represents the open pore current of MspA-NTA-Ni. The measurements were carried out using MspA-NTA-Ni in a 1.5 M KCl buffer (1.5 M KCl, 10 mM CHES, pH 9) . A voltage of +100 mV was continually applied. Each glycine peptide demonstrates a distinct kind of event features.
[0063] Figure 11. Single molecule sensing of glycylglycine (GG) and N-acetylglycylglycine (N-Ac-GG) . (a) The chemical structure of glycylglycine (GG) . (b) A representative trace acquired with GG using MspA-NTA-Ni (Methods) . Successive appearance of GG events was observed, as marked by the circles. Events caused by GG are generally consistent in the event blockage amplitude, indicating a strong interaction between the pore and peptide. (c) The chemical structure of N-acetylglycylglycine (N-Ac-GG) . N-Ac-GG has an acetyl group at the terminal amino group. (d) A representative trace acquired with N-Ac-GG using MspA-NTA-Ni (Methods) . No events with well-defined event features were observed, indicating that the terminal amino-N of peptide is critical in the generation of peptide events. Here, all measurements were carried out using MspA-NTA-Ni in a 1.5 M KCl buffer (1.5 M KCl, 10 mM CHES, pH 9) . A voltage of +100 mV was continually applied. Each analyte was separately added to the cis chamber with a final concentration of 1 mM.
[0064] Figure 12. Sensing of glycine peptides with M2 MspA. (a-e) Representative traces acquired with GG (a) , GGG (b) , GGGG (c) , GGGGG (d) and GGGGGG (e) using M2 MspA. All measurements were performed with an octameric M2 MspA in a 1.5 M KCl buffer (1.5 M KCl, 10 mM CHES, pH 9) . A voltage of +100 mV was continually applied. Each analyte was separately added to the cis chamber with a final concentration of 1 mM. No binding events were observed for GG, GGG, GGGG or GGGGG. However, some events were observed for GGGGGG, which might be generated by non-specific binding between nanopore and peptide. Since an octameric M2 MspA doesn’ t have any NTA-Ni adapter, these results further indicate that the Ni-NTA modification made to the pore constriction is critical in the generation of peptide events.
[0065] Figure 13. The definition of core event parameters. (a) A representative trace containing peptide events. Here, the measurement was carried out using MspA-NTA-Ni. GG is treated as the model analyte. I0 represents the open pore current of MspA-NTA-Ni. Ipep represents the blockage level of peptide. ΔI is defined as ΔI=Ipep-I0. SD is the standard deviation of the blockage level caused by peptide binding. toff is the event dwell time. (b) The histogram of ΔI. (c) The histogram of SD. The mean and the mean are respectively derived from results of Gaussian fitting, as described in b, c. (d) The histogram of toff. The mean toff (τoff) is derived from results of single-exponential fitting, following the equation of y=a*exp (-x / τ) .
[0066] Figure 14. The event scatter plot of ΔI versus SD of events acquired with peptides composed of different lengths of glycine. 200 events of each peptide were used to generate the plot. Five event clusters, respectively corresponding to different peptide types, are clearly seen in the plot.
[0067] Figure 15. Machine learning assisted glycine peptide identification. (a) The machine learning workflow. Briefly, 300 events of each glycine peptide were collected to form a dataset. Then, five event features, including ΔI, SD, toff, skew, and kurt were extracted to generate a feature matrix for model training. The cubic SVM model has demonstrated the highest cross-validation accuracy of 99.8%. It was then used to predict unclassified events. (b) The confusion matrix of glycine peptide classification generated by the cubic SVM model. (c) The event scatter plot of ΔI versus SD generated by results acquired during simultaneous sensing of GG, GGG, GGGG, GGGGG and GGGGGG (n=1480) . (d) The event scatter plot of ΔI versus SD labelled with machine learning prediction results. All events were identified by the trained cubic SVM model. (e) A representative trace acquired during simultaneous sensing of GG, GGG, GGGG, GGGGG and GGGGG. Experimentally, five types of peptides were simultaneously added to cis with a final concentration of 0.5 mM for each peptide. Events of different glycine peptides were automatically predicted by the previously trained cubic SVM model and marked with corresponding identities.
[0068] Figure 16. Simultaneous sensing of glycine and glycine peptides. (a) A representative trace acquired with glycine (G) . The final concentration of glycine in cis is 0.5 mM. I0 represents the open pore current of MspA-NTA-Ni. The blockage level of glycine events is marked with a dashed line. The glycine events were marked with circles. (b) A representative trace acquired by simultaneous sensing of G, GG, GGG, GGGG, GGGGG and GGGGGG. Events from glycine and glycine peptides can be easily discriminated based on the difference of their event features. The measurements were carried out using MspA-NTA-Ni in a 1.5 M KCl buffer (1.5 M KCl, 10 mM CHES, pH 9) . A voltage of +100 mV was continually applied. G, GG, GGG, GGGG, GGGGG and GGGGGG were simultaneously added to the cis chamber. The final concentration of each analyte was 0.5 mM. (c) The event scatter plot of ΔI versus SD acquired with glycine. 137 events were employed to generate the plot. (d) The event scatter plot of ΔI versus SD acquired by simultaneous sensing of G, GG, GGG, GGGG, GGGGG and GGGGGG. A total of 865 events were employed to generate the plot.
[0069] Figure 17. The comparison of events generated by glycine and glycine peptides. (a) Representative events respectively acquired with glycine and different glycine peptides. (b) The histograms of toff acquired with glycine or glycine peptide. The mean toff (τoff) was derived from results of single-exponential fitting, following the equation of y=a*exp (-x / τ) . According to the τoff results, glycine reports significantly longer-residing events than those of glycine peptides.
[0070] Figure 18. Nanopore events of DQ, DQAR, FQAR, DQARNKR, DQARNKRS, and ARNKRS. (a-f) Representative traces respectively acquired with DQ (a) , DQAR (b) , FQAR (c) , DQARNKR (d) , DQARNKRS (e) and ARNKRS (f) . Different peptide events are labeled with different color shadings. Zoomed-in views of representative peptide events are placed to the right of each corresponding trace. Each representative event is taken from the source trace, as indicated by arrow. I0 represents the open pore current of MspA-NTA-Ni. The measurements were carried out using MspA-NTA-Ni in a 1.5 M KCl buffer (1.5 M KCl, 10 mM CHES, pH 9) . A voltage of +100 mV was continually applied. Each peptide was separately added to the cis chamber. The final concentration of FQAR was 0.5 mM. The final concentrations of DQAR, DQARNKR or DQARNKRS were set at 1 mM. The final concentration of DQ was set at 1.5 mM. The final concentration of ARNKRS was set at 2 mM. (g-l) The event scatter plots of ΔI versus SD generated by events respectively acquired with DQ (g) , DQAR (h) , FQAR (i) , DQARNKR (j) , DQARNKRS (k) or ARNKRS (l) . 100 events for each condition were used to generate each scatter plot.
[0071] Figure 19. Nanopore events of GI, GL, GD, GP, and GH. (a-e) Representative traces respectively acquired with GI (a) , GL (b) , GD (c) , GP (d) and GH (e) . Different peptide events are labeled with different color shadings. Zoomed-in views of representative peptide events are placed to the right of each corresponding trace. Each representative event is taken from the source trace, as indicated by arrow. I0 represents the open pore current of MspA-NTA-Ni. The measurements were carried out in a 1.5 M KCl buffer (1.5 M KCl, 10 mM CHES, pH 9) . A voltage of +100 mV was continually applied. Each peptide was separately added to the cis chamber with a final concentration of 1 mM. (f-j) The event scatter plots of ΔI versus SD generated by events respectively acquired with GI (f) , GL (g) , GD (h) , GP (i) , and GH (j) . 100 events for each condition were used to generate each scatter plot. For GH and GP, multiple types of events are observed (d, i, e, j) .
[0072] Figure 20. Nanopore events of GYI, GYIK, GYLK, and IK. (a-d) Representative traces respectively acquired with GYI (a) , GYIK (b) , GYLK (c) , and IK (d) . Different peptide events are labeled with different color shadings. Zoomed-in views of representative peptide events are placed to the right of each corresponding trace. Each representative event is taken from the source trace, as indicated by the arrow. I0 represents the open pore current of MspA-NTA-Ni. The measurements were carried out in a 1.5 M KCl buffer (1.5 M KCl, 10 mM CHES, pH 9) . A voltage of +100 mV was continually applied. Each peptide was separately added to the cis chamber. The final concentration of GYI was 0.5 mM. The final concentrations of GYI and GYIK were 1 mM. The final concentration of IK was 1.5 mM. (e-h) The event scatter plots of ΔI versus SD generated by events respectively acquired with GYI (e) , GYIK (f) , GYLK (g) , and IK (h) . 100 events for each condition were used to generate each scatter plot.
[0073] Figure 21. Nanopore events of GR, GGR, RG, LRG, and LRGY. (a-e) Representative traces respectively acquired with GR (a) , GGR (b) , RG (c) , LRG (d) and LRGY (e) . Different peptide events are labeled with different color shadings. Zoomed-in views of representative peptide events are placed to the right of each corresponding trace. Each representative event is taken from the source trace, as indicated by arrow. I0 represents the open pore current of MspA-NTA-Ni. The measurements were carried out in a 1.5 M KCl buffer (1.5 M KCl, 10 mM CHES, pH 9) . A voltage of +100 mV was continually applied. Each peptide was separately added to the cis chamber. The final concentrations of LRG and LRGY were 0.5 mM. The final concentrations of GGR and RG were 1 mM. The final concentration of GR was 1.5 mM. (f-j) The event scatter plots of ΔI versus SD generated by events respectively acquired with GR (f) , GGR (g) , RG (h) , LRG (i) and LRGY (j) . 100 events for each condition were used to generate each scatter plot.
[0074] Figure 22. Nanopore events of NK, NKR and NMR. (a-c) Representative traces respectively acquired with NK (a) , NKR (b) and NMR (c) . Different peptide events are labeled with different color shadings. Zoomed-in views of representative peptide events are placed to the right of each corresponding trace. Each representative event is taken from the source trace, as indicated by arrow. I0 represents the open pore current of MspA-NTA-Ni. The measurements were carried out in a 1.5 M KCl buffer (1.5 M KCl, 10 mM CHES, pH 9) . A voltage of +100 mV was continually applied. Each peptide was separately added to the cis chamber. The final concentration of NK was 2 mM. The final concentration of NKR was 1 mM. The final concentration of NMR was 0.5 mM. (d-f) The event scatter plots of ΔI versus SD generated by events respectively acquired with NK (d) , NKR (e) and NMR (f) . 100 events for each condition were used to generate each scatter plot.
[0075] Figure 23. Nanopore events of RSLR and SLR. (a-b) Representative traces respectively acquired with RSLR (a) and SLR (b) . Different peptide events are labeled with different color shadings. Zoomed-in views of representative peptide events are placed to the right of each corresponding trace. Each representative event is taken from the source trace, as indicated by the arrow. I0 represents the open pore current of MspA-NTA-Ni. The measurements were carried out in a 1.5 M KCl buffer (1.5 M KCl, 10 mM CHES, pH 9) . A voltage of +100 mV was continually applied. Each peptide was separately added to the cis chamber with a final concentration of 1 mM. (c-d) The scatter plots of ΔI versus SD generated by events respectively acquired with RSLR (c) and SLR (d) . 100 events for each condition were used to generate each scatter plot.
[0076] Figure 24. Zoomed-in views of peptide events with similar event features. Peptides that report visually similar event features are quantitatively compared. (a) Representative events of FQAR, RSLR, ARNKRS and NKR. (b) The event scatter plot of ΔI versus SD generated by events acquired with FQAR, RSLR, ARNKRS and NKR. (c) Representative events of IK, GGGGGG, and NK. (d) The event scatter plot of ΔI versus SD generated by events acquired with IK, GGGGGG and NK. (e) Representative events of DQARNKR, DQARNKRS and LRGY. (f) The event scatter plot of ΔI versus SD generated by events acquired with DQARNKR, DQARNKRS and LRGY. (g) Representative events of NMR, GYIK and LRG. (h) The event scatter plot of ΔI versus SD generated by events acquired with NMR, GYIK and LRG. (i) Representative events of GI, GL, GG and GH (Type I) . (j) The event scatter plot of ΔI versus SD generated by events acquired with GI, GL, GG and GH (Type I) . 100 events for each peptide type were used to generate the scatter plots. Though the event features of each group of peptides above appear similar, they are still thoroughly discriminated by MspA-NTA-Ni when their event features are quantitatively compared.
[0077] Figure 25. Cluster analysis of peptide events. DBSCAN, a density-based clustering algorithm, is used to remove all non-clustered points in the scatter plot. Here, events acquired with glycine peptides were used as representative data to demonstrate cluster analysis. (a-e) Left: The event scatter plots of ΔI versus SD generated by events respectively acquired with GG (a, n=265) , GGG (b, n=311) , GGGG (c, n=276) , GGGGG (d, n=338) and GGGGGG (e, n=410) before cluster analysis. Right: The event scatter plots of ΔI versus SD generated by events respectively acquired with GG (a, n=239) , GGG (b, n=272) , GGGG (c, n=250) , GGGGG (d, n=326) and GGGGGG (e, n=399) after cluster analysis. The epsilon was set to 0.5 and the min samples were set to 10 for each condition.
[0078] Figure 26. Comparison of GR and RG events. (a) Representative traces acquired with GR (left) or RG (right) using MspA-NTA-Ni. I0 represents the open pore current of MspA-NTA-Ni.The peptide events were marked with circles. (b) Representative events of GR (left) and RG (right) . The measurements were carried out using MspA-NTA-Ni in a 1.5 M KCl buffer (1.5 M KCl, 10 mM CHES, pH 9) . A voltage of +100 mV was continually applied. Each peptide was separately added to cis with a final concentration of 1 mM for RG and 1.5 mM for GR. (c) The event scatter plot of ΔI versus SD generated by events acquired with GR and RG. 200 events acquired from each peptide are employed to generate the scatter plot. The events of GR and RG appear as two fully separated event clusters.
[0079] Figure 27. Rapid quality control of synthetic peptides. NK and RSLR peptide sourced from different companies are used as model peptides for this demonstration. Each peptide was reported to have a chromatographic purity exceeding 95%. (a-b) Representative traces respectively measured with NK obtained from Company A (a) or Company B (b) . The peptide was separately added to cis with a final concentration of 1.5 mM. The measurements were carried out using MspA-NTA-Ni in a 1.5 M KCl buffer (1.5 M KCl, 10 mM CHES, pH 9) . A voltage of +100 mV was continually applied. represents the open pore current of MspA-NTA-Ni. Events from NK and impurities are respectively labelled with dots. (c-d) The scatter plots of ΔI versus SD generated by events acquired with NK from Company A (c, n=346) and Company B (d, n=171) . DBSCAN analysis was performed to identify cluster formed populations. The epsilon was set to 0.3 and the min samples were set to 10. The arrow indicates the correct event population. (e) The ratio of correct NK events against impurity events. (f-g) Representative traces measured with RSLR obtained from Company A (f) and Company B (g) . The peptide was separately added to cis with a final concentration of 1 mM for each condition. The correct RSLR events were labelled with dot. (h-i) The scatter plots of ΔI versus SD generated by events acquired with RSLR from Company A (h, n=246) and Company B (i, n=121) . DBSCAN analysis was performed to identify cluster formed populations. The epsilon was set to 0.25 and the min samples were set to 10. The arrow indicates the correct event population. (j) Left: A representative event of SLR, which is a truncated product of RSLR and is considered as an impurity. Right: The scatter plot of ΔI versus SD generated by events acquired with standard SLR. (k) The ratio of correct RSLR events against impurity events.
[0080] Figure 28. Machine learning-assisted peptide identification. (a) The general workflow of machine learning. Sensing events respectively acquired with 30 peptides were collected to form a labeled dataset. Five event features (ΔI, SD, skew, kurt, toff) were extracted from all events to form a feature matrix. The feature matrix was then used to train the machine learning model. According to the ten-fold cross-validation results, the quadratic SVM model has reported the best validation accuracy of 99.0% (Table 3) . (b) The corresponding confusion matrix result generated by the quadratic SVM model.
[0081] Figure 29. Analysis of events acquired during simultaneous sensing of SLR, NKR, LRG, IK and DQ. (a) The event scatter plot of ΔI versus SD generated by results of simultaneous sensing of SLR, NKR, LRG, IK and DQ (n=1048) . (b) The event scatter plot of ΔI versus SD of results shown in a, however after cluster analysis treatment (n=960) . The epsilon was set to 0.08 and the min samples were set to 30. (c) Machine learning prediction of the peptides (n=876) . The trained quadratic SVM model was employed to predict the event identities. Events with a prediction score less than 0.99 were excluded. (d) The frequency distribution histogram of corresponding peptide events, according to which events of SLR, NKR, LRG, IK and DQ were clearly identified.
[0082] Figure 30. Representative events of twenty proteinogenic amino acids and four representative post-translational modified amino acids. The measurements were performed using MspA-NTA-Ni in a 1.5 M KCl buffer (1.5 M KCl, 10 mM CHES, pH 9) . Each amino acid was separately sensed by MspA-NTA-Ni. A voltage of +100 mV was continually applied. Histidine and proline each report two types of sensing events, while each remaining amino acid report a single type of event.
[0083] Figure 31. Simultaneous identification of peptides and amino acids by MspA-NTA-Ni and machine learning. (a) The workflow of machine learning-assisted classification of peptides and amino acids. To further expand the database, events from all twenty proteinogenic amino acids, four post-translational modified amino acids, and thirty peptides were merged to form a bigger database. Five event features (ΔI, SD, skew, kurt, toff) were extracted from events to form a feature matrix. The feature matrix was then fed to the machine learning algorithm for training. According to the ten-fold cross-validation results, the quadratic SVM model has reported the best validation accuracy of 98.1% (Table 4) . (b) The confusion matrix generated by the quadratic SVM model. The model can simultaneously deal with sensing of amino acids and peptides.
[0084] Figure 32. The true positive rate (TPR) of machine learning classification of thirty peptides and twenty-four amino acids. The results were generated by the quadratic SVM model (Figure 30b) . Most classes show a TPR higher than 90%.
[0085] Figure 33. Nanopore analysis of the trypsin hydrolysate of angiotensin III. (a) Representative traces acquired with the trypsin hydrolysate of angiotensin III. All nanopore events are marked with circles above the traces. Angiotensin III was hydrolyzed by trypsin (~10000 BAEE U / mg) at an enzyme / peptide ratio of 1: 50 (w / w) at 37℃ for 8 h. The measurements were carried out using MspA-NTA-Ni in a 1.5 M KCl buffer (1.5 M KCl, 10 mM CHES, pH 9) . A voltage of +100 mV was continually applied. 20 μL of the digestion product was added to the cis chamber to initiate the measurement. I0 represents the open pore current of MspA-NTA-Ni. (b) Representative events acquired with the trypsin digestion product of angiotensin III. Angiotensin III itself can also be detected in the hydrolysates (Figure 62) . (c) The event scatter plot of ΔI versus SD of events acquired with the trypsin hydrolysate of angiotensin III (n=2115) . (d-e) Separation of events of angiotensin III (d) and its digestion fragments (e) . Events acquired with angiotensin III were used to train the One-Class SVM model. (d) The event scatter plot of ΔI versus SD for inlier events acquired with the hydrolysate (n=603) . These inlier events were recognized as events of unhydrolyzed angiotensin III. (e) The event scatter plot of ΔI versus SD of events acquired with the hydrolysate after removing the unhydrolyzed angiotensin III events (n=1512) . (f) The event scatter plot of ΔI versus SD of events in e, however after cluster analysis treatment (n=1321) . The epsilon was set to 0.1 and the min samples were set to 8.
[0086] Figure 34. Nanopore analysis of the trypsin digestion product of secretin. (a) Representative traces acquired with the trypsin digestion product of secretin. Secretin was hydrolyzed by trypsin (~10000 BAEE U / mg) at an enzyme / peptide ratio of 1: 50 (w / w) at 37℃ for 8 h. The measurements were carried out using MspA-NTA-Ni in a 1.5 M KCl buffer (1.5 M KCl, 10 mM CHES, pH 9) . A voltage of +100 mV was continually applied. 20 μL digestion product was added to the cis chamber to initiate the measurement. I0 represents the open pore current of MspA-NTA-Ni. All nanopore events are marked with circles above the traces. (b) Representative events acquired with the trypsin digestion product of secretin. (c) The event scatter plot of ΔI versus SD of events acquired with the trypsin digestion product of secretin (n=590) . (d) The event scatter plot of ΔI versus SD of events in c, however after cluster analysis treatment (n=536) . The epsilon was set to 0.08 and the min samples were set to 10. (e) The event scatter plot of ΔI versus SD of events in d, however after secondary cluster analysis treatment for events within the grey background area (n=504) . The epsilon was set to 0.2 and the min samples were set to 10.
[0087] Figure 35. Nanopore analysis of the trypsin digestion product of glucagon. (a) Representative traces acquired with the trypsin digestion product of glucagon. Glucagon was hydrolyzed by trypsin (~10000 BAEE U / mg) at an enzyme / peptide ratio of 1: 50 (w / w) at 37℃ for 8 h. The measurements were carried out using MspA-NTA-Ni in a 1.5 M KCl buffer (1.5 M KCl, 10 mM CHES, pH 9) . A voltage of +100 mV was continually applied. 20 μL of the digestion product was added to the cis chamber to initiate the measurement. I0 represents the open pore current of MspA-NTA-Ni. All nanopore events are marked with circles above the traces. (b) Representative events acquired with the trypsin digestion product of glucagon. (c-d) The event scatter plot of ΔI versus SD of events acquired with the trypsin digestion product of glucagon before (c, n=952) and after cluster analysis (d, n=670) . The epsilon was set to 0.1 and the min samples were set to 10. (e) A zoomed-in view of the area in d, which is marked with a box.
[0088] Figure 36. Nanopore analysis of the trypsin digestion product of adrenocorticotropic hormone (ACTH) . (a) Representative traces acquired with the trypsin digestion product of ACTH. ACTH was hydrolyzed by trypsin (~10000 BAEE U / mg) at an enzyme / peptide ratio of 1: 50 (w / w) at 37℃ for 8 h. The measurements were carried out using MspA-NTA-Ni in a 1.5 M KCl buffer (1.5 M KCl, 10 mM CHES, pH 9) . A voltage of +100 mV was continually applied. 20 μL of the digestion product was added to the cis chamber to initiate the measurement. I0 represents the open pore current of MspA-NTA-Ni. All nanopore events are marked with circles above the traces. (b) Representative events acquired with the trypsin digestion product of ACTH. (c) The event scatter plot of ΔI versus SD of events acquired with the trypsin digestion product of ACTH (n=803) . (d) The event scatter plot of ΔI versus SD of events in c, however after cluster analysis treatment (n=746) . The epsilon was set to 0.2 and the min samples were set to 10. (e) The event scatter plot of ΔI versus SD of events in d, however after secondary cluster analysis treatment for data within the grey background area (n=582) . The epsilon was set to 0.2 and the min samples were set to 15.
[0089] Figure 37. An overview of nanopore peptide sequencing by assembly. The sequencing strategy contains four steps: (1) Database building. Model peptides in the peptide pool are first sensed by MspA-NTA-Ni. All nanopore events generated by different peptides are collected to form a peptide database. This peptide database is then merged with the amino acid database to create an integrated database. (2) Amino acid component identification. The peptide to be sequenced is thoroughly treated by exopeptidase. The generated amino acids are then sensed by MspA-NTA-Ni. The amino acid identities are then predicted by machine learning so that the amino acid composition of the peptide is known. (3) Database simplification. The integrated database is simplified based on the amino acid composition reported in step (2) . All data corresponding to unidentified amino acids or relevant peptides are excluded from the database to form a simplified database. (4) Peptide fragmentation and sequence assembly. The peptide to be sequenced is separately treated by two types of endopeptidases featuring orthogonal cleavage sites. Then, the hydrolysates of these two endopeptidases are respectively sensed by MspA-NTA-Ni. The fragmented components are identified by machine learning, based on the simplified peptide database described in step (3) . Finally, the identified peptide sequence fragments are assembled to reconstruct its original peptide sequence.
[0090] Figure 38. Representative traces acquired with LAP hydrolysate of the reference peptide (DQARNKRSLRGYIK) . The reference peptide was treated with LAP (>7 U / mg) at an enzyme / peptide ratio of 1: 28 (w / w) at 37℃ for 8 h. All events were automatically identified by an amino acid classifier, and the event identities were marked above the continuous traces. I0 represents the open pore current of MspA-NTA-Ni. The measurements were carried out using MspA-NTA-Ni in a 1.5 M KCl buffer (1.5 M KCl, 10 mM CHES, pH 9) . A voltage of +100 mV was continually applied. 20 μL digestion product was added to the cis chamber to initiate the measurement.
[0091] Figure 39. Identification of amino acid composition of reference peptide (DQARNKRSLRGYIK) . (a) The event scatter plot of ΔI versus SD of events acquired with the LAP hydrolysate of reference peptide (DQARNKRSLRGYIK) (n=1130) . (b) The event scatter plot of ΔI versus SD of events in a, however after cluster analysis treatment (n=1004) . The epsilon was set to 0.2 and the min samples were set to 10. (c) The event scatter plot of ΔI versus SD in b, however after the One-Class SVM data treatment (n=961) . Only events that were identified as amino acids were considered as inliers. (d) Machine learning prediction of amino acid events (n=888) . All event identities were predicted by the previously trained amino acid classifier. Events with a prediction score less than 0.99 were removed. A total of eleven amino acid types were identified, consistent with the amino acid composition of the reference peptide DQARNKRSLRGYIK.
[0092] Figure 40. The confusion matrix result of the simplified database produced by the quadratic SVM model. The integrated database containing data of both peptides and amino acids was further simplified based on the identified amino acid components of the reference peptide (Figure 3a) . Any unidentified amino acids and peptides containing unidentified amino acid constituents were excluded from the database, yielding a further refined database tailored for the reference peptide. This simplified database was then fed to the machine learning classification system for model training. The quadratic SVM model has demonstrated the best validation accuracy of 99.0%, and the corresponding confusion matrix result is demonstrated here.
[0093] Figure 41. Representative traces acquired with the trypsin hydrolysate of the reference peptide. The measurements were carried out using MspA-NTA-Ni in a 1.5 M KCl buffer (1.5 M KCl, 10 mM CHES, pH 9) . A voltage of +100 mV was continually applied. The reference peptide (DQARNKRSLRGYIK) was treated with trypsin (~10000 BAEE U / mg) at an enzyme / peptide ratio of 1: 50 (w / w) at 37℃ for 8 h. 20 μL of the digestion product was added to the cis chamber to initiate the measurement. All events were automatically identified by the trained quadratic SVM model. The predicted event identities are marked above the continuous traces. I0 represents the open pore current of MspA-NTA-Ni.
[0094] Figure 42. Representative traces acquired with thermolysin hydrolysate of the reference peptide. The measurements were carried out in a 1.5 M KCl buffer (1.5 M KCl, 10 mM CHES, pH 9) . A voltage of +100 mV was continually applied. The reference peptide is DQARNKRSLRGYIK. The reference peptide was treated with thermolysin (30-350 U / mg) at an enzyme / peptide ratio of 1: 50 (w / w) at 37℃ for 8 h. 20 μL of the digestion product was added to the cis chamber to initiate the measurement. All events were automatically identified by the trained quadratic SVM model. The predicted event identities are marked above the traces. I0 represents the open pore current of MspA-NTA-Ni.
[0095] Figure 43. Analysis of events acquired with the tryptic digestion product of the reference peptide. (a) The event scatter plot of ΔI versus SD of events acquired with trypsin digestion product of the reference peptide (DQARNKRSLRGYIK) (n=1013) . (b) The event scatter plot of ΔI versus SD of events shown in a, however after cluster analysis treatment (n=880) . The epsilon was set to 0.1 and the min samples were set to 10. (c) Left: Machine learning prediction of the peptides (n=807) . The trained quadratic SVM model was employed to predict the event identities. Events with a prediction score less than 0.99 were excluded. Right: The frequency distribution histogram of peptide events, according to which events of DQAR, GYIK, NKR and SLR were successfully identified.
[0096] Figure 44. Analysis of events acquired with the thermolysin hydrolysate of the reference peptide. (a) The event scatter plot of ΔI versus SD of events acquired with thermolysin digestion product of the reference peptide (n=1147) . (b) The event scatter plot of ΔI versus SD of events shown in a, however after cluster analysis treatment (n=1001) . The epsilon was set to 0.1 and the min samples were set to 15. (c) Left: Machine learning prediction of the peptides (n=970) . The trained quadratic SVM model was employed to predict the event identities. Events with a prediction score less than 0.99 were excluded. Right: The frequency distribution histogram of events, according to which peptide events of DQARNKRS, LRGY, LRG and IK were clearly identified. Amino acid events of Y were also detected, again demonstrating the technical advantage of MspA-NTA-Ni by having a capacity for simultaneous identification of peptides and amino acids.
[0097] Figure 45. LC-MS / MS analysis of DQARNKRSLRGYIK. (a) The total ion chromatogram (TIC) of the reference peptide. The measurement was performed as described in Methods. (b) Representative MS spectrum of the reference peptide. (c) Representative MS / MS spectrum of the reference peptide. The spectrum is dominated by the intact precursor peptide ions. Only 11 / 26 of the fragment ions (b-and y-type ions) derived from cleavage were observed, but with very low intensity. This observation is primarily attributed to the presence of several basic amino acids in the peptide sequence, which impede random fragmentation of the peptide backbone during collision-induced dissociation. As a result, the fragmentation of DQARNKRSLRGYIK is inefficient, leading to insufficient sequence ions for sequence analysis.
[0098] Figure 46. Representative traces acquired with the trypsin digestion product of the D1F peptide. The measurements were carried out in a 1.5 M KCl buffer (1.5 M KCl, 10 mM CHES, pH 9) . A voltage of +100 mV was continually applied. The D1F peptide is FQARNKRSLRGYIK. The D1F peptide was treated with trypsin (~10000 BAEE U / mg) at an enzyme / peptide ratio of 1: 50 (w / w) at 37℃ for 8 h. 20 μL of the digestion product was added to the cis chamber to initiate the measurement. All events were automatically identified by the trained quadratic SVM model, and the predicted event identities were marked above the continuous traces. I0 represents the open pore current of MspA-NTA-Ni.
[0099] Figure 47. Representative traces acquired with the trypsin digestion product of the K6M peptide. The measurements were carried out in a 1.5 M KCl buffer (1.5 M KCl, 10 mM CHES, pH 9) . A voltage of +100 mV was continually applied. The K6M peptide is DQARNMRSLRGYIK. The K6M peptide was treated with trypsin (~10000 BAEE U / mg) at an enzyme / peptide ratio of 1: 50 (w / w) at 37℃ for 8 h. 20 μL of the digestion product was added to the cis chamber to initiate the measurement. All events were automatically identified by the trained quadratic SVM model, and the predicted event identities were marked above the continuous traces. I0 represents the open pore current of MspA-NTA-Ni.
[0100] Figure 48. Representative traces acquired with the trypsin digestion product of the I13L peptide. The measurements were carried out in a 1.5 M KCl buffer (1.5 M KCl, 10 mM CHES, pH 9) . A voltage of +100 mV was continually applied. The I13L peptide is DQARNKRSLRGYLK. The I13L peptide was treated with trypsin (~10000 BAEE U / mg) at an enzyme / peptide ratio of 1: 50 (w / w) at 37℃ for 8 h. 20 μL of the digestion product was added to the cis chamber to initiate the measurement. All events were automatically identified by the trained quadratic SVM model, and the predicted event identities were marked above the continuous traces. I0 represents the open pore current of MspA-NTA-Ni.
[0101] Figure 49. Representative traces acquired with the trypsin digestion product of the K14 deletion peptide. The measurements were carried out in a 1.5 M KCl buffer (1.5 M KCl, 10 mM CHES, pH 9) . A voltage of +100 mV was continually applied. The K14 deletion peptide is DQARNKRSLRGYI. The K14 deletion peptide was treated with trypsin (~10000 BAEE U / mg) at an enzyme / peptide ratio of 1: 50 (w / w) at 37℃ for 8 h. 20 μL of the digestion product was added to the cis chamber to initiate the measurement. All events were automatically identified by the trained quadratic SVM model, and the predicted event identities were marked above the continuous traces. I0 represents the open pore current of MspA-NTA-Ni.
[0102] Figure 50. Analysis of events acquired with trypsin digestion products of mutated peptides. (a, d, g, j) The event scatter plots of ΔI versus SD of events acquired with tryptic digestion products of D1F peptide (a, n=1290) , K6M peptide (d, n=852) , I13L peptide (g, n=1155) and K14 deletion peptide (j, n=692) . (b, e, h, k) Cluster analysis of sensing results respectively acquired with trypsin hydrolysate of D1F peptide (b, n=1134) , K6M peptide (e, n=701) , I13L peptide (h, n=1004) and K14 deletion peptide (k, n=601) . Cluster analysis was performed by the DBSCAN algorithm. The epsilon was set to 0.05 (K6M) or 0.1 (D1F, I13L and K14deletion) . The min samples were set to 10 (D1F) or 15 (K6M, I13L and K14deletion) . (c, f, i, l) Machine learning prediction of the peptides. Left: The event scatter plots of ΔI versus SD of events acquired with tryptic digestion products of D1F peptide (c, n=1126) , K6M peptide (f, n=693) , I13L peptide (i, n=856) and K14 deletion peptide (l, n=537) , however with machine learning predicted labels. The previously trained quadratic SVM model was employed to predict the event identities. The threshold of the predict score was set as 0.95 (D1F and K6M) or 0.99 (I13L and K14deletion) . Events with a predict score less than the threshold value were excluded. Right: The corresponding frequency distribution histograms of peptide events.
[0103] Figure 51. Nanopore measurement of (pS) LR, (pS) LRGYIK and NKR (pS) LR. (a-c) Representative traces acquired with (a) (pS) LR, (b) (pS) LRGYIK and (c) NKR (pS) LR. All peptide events are marked with corresponding shadings. Zoomed-in views of representative peptide events were placed to the right of the source traces. The representative events in the source trace were marked with arrows. I0 represents the open pore current of MspA-NTA-Ni. The measurements were carried out using MspA-NTA-Ni in a 1.5 M KCl buffer (1.5 M KCl, 10 mM CHES, pH 9) . A voltage of +100 mV was continually applied. Each peptide was separately added to the cis chamber. The final concentrations of (pS) LR and NKR (pS) LR were set at 1 mM. The final concentration of (pS) LRGYIK was set at 0.5 mM. (d-f) The event scatter plots of ΔI versus SD generated by events of (d) (pS) LR, (e) (pS) LRGYIK and (f) NKR (pS) LR. 100 events of each type of peptide were used to generate the scatter plot.
[0104] Figure 52. Nanopore measurement of GYI (acK) , SLRGYI (acK) and NKRSLRGYI (acK) . (a-c) Representative traces acquired with GYI (acK) (a) , SLRGYI (acK) (b) and NKRSLRGYI (acK) (c) . All peptide events are marked with corresponding shadings. Zoomed-in views of representative peptide events were placed to the right of the source traces. The representative events in the source trace were marked with arrows. I0 represents the open pore current of MspA-NTA-Ni. The measurements were carried out in a 1.5 M KCl buffer (1.5 M KCl, 10 mM CHES, pH 9) . A voltage of +100 mV was continually applied. Each peptide was separately added to the cis chamber with a final concentration of 0.5 mM. (d-f) The event scatter plots of ΔI versus SD generated by events of (d) GYI (acK) , (e) SLRGYI (acK) and (f) NKRSLRGYI (acK) . 100 events of each type of peptide were used to generate the scatter plot.
[0105] Figure 53. The confusion matrix of peptides with phosphorylation modification. Three phosphorylated peptides ( (pS) LR, (pS) LRGYIK, NKR (pS) LR, Table 5) were added to the existing peptide database to create a larger peptide database containing phosphorylation modifications. The new database was then passed to the machine learning classification algorithm for further training. The ten-fold cross-validation accuracy and the testing accuracy reported by the quadratic SVM model are 99.0%and 98.5%, respectively.
[0106] Figure 54. The confusion matrix of peptides with acetylation modification. Three acetylated peptides (GYI (acK) , SLRGYI (acK) , NKRSLRGYI (acK) , Table 5) were added to the existing peptide database to create a larger peptide database containing acetylation modifications. The new database was then passed to the machine learning classification algorithm for further training. The ten-fold cross-validation accuracy and the testing accuracy reported by the quadratic SVM model are 99.0%and 98.6%, respectively.
[0107] Figure 55. Representative traces acquired with the trypsin digestion product of the pS8 peptide. The measurements were carried out in a 1.5 M KCl buffer (1.5 M KCl, 10 mM CHES, pH 9) . A voltage of +100 mV was continually applied. The pS8 peptide is DQARNKR (pS) LRGYLK. The pS8 peptide was treated with trypsin (~10000 BAEE U / mg) at an enzyme / peptide ratio of 1: 50 (w / w) at 37℃ for 8 h. 20 μL of the digestion product was added to the cis chamber to initiate the measurement. All events were automatically identified by the trained quadratic SVM model, and the predicted event identities were marked above the continuous traces. I0 represents the open pore current of MspA-NTA-Ni.
[0108] Figure 56. Representative traces acquired with the trypsin digestion product of the acK14 peptide. The measurements were carried out in a 1.5 M KCl buffer (1.5 M KCl, 10 mM CHES, pH 9) . A voltage of +100 mV was continually applied. The acK14 peptide is DQARNKRSLRGYL (acK) . The acK14 peptide was treated with trypsin (~10000 BAEE U / mg) at an enzyme / peptide ratio of 1: 50 (w / w) at 37℃ for 8 h. 20 μL of the digestion product was added to the cis chamber to initiate the measurement. All events were automatically identified by the trained quadratic SVM model, and the predicted event identities were marked above the continuous traces. I0 represents the open pore current of MspA-NTA-Ni.
[0109] Figure 57. Analysis of events acquired with the trypsin hydrolysate of post-translationally modified peptides. (a, d) The event scatter plots of ΔI versus SD of events acquired with tryptic digestion products of pS8 peptide (a, n=806) and acK14 peptide (d, n=915) . (b, e) Cluster analysis of sensing results respectively acquired with trypsin digestion products of pS8 peptide (b, n=768) and acK14 peptide (e, n=839) . Cluster analysis was performed by the DBSCAN algorithm. For the digestion product of pS8 peptide, the epsilon was set to 0.2 and the min samples were set to 20. For the digestion product of acK14 peptide, the epsilon was set to 0.08 and the min samples were set to 8. (c, f) Machine learning prediction of the peptides. Left: The event scatter plots of ΔI versus SD of events respectively acquired with the tryptic digestion products of pS8 peptide (c, n=701) and acK14 peptide (f, n=783) , however with machine learning predicted labels. The trained quadratic SVM model was employed to predict the event identities. Events with a predict score less than 0.99 were excluded. Right: The corresponding frequency distribution histograms of peptide events.
[0110] Figure 58. Nanopore events of aspartame. (a) The chemical structure of aspartame. (b) A representative trace acquired using MspA-NTA-Ni with aspartame. I0 represents the open pore current of MspA-NTA-Ni. All events of aspartame are marked with pentacles. The measurements were carried out using an MspA-NTA-Ni in a 1.5 M KCl buffer (1.5 M KCl, 10 mM CHES, pH 9) .A voltage of +100 mV was continually applied. Aspartame was added to the cis chamber with a final concentration of 0.5 mM. (c) A representative aspartame event. (d) The event scatter plot of ΔI versus SD generated by events acquired with aspartame (n=240) .
[0111] Figure 59. Nanopore events of cyclo (His-Pro) . (a) The chemical structure of cyclo (His-Pro) . (b) A representative trace acquired with cyclo (His-Pro) . I0 represents the open pore current of MspA-NTA-Ni. Cyclo (His-Pro) reports three types of events. Different types of events were respectively marked with Roman numerals above the trace. The measurements were carried out using MspA-NTA-Ni in a 1.5 M KCl buffer (1.5 M KCl, 10 mM CHES, pH 9) . A voltage of +100 mV was continually applied. Cyclo (His-Pro) was added to the cis chamber with a final concentration of 1 mM. (c) Representative events of cyclo (His-Pro) . (d) The event scatter plot of ΔI versus SD generated by events acquired with cyclo (His-Pro) (n=557) .
[0112] Figure 60. Nanopore events of RGD. (a) The chemical structure of RGD peptide. (b) A representative trace acquired using MspA-NTA-Ni with RGD. I0 represents the open pore current of MspA-NTA-Ni. The events of RGD were marked with pentacles. The measurements were carried out using an MspA-NTA-Ni in a 1.5 M KCl buffer (1.5 M KCl, 10 mM CHES, pH 9) . A voltage of +100 mV was continually applied. RGD was added to the cis chamber with a final concentration of 1.5 mM. (c) A representative RGD event. (d) The event scatter plot of ΔI versus SD generated by events acquired with RGD (n=180) .
[0113] Figure 61. Nanopore events of GSH and GSSG. (a, f) The chemical structure of GSH and GSSG. (b, f) Representative traces acquired with GSH (b) or GSSG (f) . I0 represents the open pore current of MspA-NTA-Ni. Events of GSH and GSSG were marked with pentacles above the corresponding traces. The measurements were carried out using MspA-NTA-Ni in a 1.5 M KCl buffer (1.5 M KCl, 10 mM CHES, pH 9) . A voltage of +100 mV was continually applied. Each peptide was separately added to the cis chamber with a final concentration of 2 mM. (c, g) Representative events of GSH (c) and GSSG (g) . (d, h) The event scatter plots of ΔI versus SD generated by events acquired with GSH (d, n=127) and GSSG (h, n=192) .
[0114] Figure 62. Nanopore events of Leu-enkephalin. (a) The chemical structure of Leu-enkephalin. (b) A representative trace acquired using MspA-NTA-Ni with Leu-enkephalin. I0 represents the open pore current of MspA-NTA-Ni. The events of Leu-enkephalin were marked with pentacles. The measurements were carried out in a 1.5 M KCl buffer (1.5 M KCl, 10 mM CHES, pH 9) . A voltage of +100 mV was continually applied. Leu-enkephalin was added to the cis chamber with a final concentration of 0.5 mM. (c) A representative Leu-enkephalin event. (d) The event scatter plot of ΔI versus SD generated by events acquired with Leu-enkephalin (n=202) .
[0115] Figure 63. Nanopore events of angiotensin III. (a) The chemical structure of angiotensin III. (b) Representative events of angiotensin III. I0 represents the open pore current of MspA-NTA-Ni.Angiotensin III reports three types of sensing events. The measurements were carried out using MspA-NTA-Ni in a 1.5 M KCl buffer (1.5 M KCl, 10 mM CHES, pH 9) . A voltage of +100 mV was continually applied. Angiotensin III was added to the cis chamber with a final concentration of 0.5 mM. (c) A representative trace acquired using MspA-NTA-Ni with angiotensin III. Different types of events were marked with Roman numerals above the trace. (d) The event scatter plot of ΔI versus SD generated by events acquired with angiotensin III (n=795) .
[0116] Figure 64. Nanopore events of somatostatin-14 cyclopeptide. (a) The chemical structure of somatostatin-14. (b) Top: A representative trace acquired with somatostatin-14. The somatostatin-14 events were marked with shadings. Bottom: Representative events of somatostatin-14. Transient blockage level transitions are occasionally observed, indicating the existence of multiple binding configurations of somatostatin-14 when captured by MspA-NTA-Ni. The measurements were carried out using MspA-NTA-Ni in a 1.5 M KCl buffer (1.5 M KCl, 10 mM CHES, pH 9) . A voltage of +100 mV was continually applied. Somatostatin-14 was added to cis with a final concentration of 0.02 mM. I0 represents the open pore current of MspA-NTA-Ni. (c) The event scatter plot of ΔI versus SD generated by somatostatin-14 events (n=140) .
[0117] Figure 65. Nanopore events of secretin. (a) The chemical structure of secretin. (b) Top: A representative trace acquired using MspA-NTA-Ni with secretin as the sole analyte. I0 represents the open pore current of MspA-NTA-Ni. Secretin demonstrates two types of sensing events. Different types of events (type I and II) are marked with Roman numerals. Bottom: A zoomed in demonstration of representative events of secretin. The measurements were carried out using MspA-NTA-Ni in a 1.5 M KCl buffer (1.5 M KCl, 10 mM CHES, pH 9) . A voltage of +100 mV was continually applied. Secretin was added to the cis chamber with a final concentration of 5 μM.(c) The event scatter plot of ΔI versus SD of events acquired with secretin (n=87) .
[0118] Figure 66. Nanopore events of glucagon. (a) The chemical structure of glucagon. (b) Top: A representative trace acquired using MspA-NTA-Ni with glucagon as the sole analyte. I0 represents the open pore current of MspA-NTA-Ni. Glucagon reports two types of sensing events. Different types of events (type I and II) are respectively marked with Roman numerals. Bottom: A zoomed-in demonstration of representative events of glucagon. The measurements were carried out using MspA-NTA-Ni in a 1.5 M KCl buffer (1.5 M KCl, 10 mM CHES, pH 9) . A voltage of +100 mV was continually applied. Glucagon was added to the cis chamber with a final concentration of 5 μM. (c) The event scatter plot of ΔI versus SD of events acquired with glucagon (n=181) .
[0119] Figure 67. Nanopore events of ACTH. (a) The chemical structure of ACTH. (b) Top: A representative trace acquired using MspA-NTA-Ni with ACTH as the sole analyte. I0 represents the open pore current of MspA-NTA-Ni. The events generated by ACTH are marked with triangles. Bottom: A zoomed-in demonstration of representative events of ACTH. The measurements were carried out using MspA-NTA-Ni in a 1.5 M KCl buffer (1.5 M KCl, 10 mM CHES, pH 9) . A voltage of +100 mV was continually applied. ACTH was added to the cis chamber with a final concentration of 5 μM. (c) The event scatter plot of ΔI versus SD of events acquired with ACTH (n=190) .
[0120] Figure 68. Simultaneous sensing of bioactive peptides. (a) The work flow of machine learning. Events respectively acquired with eleven bioactive peptides were used to train the model. The bagged trees model has reported the best with a validation accuracy of 97.8%. (b) The confusion matrix result generated by the trained bagged trees model. (c-d) Top: Representative traces acquired during simultaneous sensing of bioactive peptides. The measurements were performed using an MspA-NTA-Ni in a 1.5 M KCl buffer (1.5 M KCl, 10 mM CHES, pH 9) . A voltage of +100 mV was continually applied. To initiate the measurements, aspartame, RGD, Leu-enkephalin, cyclo (His-Pro) , angiotensin III, GSH and GSSG were simultaneously added to cis. The final concentration of angiotensin III was 0.1 mM. The final concentrations of all other peptides were 0.5 mM for each component. Bottom: Zoomed-in view of representative events. (e) The raw event scatter plot of ΔI versus SD generated by results acquired during simultaneous sensing of seven bioactive peptides (n=1113) . (f) The event scatter plot of ΔI versus SD of events in e, however after cluster analysis treatment (n=1082) . The epsilon was set to 0.2 and the min samples were set to 10. (g) The event scatter plot of ΔI versus SD of events in f however with machine learning predicted labels (n=971) . All events were identified by the trained bagged trees model. Events with a prediction score less than 0.8 were excluded.
[0121] Figure 69. Comparison of events of RGD, GDR and DRG. (a) Representative events of RGD, GDR and DRG. I0 represents the open pore current of MspA-NTA-Ni. The measurements were carried out using an MspA-NTA-Ni in a 1.5 M KCl buffer (1.5 M KCl, 10 mM CHES, pH 9) .A voltage of +100 mV was continually applied. Each peptide was separately added to the cis chamber with a final concentration of 1.5 mM. (b) The event scatter plot of ΔI versus SD generated by events acquired with RGD, GDR and DRG. 150 events acquired with each type of peptide were employed to generate the plot. Clearly, three populations of events, respectively corresponding to RGD, GDR and DRG, are observed. These results demonstrate the high resolution of MspA-NTA-Ni which enables direct discrimination of peptide isomers however composed of different sequences.
[0122] Figure 70. The true positive rate (TPR) of machine learning classification of seventy-three types of peptides and amino acids. The results were generated by the quadratic SVM model (Figure 6) . Most classes show a TPR higher than 90%.
[0123] Figure 71. The discrimination capacity of this sensing strategy against an expanding database size. (a) The accuracy of machine learning during the gradual expansion of the database size. (b) The plot of validation accuracy versus the database size. The quadratic SVM model and events acquired with 49 peptides were used for this demonstration. The quadratic SVM model was respectively trained with different numbers of peptide types, ranging from two to forty-nine peptides. Consequently, the cross-validation accuracy corresponding to each database size was derived. The data was fitted to a power function. The accuracy of further expanded database can be extrapolated from the fitting results.
[0124] Figure 72. Nanopore recognition of neoantigens. (a-b) Left: Representative traces acquired with BRAFV600E neoantigen (sequence: LATEKSRWSG) and its wide-type peptide (BRAFWT, sequence: LATVKSRWSG) . The measurements were carried out using an MspA-NTA-Ni in a 1.5 M KCl buffer (1.5 M KCl, 10 mM CHES, pH 9) . A voltage of +100 mV was continually applied. Each peptide was separately added to the cis chamber with a final concentration of 0.2 mM.I0 represents the open pore current of MspA-NTA-Ni. Right: Representative events of BRAFV600E and BRAFWT peptide. Each event is taken from the corresponding source trace highlighted with shadings. (c-d) Left: Representative traces acquired with KRASG12D neoantigen (sequence: VVVGADGVGK) and its wild-type counterpart (KRASWT, sequence: VVVGAGGVGK) . Each peptide was separately added to the cis chamber with a final concentration of 0.4 mM. Right: Representative events of KRASG12D and KRASWT. Each event is taken from the corresponding source trace highlighted with shadings. (e) Top: A representative trace acquired during simultaneous sensing of BRAFV600E and BRAFWT. The two peptides were simultaneously added to cis with a final concentration of 0.4 mM. Bottom: Representative events of BRAFV600E and BRAFWT. Each event is taken from the source trace, as indicated by the grey shadings. (f) The scatter plot of ΔI versus SD generated by events acquired during simultaneous sensing of BRAFV600E and BRAFWT (n=434) . (g) A representative trace acquired during simultaneous sensing of KRASG12D and KRASWT. The two peptides were simultaneously added to cis with a final concentration of 0.8 mM. Bottom: Representative events of KRASG12D and KRASWT. Each event is taken from the source trace, as indicated by the grey shadings. (h) The scatter plot of ΔI versus SD generated by results acquired during simultaneous sensing of KRASG12D and KRASWT (n=249) .
[0125] Figure 73. Nanopore peptide sequencing integrated with Edman degradation. (a) Schematic workflow of the hybrid sequencing strategy using GGGGGG as a model peptide. The model peptide first reacts with phenylisothiocyanate (PITC) under mild conditions, followed by acidic cleavage to remove the modified N-terminal residue. The truncated peptide is then partitioned for parallel processing, with one aliquot analyzed by nanopore while the other subjected to cycles of labeling and cleavage reactions, allowing for sequential identification of the model peptide. (b-f) Representative current segments acquired during Edman degradation of GGGGGG. Nanopore measurements were performed with MspA-NTA-Ni in a 1.5 M KCl buffer (1.5 M KCl, 10 mM CHES, pH 9.0) . A voltage of +100 mV was continually applied. I0 represents the open pore current of MspA-NTA-Ni. Ib represents the blockage amplitude of the binding events. (g-k) The scatter plots of ΔI versus SD generated by results acquired Edman degradation of GGGGGG. The progressively truncated peptides generate unique nanopore signatures and can be clearly distinguished.DETAILED DESCRIPTION OF THE INVENTION
[0126] Unless defined otherwise, all technical and scientific terms used herein have the same meaning as is commonly understood by one of ordinary skill in the art to which this invention belongs. All publications, patent applications, patents, and other references mentioned herein are incorporated by reference in their entirety. In case of conflict, the present specification, including definitions, will control. In addition, the materials, methods, and examples herein are illustrative only and not intended to be limiting.
[0127] The terms “about” and “approximate, ” when used along with a numerical variable, generally means the value of the variable and all the values of the variable within a measurement or an experimental error (e.g., 95%confidence interval for the mean) or within a specified value within a broader range (e.g., ± 10%) .
[0128] As used herein, the singular forms “a” , “an” and “the” are intended to include the plural forms as well, unless the context clearly indicates otherwise.
[0129] The term “comprise” and variations thereof, such as “comprises” and “comprising” , as well as “contain” , “containing” , “have” , “having” , “include” and “including” means including the recited steps or elements, but not excluding other steps or elements. “Consisting of” means excluding any step or element not specified. “Consisting essentially of” means not excluding steps or elements that do not materially affect the basic and novel characteristics of the claimed invention. The term “comprise" and its variants also include the cases of “consisting of ......” and “consisting essentially of ......” .
[0130] Where a range of values is provided, it is understood that the upper and lower limits, and each smaller range between the upper limit (or the lower limit) and any intervening value, or between any two intervening values in that range, shall be considered to be specifically disclosed. Any intervening range and all individual value in the stated range of value may be excluded from said range of value.
[0131] The term “and / or” refers to any one, several or all of the elements connected by the term.
[0132] Unless otherwise indicated, nucleic acids are written left to right in 5’ to 3’ orientation; amino acid sequences are written left to right in N-terminus to C-terminus orientation, respectively.
[0133] The terms “first” , “second” , “third” , “fourth” etc. are used only to distinguish between elements or steps and do not imply a specific order of precedence or location relationship. It should be understood that when “second” element or step is mentioned, it is meant to be literally distinguished from other elements or steps and it is not necessary to have a corresponding “first” element or step.
[0134] Further, it should be understood that although the specification may have presented the method and / or process of the present invention as a particular sequence of steps. However, to the extent that the method or process does not rely on the particular order of steps set forth herein, the method or process should not be limited to the particular sequence of steps described. As one of ordinary skill in the art would appreciate, other sequences of steps may be possible. Therefore, the particular order of the steps set forth in the specification should not be construed as limitations on the claims.
[0135] It should be understood that the method of the present invention may be performed in vivo, in vitro, or ex vivo. The method of the present invention may be not for the purpose of disease treatment, and / or not for the purpose of disease diagnosis.
[0136] The inventors have constructed a nanopore with a metal ion as a sensing moiety, the interaction between a peptide and the metal ion can reinforce the measurable ionic current change (blockage of the ionic current) across the nanopore, make the distinctive features of ionic current change more pronounced and recognizable, and allow the identification of different peptides at an unprecedentedly high resolution. This nanopore can be used to directly detect the intact molecule of a peptide and obtain information about its fingerprint, sequence, conformation or structure without resorting to labelling, traction or linearization of the peptide, or local detection of the peptide.
[0137] In a first aspect, the present invention provides a method of characterizing one or more target analytes, comprising:
[0138] (i) providing a nanopore incorporating at least one sensing moiety which is capable of interacting with a peptide, and causing the target analyte to contact the sensing moiety in a region that allows the target analyte to be detected based on a change of an ionic current across the nanopore; and
[0139] (ii) measuring the ionic current as the analyte migrates to provide a representative current pattern of the target analyte, wherein the representative current pattern indicates a characteristic of the target analyte;
[0140] wherein the one or more target analytes comprise a peptide. In some embodiments, the one or more target analytes are one or more proteinogenic analytes (which may comprise one or more peptide, and may further comprise one or more amino acids) .
[0141] The nanopore used in the method incorporates at least one sensing moiety (which can also be called reactive site or sensing site herein) which is attached to the inner surface of the channel of the nanopore, wherein the sensing moiety can interact with the peptide, which allows the nanopore to identify the peptide and optionally other molecules that can be identified by said nanopore (e.g., an amino acid) .
[0142] In some embodiments, the nanopore may be a biological nanopore or a solid-state nanopore or a combination thereof.
[0143] The biological nanopore may be a protein nanopore, e.g., a porin; or a DNA origami structure. The solid-state nanopore may include or be formed from SiNx, SiO2, CNT, graphene, nanopipettes, etc.
[0144] The channel of a protein nanopore typically comprises a narrowed region, called the constriction region. When an analyte enters the constriction region, the ionic current is blocked at the constriction region, thereby causing a detectable change of the current, which can reflect the characteristic of the analyte. In some embodiments, other regions of a protein nanopore can also be used to detect an analyte, such as vestibule or entrance, i.e., changes in the ionic current when an analyte enters such region are also capable of reflecting the characteristics of the analyte. If a detectable change of the ionic current across the nanopore can be induced when an analyte migrates into a certain region of the nanopore, these regions may be considered to be the detection region of said nanopore, which include, but are not limited to, the constriction zone, vestibule, and entrance of the nanopore, which are known to a person of skill in the art and can be determined by any method for detecting an ionic current across the nanopore.
[0145] The nanopore (e.g., protein nanopore) may be of any suitable shape, such as a conical shaped nanopore or a cylindrical shaped nanopore. in some embodiments, a conical shaped nanopore is preferred.
[0146] In some embodiments, the protein nanopore may comprise a α-helical structure or a β-barrel structure. In some embodiments, a protein nanopore having a β-barrel structure is preferred, which may provide rigid scaffolds for the transport of molecules across the nanopore.
[0147] Preferably, the protein nanopore used in the present invention does not gate spontaneously, even at 150mV-200mV or more. For some protein nanopore, the probability of gating increases with the application of higher voltages. Typically, the protein becomes less conductive during gating, and conductance may permanently stop (i.e., the tunnel may permanently shut) as a result, such that the process is irreversible. Optionally, gating refers to the conductance through the tunnel of a protein spontaneously changing to less than 75%of its open state current.
[0148] In some embodiments, the protein nanopore may selected from alpha-hemolysin (α-HL) , Mycobacterium smegmatis porin A (MspA) , Aerolysin (Ael) , curli production assembly / transport component (CsgG) , outer membrane porin F (OmpF) , Cytolysin A (ClyA) , ferric hydroxamate uptake component A (FhuA) , Fragaceatoxin C (FraC) , Cytotoxin K (CytK) , Pleurotolysin A (PlyA) / Pleurotolysin B (PlyB) , Curli production assembly / transport component CsgG (CsgG) , Phi29 connector protein, e.g., wild-type ones and any variants thereof.
[0149] In some embodiments, the protein nanopore may be a MspA, which may be a wild-type MspA or a variant of wild-type of MspA, e.g., M1 MspA or M2 MspA.
[0150] A wild-type MspA is an octameric protein nanopore in which each monomer has the following sequence:
[0151] Variants of a MspA include, but are not limited to, an octameric protein nanopore in which each monomer has a mutation of D90N / D91N / D93N (M1 MspA) or D93N / D91N / D90N / D118R / D134R / E139K (M2 MspA) compared to the wild-type one. The description of the mutation means that the variant comprises simultaneously all of listed mutations compared to the wild-type one, wherein the amino acid numbering is with reference to wild-type MspA.
[0152] The nanopore comprises at least one sensing moiety, and preferably, only one sensing moiety.
[0153] The sensing moiety may be attached to the nanopore directly or via a linker. In some embodiments, the sensing moiety may be attached to an amino acid residue of the protein nanopore, which may be called a reactive amino acid residue. In some embodiments, the sensing moiety may be attached to the reactive amino acid residue via a ligand. The sensing moiety or both the sensing moiety and the linker (e.g., ligand) may be called an adapter herein, which is assembled onto a nanopore for characterization of peptides, amino acids, and / or other analytes. In some embodiments, the nanopore described herein incorporates at least one adapter; preferably, only one adapter.
[0154] In some embodiments, the adapter may be attached to a site within the detection region of the nanopore, such as the constriction region, vestibule, or entrance of a protein nanopore. In some embodiments, the adapter may be attached to a site of the nanopore such that the sensing moiety is located within the detection region of the nanopore, e.g., during the detection process.
[0155] The nanopore may comprise one or more reactive amino acid residues. In some embodiments, the nanopore comprise only one reactive amino acid residue. In some embodiments, each reactive amino acid residue attaches to one sensing moiety.
[0156] The reactive amino acid residue may be located on the inner surface of the channel of the nanopore. In some embodiments, the reactive amino acid residue may be located at the constriction region, vestibule, or entrance of a protein nanopore.
[0157] It should be understood that in some cases, a protein nanopore inherently comprises suitable reactive amino acid residue that can be attached with the sensing moiety or the linker.
[0158] In some other cases, a protein nanopore may not have a suitable reactive amino acid residue, in which case a suitable reactive amino acid residue can be created by modification of the protein nanopore (e.g., insertion, substitution and / or deletion of one or more amino acids) .
[0159] In some other cases, a protein nanopore may have multiple reactive amino acid residues, in which case one or more but not all reactive amino acid residues may be eliminated by modification to reduce the number of sensing moieties attached to the protein nanopore.
[0160] For example, an amino acid residue that is not suitable for a reactive amino acid residue (i.e., a non-reactive amino acid residue) in the protein nanopore may be replaced with a reactive amino acid residue to create a reactive amino acid residue; ; or a reactive amino acid residue in the protein nanopore may be replaced with a non-reactive amino acid residue, which may be achieved by well-known techniques, such as solid-phase synthesis of polypeptides or DNA recombination and protein expression techniques.
[0161] The protein nanopore that have suitable reactive amino acid residue after modification may be called a variant of the parental protein nanopore and may be referred to as being derived from the parental protein nanopore.
[0162] The parental protein nanopore may be a wild-type protein nanopore or a variant thereof, such as any of the aforementioned protein nanopores.
[0163] Typically, a protein nanopore consists of two or more monomers.
[0164] In some embodiments, the protein nanopore used in the method of the present invention may be a heterogeneous protein nanopore in which one or more but not all monomers incorporate the sensing moiety and the other monomers do not incorporate the sensing moiety.
[0165] A monomer incorporating the sensing moiety may be called reactive monomer. It should be understood that a reactive monomer comprises a reactive amino acid residue. In some embodiments, one reactive monomer may incorporate only one sensing moiety. In some embodiments, only one monomer incorporates a sensing moiety (e.g., only one sensing moiety) .
[0166] In some embodiments, all of the reactive monomers in the protein nanopore are the same as or different from each other, and each reactive monomer may independently be a wild-type monomer or a variant of wild-type monomer.
[0167] The monomer that does not comprise the sensing moiety may be called non-reactive monomer. It should be understood that a non-reactive monomer does not comprise a reactive amino acid residue. In some embodiments, all of the non-reactive monomers in the protein nanopore are the same as or different from each other, and each non-reactive monomer may independently be a wild-type monomer or a variant of wild-type monomer.
[0168] Compared to the non-reactive monomer, the reactive monomer may comprise an additional amino acid modification that provides attachment to the sensing moiety, and optionally comprises or does no comprise other amino acid modification.
[0169] The heterogeneous protein nanopore of the present invention can be prepared by providing one or more reactive monomers and one or more non-reactive monomers (such as by co-expression of them) , and enabling them to assemble into a protein nanopore under appropriate conditions.
[0170] The reactive monomer and the non-reactive monomers may be prepared by by conventional methods known in the art, such as chemical synthesis or genetic recombination.
[0171] The reactive amino acid residue may be located on the surface of the channel. The reactive amino acid residue may be located at any position on the surface of the nanopore channel, such as the constriction zone, which is the narrowest portion of the nanopore channel, or the vestibule, which is at one end of the nanopore channel and has a larger diameter than the constriction zone.
[0172] When the protein nanopore is derived from MspA or variant thereof, one or more reactive amino acid residues are located at one or more positions selected from 83-111, preferably 90, 91, 92 and 93, wherein the position of the amino acid residue is with reference to the wild-type MspA. In some embodiments, the one or more reactive amino acid residues comprise one or more selected from cysteine, methionine, histidine and lysine. In some embodiments, the reactive amino acid residue is cysteine or methionine, which may be located at positions selected from 90, 91, 92 and 93.
[0173] In some embodiments, the heterogeneous protein nanopore that is used in the method of the present invention comprises at least one amino acid mutation in one or more monomers compared to MspA or M2 MspA. In some embodiments, the mutation comprises mutation to cysteine, methionine, histidine and / or lysine, preferably at one or more positions selected from 83-113, preferably 90, 91, 92 and 93.
[0174] In some embodiments, the heterogeneous protein nanopore that is used in the method of the present invention comprise a single reactive monomer which comprise a single reactive amino acid residue, wherein the single reactive amino acid residue is located at position 90, 91, 92 or 93 and selected from cysteine and methionine. In some embodiments, the heterogeneous protein nanopore that is used in the method of the present invention has a mutation of N90C, N90M, N91M and / or N91C in one or more monomers compared to wild-type MspA or a variant thereof (e.g., M2 MspA) .
[0175] The sensing moiety is able to interacting with a peptide and / or an amino acid. In some embodiments, the sensing moiety is able to interacting with a moiety, a group or an atom of the peptide and / or the amino acid. The interaction may be a chemical reaction, such as a reversible or irreversible chemical reaction, which can form a chemical bond (e.g., a coordination bond and / or a covalent bond) between the sensing moiety and the peptide and / or amino acid. The interaction may occur automatically when the peptide and / or amino acid contacts the sensing moiety. The interaction (e.g., chemical reaction) may include bond-forming and bond-breaking processes between the sensing moiety and the peptide and / or amino acid.
[0176] In some preferred embodiments, the bonding time of the chemical bond between the sensing moiety and the peptide and / or the amino acid is no shorter than about 1 ms. In some preferred embodiments, the bonding time of the chemical bond between the sensing moiety and the peptide and / or the amino acid is no longer than about 10 s. Appropriate bonding time can be more conducive to achieve a higher resolution for the analyte. The bonding time can be adjusted by a number of parameters, such as temperature, pH, voltage, and so on. The bonding time can be learned from the measured current pattern.
[0177] The sensing moiety constructed on the inner surface of the nanopore and the interaction between the sensing moiety and the peptide and / or amino acid not only amplifies the blockage signal in the nanopore, but also enhances the capture of the peptide and / or amino acid and promotes their retention at the location of the sensing moiety for signal detection. The interaction between the sensing moiety and the peptide constrains the spatial configuration (or the degree of freedom of its spatial flip) of the peptide when it is being detected, which can improve the consistency of the peptide detection signal from time to time and reduce the difficulty of data processing, and reduce the detection error.
[0178] In some embodiments, the interaction between the sensing moiety and the peptide and / or amino acid is reversible and dynamic, along with continuous bonding and bond-breaking processes between the peptide and / or amino acid and the sensing moiety, thereby making it possible to use a single nanopore to detect a variety of peptides and / or amino acids simultaneously, e.g., a mixture of them. It should be understood that when characterizing a plurality of analytes, each analyte, in turn, separately interacts with the sensing moiety to produce a measurable and unique signal to characterize each analyte.
[0179] In some embodiments, the sensing moiety may be a metal ion, such as a metal cation (e.g., a transition metal ion) , which can coordinate with a peptide and / or an amino acid (e.g., a moiety, a group or an atom thereof) . Examples of such metal ion include Ni2+, Cu2+, Co2+, Zn2+, Cd2+, Ag+, Mn2+, Pb2+, Fe2+ or Fe3+.
[0180] The sensing moiety may be attached to the nanopore (in some embodiments, attached to the reactive amino acid residue) directly or via a linker.
[0181] In some embodiments, the linker is attached to the reactive amino acid residue by a chemical bond, e.g., a coordination bond and / or a covalent bond.
[0182] In some embodiments, the sensing moiety can be attached to the linker by a chemical bond, e.g., a coordination bond and / or a covalent bond.
[0183] In some embodiments, the linker may be a ligand. In some embodiments, the ligand and the sensing moiety may form a coordinate complex.
[0184] The linker may be attached to the reactive amino acid residue by any suitable approaches, such as a chemical reaction, e.g., a click reaction. Examples of the such chemical reaction may include, but not limited to, a copper (I) -catalyzed alkyne-azide cycloaddition (CuAAC) , such as a reaction between azide and alkyne; a copper free alkyne-azide cycloaddition, such as a reaction between azide and dibenzocyclooctyne (DBCO) or difluorinated cyclooctyne; a staudinger ligation, such as a reaction between azide and phosphine; a radical addition, such as between a reaction thiol and alkene; a Michael addition, such as a reaction between thiol and maleimide; a nucleophilic substitution, such as a reaction between amine and para-fluoro (Becer, Hoogenboom, and Schubert, Click Chemistry beyond Metal-Catalyzed Cycloaddition, Angewandte Chemie International Edition, 2009, 48: 490-4908; Rostovtsev, V. V. et al., 2002, A stepwise Huisgen cycloaddition process: Copper (I) -catalyzed regioselective “ligation” of azides and terminal alkynes. Angew. Chem., Int. Ed. 41, 2596-2599; Torne, C. W. et al., 2002, Peptidotriazoles on solid phase: [1, 2, 3] -Triazoles by regiospecific copper (I) -catalyzed 1, 3-dipolar cycloadditions of terminal alkynes to azides. J. Org. Chem. 67, 3057-3064; Agard, N. J. et al., 2004, A strainpromoted [3+2] azide-alkyne cycloaddition for covalent modification of blomolecules in living systems. J. Am. Chem. Soc. 126, 15046-15047; Kohn, M., and Breinbauer, R., 2004, The Staudinger ligation: A gift to chemical biology. Angew. Chem., Int. Ed. 43, 3106-3116) .
[0185] In some embodiments, the ligand may be a metal chelating agent, such as nitrilotriacetic acid (NTA) , diethylenetriaminepentaacetic acid (DTPA) , ethylene diamine tetraacetic acid (EDTA) , N- (2-hydroxyethyl) ethylenediaminetriaceticacid (HEDTA) , ethylenediamine (EDA) , bipyridine, 1, 10-phenanthroline monohydrate or iminodiacetic acid (IDA) , which can be attached to the reactive amino acid residue by a chemical reaction, for example, a reaction between thiol and maleimide.
[0186] According to the present invention, a suitable linker and / or a suitable reactive amino acid residue may be selected according to the particular species of the sensing moiety. The reactive amino acid residue should be able to be linked to the linker or the sensing moiety, and the linker should be able to be linked to the reactive amino acid residue and the sensing moiety, respectively.
[0187] In some preferred embodiments, the protein nanopore (especially the heterogeneous protein nanopore) of the present invention comprises Ni2+ as a sensing moiety that is attached to a reactive amino acid residue via NTA, wherein NTA and Ni2+ forms a coordination complex that can be called NTA-Ni. The protein nanopore comprising NTA-Ni can also be called a protein nanopore modified by NTA-Ni. In some embodiments, the protein nanopore (especially the heterogeneous protein nanopore) of the present invention comprises a single reactive amino acid residue and comprises a single sensing moiety that is attached to the reactive amino acid residue via a single ligand. In more preferred embodiments, the protein nanopore (especially the heterogeneous protein nanopore) of the present invention comprises a single reactive amino acid residue and comprises a single NTA-Ni attached to the single reactive amino acid residue, wherein “single NTA-Ni” refers to a coordination complex consisting of a single NTA and a single Ni2+. In some embodiments, the single reactive amino acid residue may be cysteine or methionine.
[0188] The method of the present invention, which allows identification of all kinds of amino acids and almost all kinds of peptides with very high resolution by means of interaction between the sensing moiety and the peptide and / or the amino acid, in particular dynamically reversible interaction accompanied by bond-forming and bond-breaking processes, and each peptide or amino acid is capable of presenting a unique ionic current signal pattern.
[0189] In the method of the present invention, means for linearizing the peptide or pulling the peptide to move through the nanopore or controlling the movement speed of the peptide are unnecessary. In some embodiments, such means are not used, e.g., the peptide is not linked to a nucleic acid or nucleic acid analog, and motor protein (e.g., helicase, etc. ) is not used to control the movement of the peptide; e.g., unfoldase (e.g., ClpX or ClpXP) is not used. In some embodiments, the peptide is not labeled (e.g., fluorescently labeled) when characterizing the peptide using the methods of the present invention. In some embodiments, in the method of the present invention, the peptide, or other molecule tested along with the peptide, when interacting with the sensing moiety, the entire molecule thereof should be accommodated in the detecting portion of the nanopore, i.e., what is being detected is the target analyte in its entirety (i.e., the intact molecule) , and the ionic current that is being measured comprises the blockade current induced by blockage of the intact molecule of the target analyte in the nanopore.
[0190] The peptide, as a target analyte that can be characterized by the method described herein, may be a cyclic peptide or a linear peptide.
[0191] In some embodiments, the peptide that can be characterized using this method may have about 2 to about 50 amino acid residues in length, e.g., 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49 or 50 amino acid residues in length.
[0192] In some embodiments, the peptide that can be characterized using this method may have a molecule weight of about 150 Da to about 10 kDa, e.g., about 150 Da to about 9.5 kDa, about 150 Da to about 9 kDa, about 150 Da to about 8.5 kDa, about 150 Da to about 8 kDa, about 150 Da to about 7.5 kDa, about 150 Da to about 7 kDa, about 150 Da to about 6.5 kDa, about 150 Da to about 6 kDa, about 150 Da to about 5.5 kDa, about 150 Da to about 5 kDa, about 150 Da to about 4.5 kDa or about 150 Da to about 4 kDa.
[0193] The peptide and / or the amino acid described herein may be modified or unmodified. The peptide may have one or more modification. In some embodiments, the modification may be a chemical modification. In some embodiments, the modification may be a post-translational modification.
[0194] In some embodiments, the modification (e.g., post-translational modification) may comprise, but be not limited to, acetylation; alkylation (such as methylation) ; amidation; phosphorylation; glycosylation (such as an addition of N-GlcNAc) ; hydroxylation; propionylation; glutathionylation; nitrosylation; sulfenylation; sulfinylation; sulfonylation; succinylation; sulfation; formation of disulfide bridge; or any combination thereof.
[0195] The peptide may comprise one or more moieties, groups and / or atoms that can interact with the sensing moiety (or the adapter) , e.g., by means of coordination therebetween. In some embodiments, the one or more moieties, groups and / or atoms may be located at the N-terminal, C-terminal, side chain, and / or backbone of the peptide. In some embodiments, the one or more moieties, groups may be selected from one or more amino groups (-NH2) , one or more amide groups (-CO-NH2) , one or more carboxyl groups (-COOH) , one or more imidazolyl groups one or more mercapto (thiol) groups (-SH) , one or more peptide bond / amide bond (-CO-NH-) moieties and any combination thereof. In some embodiments, the one or more atoms may be selected from one or more oxygen atoms (O) , one or more nitrogen atoms (N) and / or one or more sulfur atoms (S) , such as provide by an amino group, an amide group, a carboxyl group, an imidazolyl group, a sulfhydryl group, and / or a peptide bond / amide bond moiety.
[0196] The present inventors have found that the N atom in the N-terminal amino group and the N atom in the N-terminal amide bond moiety (i.e., the first amide bond moiety from the N-terminal of a peptide) of a peptide can be reversibly coordinated with a sensing moiety (e.g., an NTA-Ni adapter) , which can result in a measurable and unique current signal in the nanopore, thereby making it possible to detect almost all peptides using said sensing moiety.
[0197] A person skilled in the art will understand that one or more amino groups (-NH2) , one or more amide groups (-CO-NH2) , one or more carboxyl groups (-COOH) , one or more imidazolyl groups one or more mercapto (thiol) groups (-SH) and / or one or more peptide bond / amide bond (-CO-NH-) moieties that is located at the C-terminal, side chain, and / or other position of the backbone of the peptide may also be reversibly coordinated with the sensing moiety through N, S and / or O atoms in these groups and / or moieties, which can also generate unique ionic current signals. These groups and / or moieties may be provided by certain amino acids (e.g., histidine, lysine, cysteine, methionine) . Furthermore, when the peptide comprises certain amino acids (e.g., proline) having certain structures (e.g., cyclic imide structures) , they can further generate unique ionic current signals; and when the peptide comprises some amino acids having different charges (e.g., positive or negative charges) , they can further generate unique ionic current signals. Such peptides can provide more information for peptide characterization and help to identify or discriminate different peptides.
[0198] In some embodiment, the peptide comprises a N atom in the N-terminal amino group and a N atom in the N-terminal amide bond moiety (i.e., the first amide bond moiety from the N-terminal of the peptide) .
[0199] In some embodiments, the peptide may comprise one or more amino groups (-NH2) , one or more amide groups (-CO-NH2) , one or more carboxyl groups (-COOH) , one or more imidazolyl groups and / or one or more mercapto (thiol) groups (-SH) that is located at the C-terminal and / or side chain of the peptide.
[0200] In some embodiments, the peptide may comprise one or more amino acids selected from histidine, lysine, cysteine, methionine and proline.
[0201] In some embodiments, the peptide does not comprise a polyhistidine, e.g., a His-Tag.
[0202] The one or more target analytes may be individual peptide.
[0203] The nanopore of the present invention can discriminate between different peptides. Therefore, the one or more target analytes may be one or two or more peptides, and the method of the present invention may be used to characterize two or more of peptides simultaneously.
[0204] The nanopore of the present invention can also discriminate a peptide from an amino acid or other molecules that can be characterized by the nanopore of the present invention. Therefore, the one or more target analytes may be one or more peptides and one or more of other molecules (such as amino acids) , the method of the present invention may be used to characterize one or more of amino acids and one or more of other molecules (such as amino acids) simultaneously.
[0205] The protein nanopore of the present invention may be disposed in a membrane that separates a first conductive liquid medium from a second conductive liquid medium, which may be called a nanopore system. The channel of the nanopore is the only path for the first conductive liquid medium and the second conductive liquid medium to communicate. Generally, a target analyte is added in at least one of the first conductive liquid medium and the second conductive liquid medium. The membrane can be an organic membrane, such as a lipid bilayer, or a synthetic membrane, such as a membrane formed of a polymeric material. The thickness of the membrane through which the nanopore extends can range from 1 nm to around 10 μm.
[0206] The preparation of a nanopore system is well known, for example, for a protein nanopore system, when a porin (such as the protein nanopore of the present invention) is placed in any one of the first conductive liquid medium and the second conductive liquid medium separated by a membrane (such as a lipid bilayer) , the porin can insert spontaneously into the membrane to form a nanopore.
[0207] The sensing moiety may be attached to the reactive amino acid residue before or after the porin insert in the membrane. For example, a sensing moiety can be attached to the reactive amino acid residue of the porin first, and then the porin comprising the sensing moiety can be inserted into the membrane, wherein the sensing moiety can be attached to the reactive amino acid residue by mix the sensing moiety and the porin together in a condition suitable for the binding of them. For another example, a porin without a sensing moiety can be inserted into the membrane first, and then a molecule comprising a sensing moiety is added in the first conductive liquid medium or the second conductive liquid medium and subsequently comes into contact with the reactive amino acid residue while moving across the nanopore and is thereby attached to the porin.
[0208] When the sensing moiety is attached to the reactive amino acid residue via a linker, the linker and the sensing moiety may be attached to the reactive amino acid residue before or after the porin insert in the membrane. For example, the linker and the sensing moiety can be attached to the reactive amino acid residue of the porin first, and then the porin comprising the sensing moiety can be inserted into the membrane to form a nanopore, wherein the linker can be attached to the reactive amino acid residue by mix the sensing moiety and the porin together in a condition suitable for the binding of them, and the sensing moiety can be bound to the linker by mix them together in a condition suitable for the interaction of them. For another example, a porin without a sensing moiety can be inserted into the membrane to form a nanopore first, and then a molecule comprising the linker is added in the first conductive liquid medium or the second conductive liquid medium and subsequently comes into contact with the reactive amino acid residue while moving across the nanopore and is thereby attached to the porin, then a molecule comprising the sensing moiety is added in the first conductive liquid medium or the second conductive liquid medium and subsequently comes into contact with the linker while moving across the nanopore and is thereby bound to the linker. The linker can be attached to the reactive amino acid residue by mix the sensing moiety and the porin together in a condition suitable for the binding of them. The sensing moiety can be bound to the linker by mix them together in a condition suitable for the interaction of them.
[0209] The analyte (e.g., a peptide) to be characterized may be added in either side of the nanopore, i.e., the first conductive liquid medium or the second conductive liquid medium. In some embodiments, the final concentration of the peptide added may range from about 1nM to about 100mM, e.g., about 10nM to about 10mM, about 100nM to about 10mM, about 0.001mM to about 10mM, about 0.01mM to about 5mM, about 0.02mM to about 5mM, 0.02mM to about 2mM, about 0.1mM, about 0.2mM, about 0.3mM, about 0.5mM, about 1.0mM, about 1.5mM. The appropriate concentration of different analytes may vary and can be determined experimentally.
[0210] It should be understood that in the method of the present invention, it is not necessary that the target analyte (e.g., a peptide and / or an amino acid) is driven to translocate through the nanopore. The target analyte (e.g., a peptide and / or an amino acid) can be characterized as long as it is able to migrate to the detection region of the nanopore and contact and interact with the sensing moiety in that region. The detection region of a nanopore is the region into which an analyte can induce a measurable change in ionic current across the nanopore when it migrates. The detection region of a nanopore may include the interior of the nanopore (e.g., a constriction region or vestibule of the nanopore) , or in the vicinity of its opening (orifice) , which are known to a person skilled in the art and can be readily detected and determined by means of measuring an ionic current signal. In some embodiments, migrating to the detection region includes entering the nanopore.
[0211] Means for migrating the analyte to the detection region are known to a person skilled in the art, for example, the analyte can be driven in the direction of entering the nanopore (although actual entry is not necessary) , which can be achieved for example by diffusion (concentration difference) , electrophoresis and / or electroosmotic flow. In some embodiments, an electrical potential difference (also called a voltage or an electric field) can be applied across the nanopore, i.e., applied between the first conductive liquid medium and the second conductive liquid medium. The electrical potential difference may be no less than about 1mV, no less than about 5mV, no less than about 10mV, no less than about 20mV, no less than about 40mV, no less than about 60mV, no less than about 80mV, no less than about 100mV, no less than about 120mV, no less than about 140mV, no less than about 160mV, no less than about 180mV or no less than about 200mV; or range from about 1mV to about 220mV, range from about 10mV to about 200mV, range from about 20mV to about 150mV, range from about 40mV to about 150mV, or range from about 60mV to about 120mV, or range from about 80mM to about 100mV.
[0212] In some embodiments, the electrical potential difference across the nanopore may vary or remain constant. Process and apparatus for applying an electric field to a nanopore are known to the person skilled in the art. For example, a pair of electrodes may be used to applying an electric field to a nanopore. As will be understood, the voltage range that can be used can depend on the type of nanopore system and the analyte being used.
[0213] The ionic current signals being measured comprise ionic current signals caused by blockage of the nanopore by the analyte (i.e., blockage of the ionic current) . The blockage of the ionic current may be related to the sequence, structure, conformation and / or modification of the analyte (e.g., a peptide and / or an amino acid) .
[0214] The measurement of the ionic current can be performed at any suitable temperature, such as -4℃-100℃, e.g., 4℃-50℃, 5℃-25℃ or room temperature.
[0215] Measurement of the ionic current through a nanopore are well known in the art and may be performed by way of optical signal or electric current signal. For example, one or more measurement electrodes could be used to measure the current through the nanopore. These can be, for example, a patch-clamp amplifier or a data acquisition device.
[0216] The liquid medium may include aqueous, organic-aqueous, and organic-only liquid media. Organic media include, e.g., methanol, ethanol, dimethylsulfoxide, and mixtures thereof. Liquids employable in methods described herein are well-known in the art. Descriptions and examples of such media, including conductive liquid media, are provided in U.S. Pat. No. 7,189,503, for example, which is incorporated herein by reference in its entirety. Salts, detergents, or buffering agents may be added to such media. Such agents may be employed to alter pH or ionic strength of the liquid medium. In some embodiments, the salt may comprise KCl. In some embodiments, the concentration of the salt may be 0.5 M-2.5 M. In some embodiments, the concentration of KCl is about 1.5 M. The buffering agent may be HEPES, MOPS, CHES or Tris, etc. The pH of the first conductive liquid medium and / or the second conductive liquid medium may range from about 1.0 to about 13.0, preferably from about 6.0 to about 9.0, preferably from about 6.0 to about 8.0, preferably from about 7.0 to about 7.4, which may depend on the desired charge properties of the analyte. In some embodiments, the first conductive liquid medium and / or the second conductive liquid medium does not contain Tris. In some embodiments, the first conductive liquid medium and / or the second conductive liquid medium comprises 1.5 M KCl, 10 mM MOPS and has a pH of about 7.0. In some embodiments, the first conductive liquid medium and / or the second conductive liquid medium comprises 1.5 M KCl, 10 mM HEPES and has a pH of about 7.0. In some embodiments, the first conductive liquid medium and / or the second conductive liquid medium comprises 1.5 M KCl, 10 mM CHES and has a pH of about 9.0.
[0217] A current pattern and a current trace, as used herein, may be used interchangeably, refer to the ionic current over time. A current pattern may comprise many individual blockade events, which may be of the same type or different types. Features about distribution, frequency, amplitude, etc. of the blockade events can be learned from the current pattern.
[0218] A variety of event features can be obtained from the current pattern. The event features may include a variety of characteristic parameters derived (or extracted) from the current pattern, which include, but are not limit to, open pore current (I0) , event current (Ipep) , event standard deviation (SD) , inter-event interval (ton) , event dwell time (toff) , mean dwell time (τoff) , mean inter-event interval (τon) , blockage amplitude (ΔI, defined as ΔI = Ipep -I0) , mean and the mean percentage blockage (defined as ΔI / Ip) , skewness (skew) , kurtosis (kurt) or any combination thereof, which can be used to determine the characteristic of the analyte (e.g., a peptide and / or an amino acid) . The event features may also be subjected to further data processing (e.g., statistical analysis, such as cluster analysis) of a set of characteristic parameters, which are able to provide additional information for determining the characteristic of the analyte. In some embodiments, the further data processing may include performing a statistical analysis of a set of characteristic parameters from multiple blockade events, e.g., a cluster analysis on a set of characteristic parameters (e.g., ΔI and SD) from multiple blockade events, which may be used to provide a fingerprint of one or more target analytes. In some embodiments, the cluster analysis may be selected from the group consisting of: hierarchical clustering, k-mean clustering, distribution-based clustering, and density-based clustering. In some embodiments, the cluster analysis may be a density-based spatial clustering of applications with noise (DBSCAN) .
[0219] The characteristic of the target analyte (e.g., a peptide and / or an amino acid) may include, but is not limited to, the identity of the target analyte, the presence or absence of the target analyte, the structure of the target analyte, the conformation of the target analyte, the sequence of the target analyte (e.g., the amino acid sequence of a peptide) , the modification of the target analyte, the purity of the target analyte, the fingerprint of the target analyte, etc., or any combination thereof.
[0220] It should be understood that when used to characterize more than one target analyte, the measured current pattern may comprise a plurality of different current patterns representing single target analyte respectively, and the measured current pattern may be considered to be a set of multiple current patterns. Current patterns representative of each target analyte may be extracted therefrom and used to characterize each target analyte. Alternatively, a fingerprint spectrum of more than one target analyte may be determined based on the set of multiple current patterns (i.e., representative of all target analytes) .
[0221] In some embodiments, the representative current pattern of the target analyte may be compared to a reference current pattern to determine the characteristic of the target analyte (e.g., a peptide and / or an amino acid) , wherein the reference current pattern is a representative current pattern of a reference analyte. By comparison, it is possible to determine whether the target analyte has the same characteristic as or different characteristics from the reference analyte.
[0222] In some embodiments, the reference analyte may be a known peptide or a known amino acid, which may or may not have a modification, and whose sequence, structure and / or modification may be known or unknown; or a mixture of one or more known peptides and / or one or more known amino acids, which optionally can be derived from fragmentation of a known polypeptide.
[0223] In some embodiments, a comparison between the representative current pattern of the target analyte and a reference current pattern may include the comparison between event features derived from the representative current pattern of the target analyte, or information obtained by further data processing, and corresponding event features derived from the reference current pattern or information obtained by further data processing thereof. In some embodiments, the reference current pattern may include the information of the event features derived from the reference current pattern (which may be called reference event features) and / or information obtained by further data processing thereof. In some embodiments, the reference current pattern may include a fingerprint of the reference analyte, which may be called reference fingerprint. The comparison may be qualitative or quantitative.
[0224] The reference current pattern can be obtained by nanopore detection of the reference analyte using the same methods as those used for detection of the target analyte, including the use of the same nanopore, and the same measurement methods.
[0225] In some embodiments, a reference database may be provided that contains reference current patterns for a plurality of different reference analytes. Searching and comparing the representative current patterns of the target analyte to the reference current patterns in this database allows determination of the characteristic of the target analyte.
[0226] In some embodiments, the step (ii) of the method comprises measuring an ionic current as the target analyte migrates to provide a representative current pattern of the target analyte, comparing the representative current pattern of the target analyte to the reference current patterns in the reference database, and determine the characteristic of the target analyte based on the comparison.
[0227] In some cases, when the number of records (i.e., reference current patterns) contained in a reference database is very large, comparing a representative current pattern of an analyte to all of the reference current patterns one by one would be a time-consuming task. Therefore, the present inventors use an assisting-assay to downsize the reference database, thereby making the data processing of the method of the present invention more efficient.
[0228] In some embodiments, the method may further comprise performing an assisting-assay on the one or more target analytes and reducing the amount of information contained in the reference database based on the results of the assisting-assay, thereby improving the efficiency of the database comparison and further optimizing the method.
[0229] The assisting-assay comprises obtaining knowledge of the amino acid composition of the one or more target analytes, i.e., obtaining the identity of each of all amino acids constituting the target one or more target analytes, wherein each amino acid may be with modifications or unmodified.
[0230] In some embodiments, the assisting-assay may comprise hydrolyzing the one or more target analytes into an amino acid mixture using an exopeptidase (e.g., an aminopeptidase and / or a carboxypeptidase) , and then detecting said amino acid mixture to obtain the identity of each of all amino acids contained therein. Examples of aminopeptidase includes, but is not limited to, proline aminopeptidase, aminopeptidase N (APN) , aminopeptidase A (APA) , leucine aminopeptidase, lysine aminopeptidase, etc. In some embodiments, one, two, or more aminopeptidases may be used to hydrolyze the target analytes, e.g., such that the target analyte is completely hydrolyzed to amino acids.
[0231] In some embodiments, the amino acid mixture is detected using nanopores. In some embodiments, the amino acid mixture is detected using a nanopore of the present invention. Nanopores with the addition of metal ions as the sensing moiety can distinguish all amino acids as well as amino acids with different modifications, as has been demonstrated in prior literature (e.g., WO2023056960A1; Wang, K., Zhang, S., Zhou, X. et al. Unambiguous discrimination of all 20 proteinogenic amino acids and their modifications by nanopore. Nat Methods 21, 92–101 (2024) . https: / / doi. org / 10.1038 / s41592-023-02021-8; which are hereby incorporated herein by reference in their entirety) .
[0232] In some embodiments, the detection of the amino acid mixture comprises causing each amino acid of the amino acid mixture to contact the sensing moiety in a region that allows each amino acid of the amino acid mixture to be detected based on a change of an ionic current across the nanopore, and measuring an ionic current as each amino acid of the amino acid mixture migrates to provide a representative current pattern of the amino acid mixture, wherein the current pattern indicates the identity of each amino acid in the mixture.
[0233] In some embodiments, the representative current pattern of the amino acid mixture may be compared to reference current patterns of specified reference amino acids to determine the identity of each amino acid in the mixture.
[0234] In some embodiments, a comparison between the representative current pattern of the amino acid mixture and reference current patterns of specified amino acids may include the comparison between event features derived from the representative current pattern of the amino acid mixture, and corresponding event features derived from the reference current pattern of each specified amino acid. The comparison may be qualitative or quantitative.
[0235] The reference current pattern of each specified amino acid can be obtained by nanopore detection of each reference amino acid using the same methods as those used for detection of the amino acid mixture, including the use of the same nanopore, and the same measurement methods.
[0236] In some embodiments, the representative current pattern of the amino acid mixture may be compared to a reference database that contains reference current patterns of all kinds of amino acids.
[0237] After identifying all amino acids constituting the target analyte, only records of amino acids and peptides containing the identified amino acid are retained (e.g., from an initial reference database) to form a simplified reference database, which can then be used as a reference database to determine the characteristic of the one or more target analytes.
[0238] In a second aspect, the present invention provides a method of characterizing a target polypeptide, comprising:
[0239] (a) fragmenting the target polypeptide to obtain a peptide fragment mixture;
[0240] (b) characterizing the peptide fragment mixture using the method of the first aspect of the present invention to determine the characteristic of the peptide fragment mixture; and
[0241] (c) determining the characteristic of the target polypeptide based one the characteristic of the peptide fragment mixture.
[0242] In some embodiments, the step (b) comprises causing each component of the peptide fragment mixture to contact the sensing moiety in a region that allows each component of the peptide fragment mixture to be detected based on a change of an ionic current across the nanopore, measuring an ionic current as each component of the peptide fragment mixture migrates to provide a representative current pattern of the peptide fragment mixture, and determining the characteristic of the peptide fragment mixture based on the representative current pattern of the peptide fragment mixture.
[0243] This method can characterize a target polypeptide that comprise a relatively large numbers of amino acid residues.
[0244] It should be understood that fragmentation of a polypeptide may produce an amino acid, and thus the peptide fragment mixture may also include one or more invidual amino acids. In some embodiments, the peptide fragment mixture may comprise a variety of peptides or may comprise a variety of peptides and a variety of amino acids. The determination of the characteristic of the peptide fragment mixture may include the determination of the characteristic of each component (i.e., each peptide and each amino acid) in the mixture. It should be understood that the reference of the term “amino acid sequence” in the context of individual amino acid refers to the identity of the individual amino acid, i.e., which specific amino acid it is. The method of the present invention allows the identification of peptides and amino acids simultaneously and thus the characterization of a variety of different mixtures of peptides and amino acids.
[0245] In some embodiments, the target polypeptide that can be characterized using this method may have at least 2, at least 3, at least 5, at least 10, at least 15, at least 20, at least 25, at least 30, at least 35, at least 40, at least 45, at least 50, at least 75, at least 100 , at least 150, at least 200, at least 300, at least 400, at least 500 amino acid residues or even more, or even thousands of amino acid residues, as long as they can be fragmented to peptide fragments of appropriate lengths that can be characterized using the method of the first aspect of the present invention, i.e., directly detected for the intact molecule thereof by the nanopore of the present invention.
[0246] In some embodiments, each peptide fragment of these peptide fragments, or at least a portion thereof, is capable of being characterized using the method of the first aspect of the present invention, i.e., being characterized as a single intact molecule.
[0247] In some embodiments, each peptide fragment of these peptide fragments, or at least a portion thereof, is an amino acid or a peptide comprising 2-50 amino acid residues in length, e.g., 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49 or 50 amino acid residues in length.
[0248] The target polypeptide may be a cyclic polypeptide or a linear polypeptide.
[0249] The target polypeptide may be modified or unmodified. The target polypeptide may have one or more modification. In some embodiments, the modification may be a chemical modification. In some embodiments, the modification may be a post-translational modification. In some embodiments, the modification (e.g., post-translational modification) may comprise, but be not limited to, acetylation; alkylation (such as methylation) ; amidation; phosphorylation; glycosylation (such as an addition of N-GlcNAc) ; hydroxylation; propionylation; glutathionylation; nitrosylation; sulfenylation; sulfinylation; sulfonylation; succinylation; sulfation; formation of disulfide bridge; or any combination thereof.
[0250] The target polypeptide may be fragmented by any suitable method, e.g., enzyme digestion or being fragmented chemically or physically or any combination thereof.
[0251] In some embodiments, the target polypeptide may be fragmented by protease (endopeptidase) hydrolysis. The protease may be selected from trypsin, thermolysin, chymotrypsin, Lys-C, Arg-C, Asp-N, Glu-C, elastase and pepsin. For different polypeptides, a suitable protease can be selected experimentally to fragment the polypeptide into peptide fragments with a suitable length. In some embodiments, the target polypeptide is digested (hydrolyzed) by one or more proteases. In some embodiments, the target polypeptide is digested (hydrolyzed) by two or more proteases. In some embodiments, the target polypeptide is digested (hydrolyzed) by trypsin and thermolysin.
[0252] In some embodiments, the target polypeptide may be fragmented chemically and / or physically, e.g., by heating or freezing, treatment with acids or bases, or combination thereof.
[0253] The representative current pattern of the peptide fragment mixture may also be considered as a representative current pattern of the target polypeptide, which comprises a plurality of different current patterns representing each single peptide fragment.
[0254] The characteristic of the target polypeptide that can be determined by this method includes, but is not limited to, the identity of the target polypeptide, the presence or absence of the target polypeptide, the amino acid sequence of the target polypeptide (e.g., full-length or partial amino acid sequence) , the amino acid alteration (including insertion, substitution and / or deletion of one or more amino acid) in the target polypeptide (e.g., relative to a specific polypeptide) , the modification on the target polypeptide (including the type of the modification and the position of the modification) , the fingerprint of the target polypeptide, etc., or any combination thereof.
[0255] In some embodiments, the representative current pattern of the peptide fragment mixture may be compared with a variety of reference current pattern to determine the characteristic of the peptide fragments. By comparison, it is possible to determine the characteristic of individual peptide fragment and / or the characteristic of the mixture as a whole (i.e., a profile) , including a fingerprint of the mixture.
[0256] In some embodiments, a reference database that contains reference current patterns for a plurality of different reference analytes may be used to determine the characteristic of the peptide fragment mixture.
[0257] In some embodiments, the step (b) of the method comprises measuring an ionic current as each component of the peptide fragment mixture migrates to provide a representative current pattern of the peptide fragment mixture, comparing representative current pattern of the peptide fragment mixture to reference current patterns in a reference database, and determine the characteristic of the peptide fragment mixture based on the comparison.
[0258] In some cases, the method may further comprise performing an assisting-assay on the target polypeptide and reducing the amount of information contained in the reference database based on the results of the assisting-assay, thereby improving the efficiency of the database comparison and further optimizing the method.
[0259] The assisting-assay comprises obtaining knowledge of the amino acid composition of the target polypeptide, i.e., obtaining the identity of each of all amino acids constituting the target polypeptide, wherein each amino acid may be with modifications or unmodified.
[0260] In some embodiments, the assisting-assay may comprise hydrolyzing the target polypeptide into an amino acid mixture using an exopeptidase (e.g., an aminopeptidase and / or a carboxypeptidase) , and then detecting the amino acid mixture to obtain the identity of each component of the amino acid mixture. In some embodiments, one, two, or more aminopeptidases may be used to hydrolyze the target polypeptide, e.g., such that the target polypeptide is completely hydrolyzed to amino acids.
[0261] In some embodiments, the amino acid mixture is detected using nanopores. In some embodiments, the amino acid mixture is detected using a nanopore of the present invention. Nanopores with the addition of metal ions as the sensing moiety can distinguish all amino acids as well as amino acids with different modifications, as has been demonstrated in prior literature (e.g., WO2023056960A1; Wang, K., Zhang, S., Zhou, X. et al. Unambiguous discrimination of all 20 proteinogenic amino acids and their modifications by nanopore. Nat Methods 21, 92–101 (2024) . https: / / doi. org / 10.1038 / s41592-023-02021-8; which are hereby incorporated herein by reference in their entirety) .
[0262] In some embodiments, the detection of the amino acid mixture comprises causing each amino acid of the amino acid mixture to contact the sensing moiety in a region that allows each amino acid of the amino acid mixture to be detected based on a change of an ionic current across the nanopore, measuring an ionic current as each amino acid of the amino acid mixture migrates to provide a representative current pattern of the amino acid mixture, wherein the current pattern indicates the identity of each amino acid in the mixture.
[0263] After identifying all amino acids constituting the target polypeptide, only records of amino acids and peptides containing the identified amino acid are retained (e.g., from an initial reference database) , which can then be used as a reference database to determine the characteristic of the peptide fragment mixture and / or the target polypeptide.
[0264] This method may be used to sequence a target polypeptide. For example, the amino acid sequence of each component or at least a portion of components of the peptide fragment mixture can be identified, which can then be assembled to obtain the amino acid sequence of the target polypeptide.
[0265] In some embodiments, the step (b) includes identifying the amino acid sequence of each component in the peptide fragment mixture using the method of the first aspect of the present invention; and the step (c) includes determining the amino acid sequence of the target polypeptide based on the amino acid sequence of each component in the peptide fragment mixture.
[0266] In some embodiments, the amino acid sequence of each component or at least a portion of the components of the peptide fragments is identified by comparing the representative current pattern of the peptide fragment mixture to a variety of reference current patterns or a reference database, wherein the reference current patterns are representative of a variety of peptides whose amino acid sequences and / or modification have been known.
[0267] If the target polypeptide to be sequenced has a modification on it, information about said modification (including the type of modification and the location of the modification) can be obtained along with the determination of the amino acid sequence of the peptide fragments, whereby information about the modification in the amino acid sequence of the target polypeptide can be determined at the same time as the target polypeptide is sequenced.
[0268] In some embodiments, the target polypeptide may be fragmented by multiple (two of more) fragmentation operations to obtain multiple peptide fragment mixtures.
[0269] The multiple peptide fragment mixtures are separately subjected to step (b) to identify the amino acid sequences of all or some of the peptide fragments in each peptide fragment mixture, and these sequences are subsequently assembled to obtain the amino acid sequence of the target polypeptide.
[0270] The multiple fragmentation operations may be consecutive, i.e. the latter fragmentation operation is performed on the product of the previous fragmentation operation; or separate, i.e. each fragmentation operation is performed on the target polypeptide; or some of the fragmentation operations are consecutive. The multiple fragmentation operations may be performed using the same means (i.e., digestion using the same protease, or degradation using the same physical or chemical means) or using different means, or some of the fragmentation operation are performed using the same means.
[0271] In some embodiments, the multiple fragmentation operations may include digestion with two or more proteases, e.g., consecutively or seperately. In some embodiments, the target polypeptide may be digested by trypsin and thermolysin, respectively.
[0272] The assembly of a polypeptide sequence based on the sequence of peptide fragments is readily achievable for a person skilled in the art, for example, the sequence of the entire polypeptide can be concluded from the order of amino acid arrangement of each peptide fragment, and the overlap between different peptide fragments, which is analogous to the process of assembling a sequence of fragments into a full-length sequence in the mass spectrometry sequencing technique for proteins.
[0273] Using the sequencing method described above, not only is it possible to obtain the amino acid sequence of the target polypeptide, but it is also possible to directly obtain the modifications contained in the sequence, which is simpler and more accurate than the protein mass spectrometry sequencing method.
[0274] This method may also be used to identify a fingerprint of the target polypeptide. In some embodiments, the fingerprint of the target polypeptide may comprise a fingerprint of the peptide fragment mixture obtained from the fragmentation of the target polypeptide. The fingerprint of the peptide fragment mixture may comprise a set of event features, which can be identified from the representative current pattern of the peptide fragment mixture. In some embodiments, the fingerprint of the target polypeptide may comprise a set of results obtained by further data processing (e.g., statistical analysis, such as cluster analysis) of a set of event features from the representative current pattern of the peptide fragment mixture. In some embodiments, the fingerprint of the target polypeptide may comprise the identity or the amino acid sequence of each component of the peptide fragment mixture obtained from the fragmentation of the target polypeptide.
[0275] In some embodiments, the fingerprint of the target polypeptide can be further compared to a fingerprint of a reference polypeptide (which can be also called a reference fingerprint) to determine a relationship between the two polypeptides, for example, whether the target polypeptide is identical to or different from the reference polypeptide (i.e., determining the identity of the target polypeptide) , or whether the target polypeptide has amino acid alteration and / or modifications relative to the reference polypeptide.
[0276] In some embodiments, the fingerprint of the target polypeptide can be further compared to fingerprint of a variety of reference polypeptides to determine the identity of the target polypeptide. The fingerprint of a variety of reference polypeptides can be provided in a form of a reference database. Searching and comparing the representative fingerprint of the target polypeptide to the reference fingerprint of known polypeptide in this database allows determination of the identity of the target polypeptide.
[0277] In some cases, it is not necessary to know the sequence of the reference polypeptide or to sequence the target polypeptide in order to perform the above-described assay.
[0278] In some other cases, if the amino acid sequence and / or modification of each component represented by the fingerprint is known, the amino acid sequence and / or modification of the target polypeptide can be determined from the fingerprint, and the difference in the amino acid sequence and / or modification of the target polypeptide relative to the reference polypeptide may also be determined.
[0279] The fingerprint of the reference polypeptide can be obtained by fragmentation and nanopore detection of the reference polypeptide using the same methods as those used for detection of the target polypeptide, including the use of the same fragmentation means (e.g., the same protease) , the same nanopore, and the same measurement methods.
[0280] In some embodiments, in order to avoid interference by the protease, after the target peptide has been cleaved (e.g., digested by an endopeptidase into a peptide fragment mixture, or by an exopeptidase into an amino acid mixture) , and prior to the subsequent protein nanopore detection, in order to avoid interference by the protease, a step of filtering the mixture to remove the protease can be included.
[0281] In a third aspect, a method for detection of the purity of a target polypeptide sample is provided, which may utilize the methods of the preceding first or second aspect. The method may be used to detect the purity of the target polypeptide sample, which can used in the quality control of a polypeptide product.
[0282] The target polypeptide sample may be any sample that comprise the target polypeptide, especially a sample comprising the target polypeptide as the primary component (e.g., a sample in which the target polypeptide constitutes more than 80%or 90%of all components) , e.g., a sample obtained for the preparation of the target polypeptide, which may be a sample obtained by chemical synthesis (e.g., solid-phase peptide synthesis, SPPS) or purification (e.g., by filtration, precipitation, electrophoresis, and / or chromatographic techniques) of the target peptide.
[0283] These kinds of samples sometimes contain impurities. The impurities may be proteinogenic, which may include a peptide different from the target polypeptide and / or an amino acid, such as fragments resulting from degradation of the target polypeptide. The presence or absence of such impurities in the target polypeptide sample can be detected by identifying the fingerprint of the target polypeptide sample.
[0284] In some embodiments, the impurities to be detected may comprise a peptide of 2-50 amino acid residues in length, e.g., 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49 or 50 amino acid residues in length.
[0285] In some embodiments, the target polypeptide in the sample may have 2-50 amino acid residues in length, e.g., 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49 or 50 amino acid residues in length. In some embodiments, the target polypeptide in the sample may have at least 2, at least 53, at least 5, at least 10, at least 15, at least 20, at least 25, at least 30, at least 35, at least 40, at least 45, at least 50, at least 75, at least 100, at least 150, at least 200, at least 300, at least 400, at least 500 amino acid residues or even more, or even thousands of amino acid residues.
[0286] In some embodiments, the purity of a target polypeptide sample can be detected using the method of the preceding first aspect, where the target polypeptide sample may comprise said one or more analytes.
[0287] The method may comprise:
[0288] (i) providing a nanopore incorporating at least one sensing moiety which is capable of interacting with a peptide, and causing each proteinogenic component of the target polypeptide sample to contact the sensing moiety in a region that allows each proteinogenic component of the target polypeptide sample to be detected based on a change of an ionic current across the nanopore;
[0289] (ii) measuring the ionic current as each proteinogenic component of the target polypeptide sample migrates to provide a representative current pattern of the target polypeptide sample, and obtain a fingerprint of the target polypeptide sample; and
[0290] (iii) comparing the fingerprint of the target polypeptide sample with a reference fingerprint of a reference polypeptide (which has a desired amino acid sequence for the target polypeptide) and determining the purity of the target polypeptide sample.
[0291] If the fingerprint of the target polypeptide sample comprises one or more additional results that are not contained in the fingerprint of the reference polypeptide, it means that the target polypeptide sample contains impurities. If the fingerprint of the target polypeptide sample comprise no additional results compared to the fingerprint of the reference polypeptide, it may mean that the target polypeptide sample contains no or few impurities. The more the additional results, the lower the purity of the target polypeptide sample, and the less the additional results, the higher the purity of the target polypeptide sample.
[0292] In some embodiments, the impurities may be detected qualitatively or quantitatively.
[0293] In some embodiments, for a sample of a target polypeptide having more amino acid residues, e.g., more than 50 amino acid residues in length, may also be fragmented into a peptide fragment mixture prior to being with nanopores, and then be detected by a nanopore, as described in the preceding method of the third aspect. In such cases, the reference polypeptide may also be fragmented in the same manner, as described in the preceding method of the third aspect.
[0294] In each aspect of the present invention, the determination of the characteristic of the target analyte (including, but not limiting to, peptide, amino acid, polypeptide, peptide fragment mixture, amino acid mixture, and other analyte, etc., as described herein) based on the representative current pattern of the target analyte may be performed using machine learning algorithm.
[0295] As can be understood by a person in the art, a reference database described herein may contain reference current patterns of a variety of peptides, amino acids, peptide fragment mixtures, polypeptides and / or other analytes, which is necessary for the determination of the characteristic of the various analytes being analyzed in the present invention.
[0296] Applications of machine learning algorithms are well known to a person skilled in the art. Generally, event features extracted from the current patterns of known analytes (e.g., a peptide and / or an amino acids) are input into a model for training, then a validation is performed to evaluate the model performance, and then the characteristic of the target analyte can be determined using the established model.
[0297] As known by a person skilled in the art, a model that can be use in the machine learning algorithm includes, but is not limited to, SVM, Bayes, hidden Markov, random forest, decision tree, bagged trees, etc.
[0298] The target analyte, such as peptide, amino acid, polypeptide and / or other analytes may be present in any suitable sample. The method of the first and second aspect of present invention may be performed on a sample that is known to contain or suspected to contain the target analyte. The method may be performed on a sample to confirm the identity of the analyte whose presence in the sample is known or expected. The sample may be a biological sample, e.g., extracted from an organism (e.g., prokaryotic or eukaryotic) ; or a chemical sample, e.g., a peptide or polypeptide sample that is synthesized by solid-phase peptide synthesis (SPPS) . In some embodiments, the sample may be a mixture of peptides or polypeptides in different form. In some embodiments, the sample may be a cell lysate. In some embodiments, the sample may be a proteinaceous food, drink, or food additive, such as protein powder, whey protein, etc.
[0299] The sample can be directly used in the methods of the present invention without any additional processing. Alternatively, the sample may be processed (pretreated) prior to being used in the methods of the present invention, for example by various purification means known in the art to isolate the target analyte. These may include, but are not limited to affinity binding methods, such as antibodies, or chromatographic methods, to isolate and purify specific components of the sample or remove unwanted background impurities. For a sample that contains a polypeptide to be fragmented into peptide fragments (e.g., for sequencing) , the method may include one or more sample preparation steps, for example, a pre-filtering step as done for other methods (e.g. Mass Spectrum) .
[0300] In another aspect, the present invention provides a method of characterizing a target polypeptide, comprising:
[0301] (1) preparing a target polypeptide sample including the target polypeptide;
[0302] (2) preparing different shortened polypeptide samples, wherein each shortened polypeptide sample including a shortened target polypeptide by truncating the target polypeptide, and lengths of the shortened target polypeptides in different shortened polypeptide samples are different;
[0303] (3) characterizing the different shortened polypeptides using the method as mentioned above, to determine the characteristic of the shortened target polypeptides separately; and
[0304] (4) determining the characteristic of the target polypeptide based on the characteristic of the different shortened polypeptides.
[0305] In some embodiments, the different shortened polypeptide samples are prepared by stepwise shortening the target polypeptide.
[0306] In some embodiments, the different shortened polypeptide samples are prepared by stepwise shortening the target polypeptide from N-terminus or C-terminus of the target polypeptide.
[0307] In some embodiments, the different shortened polypeptide samples are prepared by Edman degradation.
[0308] In some embodiments, the characteristic of the target polypeptide is the identity of the target polypeptide, the presence or absence of the target polypeptide, the amino acid sequence of the target polypeptide, the amino acid alteration in the target polypeptide, the modification in the target polypeptide, the fingerprint spectrum of the target polypeptide, the purity of the target polypeptide, or any combination thereof.
[0309] In another aspect, the present application provides a use of the method as mentioned above for determining characteristic of a neoantigen.
[0310] In some embodiments, the characteristic of the neoantigen is the identity of the neoantigen, the presence or absence of the neoantigen, the amino acid sequence of the neoantigen, the amino acid alteration in the neoantigen, the modification in the neoantigen, the fingerprint spectrum of the neoantigen, the purity of the neoantigen, or any combination thereof. TERM DEFINITION
[0311] The term “nanopore” generally refers to a pore, channel or passage which has a very small diameter on the order of nanometers and extends through a membrane. A nanopore may have a characteristic width or diameter on the order of 0.1 nanometers (nm) to about 1000 nm.
[0312] The term “protein nanopore” refers to a polypeptide subunit or a multimer of polypeptide subunits (each subunit may be called a monomer of the protein nanopore) that can form a channel through a membrane. The term “protein nanopore” includes wild-type nanopore or a variant of a wild-type nanopore. Sequences of wild type protein nanopore can be found in GenBank on https: / / www. ncbi. nlm. nih. gov / . A variety of variants of the above protein nanopores have been establish in recent years.
[0313] The term “variant” of a protein nanopore or a monomer refers to that the monomer, or one or more monomers of the nanopore have one or more additions, substitutions and / or deletions of amino acids compared to their parental ones, or may have a sequence identity of at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99%compared to their parental ones, wherein the parental protein or monomer may be a wild-type one, or homolog or variant thereof, and retains tunnel-forming capability. The term “variant” of a protein nanopore or a monomer may also refer to that the monomer, or one or more monomers of the nanopore have other modification (e.g., chemical modification) .
[0314] The term “identity” of an analyte refers to a relationship between the analyte and a known substance, e.g., whether the analyte is identical to a known substance.
[0315] The term “sequence identity” refers to a relationship between the sequences of two or more peptides or polynucleotides. The term “sequence identity” refers to the percentage of identical nucleotide or amino acid residues at corresponding positions in two or more sequences when the sequences are aligned to maximize sequence matching, i.e., taking into account gaps and insertions. The alignment of the sequences and the calculation of percentage of the sequence identity can be carried out with suitable computer programs known in the art. Such programs include, but are not limited to, BLAST, ALIGN, ClustalW, EMBOSS Needle, etc. An example of a local alignment program is BLAST (Basic Local Alignment Search Tool) with default parameters, which is available from the webpage of National Center for Biotechnology Information which can currently be found at http: / / www. ncbi. nlm. nih. gov / / and which was firstly described in Altschul et al. (1990) J. Mol. Biol. 215; 403-410. Examples of a global alignment program (which optimizes the alignment over the full-length of the sequences) are EMBOSS Needle and EMBOSS Stretcher programs based on the Needleman-Wunsch algorithm (Needleman, Saul B.; and Wunsch, Christian D. (1970) with default parameters, "A general method applicable to the search for similarities in the amino acid sequence of two proteins" , Journal of Molecular Biology 48 (3) : 443-53) , which are both available at http : / / www. ebi. ac. uk / Tools / psa / .
[0316] The term “moiety” refers to a chemical molecule or any part of a chemical molecule, including a functional group, an atom or group of chemically bonded atoms that is attached to another atom or molecule by one or more chemical bonds thereby forming part of a molecule, or an ion.
[0317] The term “sensing moiety” refers to a moiety which is capable of interacting with single molecule of a target analyte.
[0318] The term “interact” or "interaction" refers to reaction or binding between the sensing moiety and the target analyte, which may be reversible or irreversible.
[0319] The term “peptide” or “polypeptide” refers to a sequence of two or more amino acid residues joined together by peptide bonds.
[0320] The term “proteinogenic” refers to peptides, polypeptides, and / or amino acids that can make up a polypeptide.
[0321] The term “representative” current pattern, event feature or fingerprint of an analyte refers to that obtained by measuring the ionic current during the time the analyte is driven to migrate and cause the blockage of the ionic current across the nanopore.
[0322] The term “attach” , “link” , “connect” , “join” or grammatical variation thereof refers to that one element is either directly joined to another element, or else indirectly joined to another element through an intervening moiety or moieties, in a covalent or non-covalent manner. Covalent association include an association through coordination bond. Examples of non-covalent association includes electrostatic forces, hydrogen bonds, hydrophobic effects, van der Waals forces, etc. In some embodiments, the attachment of one element to another element can be reversible or irreversible.
[0323] The term “α-helical structure” refers to a type of structure of proteins that are twisted into a coil (ahelix) .
[0324] The term “β-barrel structure” refers to a type of structure of proteins formed by a large beta-sheet that twists and coils to form a closed structure in which the first and last strands are bonded by hydrogen bonds.
[0325] The term “to gate” or “gating” refers to the spontaneous change of electrical conductance through the tunnel of the protein that is usually temporary (e.g., lasting for as few as 1-10 milliseconds to up to a second) .
[0326] The term “ligand” refers to a compound capable of coordinating to a metal atom or ion.
[0327] The term “heterogeneous protein nanopore” refers to a protein nanopore in which at least one of the multiple monomers has a different structure (e.g., amino acid sequence and / or chemical modifications) from the other monomers.
[0328] The term “modified” or "modification” refers to a changed state or structure of a molecule. Molecules may be modified in many ways, including chemically, structurally, and functionally, for example, by replacement of the original molecule or a group with a different molecule or a group, or by introduction of a molecule or a group by covalent or coordinated attachment. In some embodiments, the term “modified” or "modification” includes amino acid alteration (including insertion, substitution and / or deletion of one or more amino acids) , and / or chemical modification of an amino acid.
[0329] The term “chemical modification” refers to addition or change of a chemical group or chemical moiety by a chemical method.
[0330] The term “reactive” refers to an ability to interact with a specific molecule. If an amino acid residue can interact with a first molecule but cannot interact with a second molecule, it is considered as being reactive to the first molecule and being non-reactive to the second molecule.
[0331] The term “coordination” refers to an interaction in which one multi-electron pair donor coordinately bonds, i.e., is “coordinated, ” to one metal ion. The term “coordination” refers to an interaction between an electron pair donor and a coordination site on a metal ion resulting in an attractive force between the electron pair donor and the metal ion. A coordinate bond may be formed between the electron pair donor and the metal ion. The electron pair donor may be a nonmetal atom, such as nitrogen, sulfur, phosphorus, carbon or oxygen, etc. A compound containing the electron pair donor may be referred as a ligand.
[0332] The term “coordination complex” " is a complex in which there is a coordinate bond between the metal ion and the electron pair donor, ligand or chelating group. Thus, ligand or chelating group is generally electron pair donor, molecule or molecular ion having unshared electron pairs available for donation to a metal ion.
[0333] The term “blockage of the ionic current” may also be called a “blockade current” , which is evidenced by a change in ionic current that is clearly distinguishable from noise fluctuations and is usually associated with the presence of an analyte molecule within the nanopore. The strength of the blockade, or change in current, will depend on a characteristic of the analyte. More particularly, “blockage” may refer to an interval where the ionic current drops to a level which is about 5-100%lower than the unblocked current level, remains there for a period of time, and returns spontaneously to the unblocked level. For example, the blockade current level may be about, at least about, or at most about 5%, 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, or 100%lower than the unblocked current level. A blockage may be called a blockade event or an event.
[0334] The term “event” , as used herein, refers to a blockage of the nanopore by a target analyte (i.e., an interval where the ionic current drops to a level which is about 5-100%lower than the first blockade current level, remains there for a period of time, and returns spontaneously to the unblocked current level) , and also refers to a current change caused by the blockage of the target analyte. The person skilled in the art know how to determine the occurrence of an event.
[0335] The term “bonding time” refers to the time from bond formation to bond breakage.
[0336] The term “linear peptide” refers to a peptide or polypeptide in which the amino acids are linked to one another via an amide bond formed between the alpha-amino group of one and the alpha-carboxylic group of another.
[0337] The term “cyclic peptide” refers to refers to a peptide having an intramolecular bond between two non-adjacent amino acids within a peptide to form a ring structure. The intramolecular bond includes, but is not limited to, a connection between the amino end and a side chain, the carboxyl end and a side chain, or two side chains or more complicated arrangements.
[0338] The term “fingerprint” or “fingerprint spectra” refers to a set of distinctive results obtained from detection of separated analytes in a sample by a nanopore. When a plurality of identical molecules of one analyte are present in the sample and more than one of them are detected, the fingerprint can comprise all the results of the identical molecules or alternatively only a representative value, e.g. average, median, of the results of identical molecules. The results may include one or more event features as described herein, or results obtained from further data processing of these event features as described herein, or may include identification results (e.g., identity or amino acid sequence, etc. ) of each analyte. The term "fingerprint" , in the context of a target polypeptide that is subjected to fragmentation, refers to a set of distinctive results obtained from separated components in the mixture obtained from the fragmentation of the target polypeptide.
[0339] The term “purity” refers to the percentage of a target peptide in the target peptide sample in relation to all the components contained in the sample (in particular all the proteogenic components) .
[0340] Inspired by the success of nanopore nucleic acid sequencing, it is widely anticipated that nanopores might also be suitable for protein sequencing. However, to date, no nanopore protein sequencing approaches have demonstrated satisfactory performance suitable for proteomics investigations. Here, a nickel-immobilized Mycobacterium smegmatis porin A (MspA) nanopore (MspA-NTA-Ni) was discovered to be suitable for the identification of a wide variety of proteomic analytes, from amino acids (aa) to peptides of 39 amino acids (1-39 aa) . Under identical conditions, 20 proteinogenic amino acids, 4 amino acids containing posttranslational modifications (PTMs) , 32 peptides, 6 peptides containing PTMs, 11 bioactive peptides and 2 neoantigens were identified. Assisted by machine learning, these analytes are fully discriminated by reporting a 97.4%validation accuracy. With MspA-NTA-Ni, a peptide can be directly sensed or hydrolysed to fragments for peptide profiling. To show nanopore peptide sequence assembly, a reference peptide is treated with exo-and endopeptidases, generating amino acid and peptide fragments with overlapping sequences. These hydrolysates are sensed, and via machine learning prediction, the amino acid composite and the sequences of the fragmented peptides are identified so that the original peptide sequence can be reconstructed. Peptide sequencing by hydrolysis is sensitive to mutations, deletions and PTMs on the original peptide, suggesting its potential applications in proteomics investigations.
[0341] EXAMPLES
[0342] The examples below are intended to be purely exemplary of the invention and should therefore not be considered to limit the invention in any way. The following examples and detailed description are offered by way of illustration and not by way of limitation.
[0343] Peptidase hydrolysis of natural proteins would result in a complex mixture of amino acids (aa) and peptide fragments. The PTM information is also retained. Clear identification of all these fragmented proteomic analytes may provide information for peptide profiling (38) and sequencing (39) , as shown in a variety of mass spectrometry-based studies (40, 41) . To further push this sensing capacity to single molecule, for example, with a nanopore, an urgent need would be the establishment of a general peptide sensor that has a high resolution and wide dynamic range so that any subtle differences in the sequence, modification and conformation of peptides would be clearly reported (Figure 7a) . However, according to previous reports of nanopore peptide sensing, peptides of shorter length (4-10 aa) consistently report shorter-residing events (24) . For peptides shorter than 7 aa, the event dwell time is too short (<10 ms) for clear data acquisition and high-resolution event recognition. Nanopore sensing of ultra-short peptides (2-3 aa) has also never been previously reported. These results clearly disclose an ambiguous zone of nanopore for high resolution identification of short peptides (2-10 aa) . Though urgently needed in nanopore proteomics, a nanopore sensor that demonstrates a full sensing coverage from an amino acid to polypeptide has never been reported.
[0344] High resolution peptide identification by MspA-NTA-Ni
[0345] MspA-NTA-Ni, a hetero-octameric Mycobacterium smegmatis porin A (MspA) nanopore modified with a single nitrilotriacetic acid-nickel (NTA-Ni) adapter at the pore constriction, has enabled simultaneous identification of all 20 proteinogenic amino acids and their post translational modifications (34) via coordination interactions. Peptides also feature abundant donor atoms capable of interacting with metal ions (42) . As reported, the terminal amino-N and amide-N of peptides are suitably positioned to form a five-membered chelates (43) , suggesting that MspA-NTA-Ni may also be suitable for peptide sensing (Figure 7b) .
[0346] To testify this, a variety of glycine peptides (Table 1) were synthesized and sensed by MspA-NTA-Ni (Figure 8) . The measurements were performed with an MspA-NTA-Ni in a 1.5 M KCl buffer with 10 mM N-cyclohexyl-2-aminoethanesulfonic acid (CHES) at pH 9.0 and a voltage of +100 mV was continually applied (Methods) . By convention, the measurement chamber which is electrically grounded is defined as cis. Whereas, its opposing chamber is defined as trans. With a single MspA-NTA-Ni inserted however without any analyte addition, a steady open pore current (I0) is reported (Figure 9) . With the addition of any glycine peptide to the cis compartment, corresponding nanopore events were immediately observed (Figure 8) , confirming the peptide sensing capacity of MspA-NTA-Ni. Briefly, each glycine peptide reports a single type of event lasting a few hundreds of ms. This extended event dwell time serves to report more details of event features for event identification (Figure 10) . Results of MspA-NTA-Ni measurement of N-acetylglycylglycine (N-Ac-GG) (Figure 11) and M2 MspA (44, 45) measurement of glycine peptides (Figure 12) further confirm that the terminal amino-N and amide-N of the peptides and the NTA-Ni adaptor are necessary for peptide event generation.
[0347] Table 1. The sequence of peptides investigated in this work.
[0348] 1. (pS) represents O-phospho-L-serine.
[0349] 2. (acK) represents N-acetyl-L-lysine.
[0350] 3. The OMe represents the carbomethoxy group.
[0351] 4. γ stands for the γ-amide bond between glutamic acid and cysteine of glutathione.
[0352] 5. [disulfide bridge: 2a-2b] represents the disulfide bridge between two cysteines of oxidized glutathione.
[0353] To quantitatively describe all peptide events, core event parameters including open pore current (I0) , event current (Ipep) , event standard deviation (SD) and dwell time (toff) are defined (Figure 13) . The blockage amplitude (ΔI) is defined as ΔI=Ipep-I0. The mean and the mean are derived from Gaussian fitting results. Whereas, the mean toff (τoff) is derived from the single exponential fitting result. Based on results of three independent measurements, a high data consistency is confirmed (Table2) . By simultaneously considering ΔI and SD, events of GG, GGG, GGGG, GGGGG and GGGGGG are thoroughly discriminated (Figure 14) , demonstrating a high-resolution of MspA-NTA-Ni in the discrimination of peptides that differ with only a single amino acid.
[0354] Table 2. Core event parameters of peptide events in the database. The measurements were performed as shown in Figures 8, 18-23 in a 1.5 M KCl buffer (1.5 M KCl, 10 mM CHES, pH 9) . A voltage of + 100 mV was continually applied during the measurements. The definition and the derivation of core event parameters are described in Figure 13. Three independent measurements (N=3) were performed for each peptide to obtain the statistics.
[0355] A machine learning algorithm was developed (Figure 15 and Methods) . Specifically, events acquired with different glycine peptides were first collected to form a dataset. Event features including ΔI, SD, toff, skewness (skew) and kurtosis (kurt) were extracted from each event for model training. A ten-fold cross-validation was performed to evaluate the model performance, according to which the cubic SVM model has reported the best validation accuracy of 99.8%and is used for downstream event prediction (Figure 15) .
[0356] To show simultaneous sensing of amino acids and peptides by MspA-NTA-Ni, glycine was added to the measurement in the presence of a variety of glycine peptides (Figure 16) . According to the event scatter plot results, all analytes including glycine and glycine peptides are fully discriminated. However, a nanopore sensor that can simultaneously identify both amino acids and peptides has never been previously reported. It was also found that the event dwell time of glycine is significantly longer than that of all glycine peptides (Figure 17) . This discrepancy may be due to the different binding mechanisms between amino acids and peptides to the NTA-Ni adapter.
[0357] To show the generality and resolution of this method, a peptide pool containing 30 model peptides, including GG, GGG, GGGG, GGGGG, GGGGGG, ARNKRS, DQ, DQAR, DQARNKR, DQARNKRS, FQAR, GD, GH, GI, GL, GP, GR, RG, GGR, GYI, GYIK, GYLK, IK, LRG, LRGY, NK, NKR, NMR, RSLR and SLR (Table 1) , is established (Figure 1a) . Each peptide in the pool was separately measured by MspA-NTA-Ni to collect corresponding event characteristics (Figures 18-23 and Table 2) , according to which each peptide reports highly discriminable event features. Representative events acquired with these peptides are demonstrated in Figure 1b. Though some peptide events report similar event appearances, they are fully discriminable when their event features are quantitatively compared (Figure 24) . Most peptides report only a single type of event, consistent with the fact that the reversible pore-peptide interaction is caused only by the coordination interaction between the N terminus of the peptide and the NTA-Ni adapter. Whereas, peptides containing histidine (His) or proline (Pro) (e.g., GH and GP) always exhibit multiple event types, suggesting that the side chain groups of these peptides may generate other ways of pore-peptide interactions. This phenomenon is consistent with that observed in our previous work, according to which His or Pro also reports multiple event types when probed by MspA-NTA-Ni (34) .
[0358] To eliminate potential interferences caused by impurities introduced during solid-phase peptide synthesis (SPPS) , such as amino acid deletion and insertion, incomplete removal of protecting groups, oxidation, reduction and diastereoisomerization (46) , results in the peptide database are also cluster analysis treated by a density-based spatial clustering of applications with noise (DBSCAN) algorithm, which removes all non-clustered events. The fundamental idea of DBSCAN is to ensure that each data in a cluster has at least a neighbouring data within a specified radius (epsilon) containing a minimum number of objects (min samples) (47) , as demonstrated by results of glycine peptides (Figure 25) .
[0359] Quantitatively, the mean blockage amplitude of these peptides measures between -62 pA and 97 pA, demonstrating a wide dynamic range of 159 pA, desired for peptide discrimination. The mean standard deviation of peptide event amplitude ranges from 2 pA to 50 pA, indicating that some peptides report large current fluctuations containing rich sensing information. The mean event dwell time (τoff) measures between 30 ms and 1100 ms (Table 2) . The event scatter plot of ΔI versus SD of results acquired with all above mentioned peptides is presented in Figures 1c-d, according to which, these peptides are fully discriminated. Since the pore-peptide interaction is established between the peptide N terminus and the NTA-Ni adapter, even a dipeptide could be clearly resolved. Notably, RG and GR, which is a pair of isomers differing only in the sequence, and difficult to be discriminated solely by conventional mass spectrometry, are immediately discriminated by MspA-NTA-Ni (Figure 26) . This also suggests potential application of this technique for rapid quality control of synthetic peptides, as demonstrated by model peptides such as NK and RSLR (Figure 27) .
[0360] A relevant machine learning algorithm is established (Figure 28) . Specifically, 300 events were randomly selected from results acquired with each peptide. Results selected from all 30 peptides were combined to form a dataset consisting of a total of 9000 events. Five event features (ΔI, SD, toff, skew, and kurt) were extracted for all events to formulate a feature matrix for training and prediction. Among various models, the quadratic SVM model has reported the highest validation accuracy of 99.0% (Table 3) and a testing accuracy of 98.4% (Figure 28b) . With an MspA-NTA-Ni, a mixture of DQ, IK, SLR, NKR and LRG was added to cis with a final concentration of 0.5 mM for each peptide. The relevant event scatter plot is generated followed with cluster analysis and machine learning labelling by the previously trained quadratic SVM model (Figure 29) . Representative traces acquired at this condition are also shown in Figure 1e.
[0361] Table 3. The accuracy of different models for peptides. The workflow of machine learning was described in Figure 28. 300 events from each type of peptide were first collected to form a dataset. Each event has a known label. Five event features were extracted from each event to form a feature matrix. The feature matrix was divided into a training set (80%) for model training and a testing set (20%) for model testing. Ten-fold cross-validation was used to evaluate the model performance. The corresponding validation accuracy and testing accuracy were respectively obtained. The quadratic SVM model has reported the highest validation accuracy of 99.0%.
[0362] Simultaneous analysis of peptides and amino acids by MspA-NTA-Ni
[0363] Since the results of peptide sensing were acquired by MspA-NTA-Ni at an identical condition previously used for sensing of all 20 proteinogenic amino acid and 4 amino acids containing PTMs (Figure 30) , to meet the need of simultaneous identification of amino acids and peptides, previous data acquired with 24 amino acids were merged with the peptide database to construct an integrated database (Figure 31) . Among various models being tested, the quadratic SVM model has reported the best validation accuracy of 98.1%for the integrated database (Table 4) . The corresponding confusion matrix is shown in Figure 31b and most classes exhibit a true positive rate (TPR) more than 95% (Figure 32) . To this end, a nanopore system for simultaneous sensing of amino acid and peptide, which has never been previously reported, is established.
[0364] Table 4. The accuracy of different models for simultaneous classification of peptides and amino acids. The workflow of machine learning was shown in Figure 31. 300 events from each amino acid and peptide were first collected to form a dataset. Each event has a known label. Five event features were extracted for each event to form a feature matrix. The feature matrix was further divided into a training set (80%) and a testing set (20%) , respectively for model training and model testing. Ten-fold cross-validation accuracy was used to evaluate the model performance. The quadratic SVM model has reported the highest validation accuracy of 98.1%.
[0365] Peptide profiling by MspA-NTA-Ni
[0366] The high peptide sensing resolution provided by MspA-NTA-Ni may also be applied for single molecule peptide profiling. In principle, a peptide can be first endo-peptidase cleaved into fragments. The generated hydrolysate is further analysed by MspA-NTA-Ni, producing distinctive fingerprint spectra, serving as a unique ID for peptide identification (Figure 2a) . To testify this, four bioactive peptides, including angiotensin III (7aa) , secretin (27aa) , glucagon (29aa) and adrenocorticotropic hormone (ACTH) (39aa) (Table 1) , were applied for this demonstration.
[0367] Experimentally, all four peptides were respectively treated with trypsin under an identical condition (Methods) . Afterwards, all resulting hydrolysates were respectively ultrafiltration (10 kDa molecular cut off) treated to remove trypsin, and the filtrate was collected. 20 μL of each filtrate was added to the cis chamber during separate nanopore measurements. The hydrolysate of each peptide reports characteristic combinations of nanopore events, generating unique fingerprints for peptide identification (Figures 33-36) . For demonstrative purpose, these results are displayed in kernel density plots (Figures 2b-e) , according to which each peptide reports a distinct pattern. Though only demonstrated with trypsin here, different combinations of endopeptidase may as well be used to produce more refined signatures for peptide profiling. Though not demonstrated in this paper, when properly treated, large proteins may as well be cleaved into peptide fragments for profiling analysis. Nanopore peptide profiling can also be carried out with complex samples like different combinations of peptides, cell lysates or whey protein.
[0368] Peptide sequencing by hydrolysis performed by MspA-NTA-Ni
[0369] The above demonstrated sensing principle may be further evolved to achieve nanopore peptide sequencing (Figure 37) . To demonstrate the proof of concept, a synthetic tetradeca-peptide (DQARNKRSLRGYIK) , also referred to as the reference peptide, was custom synthesized. Its sequence is designed to be suitable for hydrolysis by common exo-and endopeptidases. The peptide to be sequenced was first thoroughly exopeptidase hydrolysed to generate amino acids (Figures 3a-b) . Leucine aminopeptidase, an exopeptidase that releases amino acids from the N-terminus of peptides, was used for this operation (Methods) . The hydrolysate was ultrafiltration (10kDa molecular cut off) treated and the filtrate was collected so that the exopeptidase was removed. During nanopore measurement, 20 μL filtrate was added to the cis chamber. Immediately afterwards, corresponding amino acid events appeared (Figure 38) . After cluster analysis, the events were identified by the previously trained machine learning algorithm (Figure 39) (34) . Eleven fully separated event populations, respectively corresponding to R, K, N, S, A, I, Y, G, Q, L and D were observed, according to which the amino acid compositions of the reference peptide is determined (Figure 3a) . The integrated database (Figure 40) was further simplified according to the amino acid compositions and only data of relevant amino acids and peptide containing the identified amino acids are still retained for further peptide prediction (Figure 3b) . After this, the simplified database includes data of 11 types of amino acids and 26 types of peptides. The previous quadratic SVM model (Figure 31) was re-trained to fit the simplified database and a validation accuracy of 99.0%was reported (Figure 40) .
[0370] To generate fragments with sequence overlaps, two endopeptidases, trypsin and thermolysin, were separately used on the reference peptide. Specifically, trypsin cleaves at the carboxylic side of arginine and lysine unless followed by proline (48) . Whereas, thermolysin hydrolyses peptide bonds at the N-terminal side of large hydrophobic residues such as valine, leucine, isoleucine and phenylalanine (49) . After endopeptidase treatment, both hydrolysates were ultrafiltration treated to remove the peptidase. In separate measurements, 20 μL hydrolysate was respectively added to the cis compartment, immediately after which, distinct peptide events were observed (Figures 3c, f, and Figures 41-42) . The corresponding results are shown in the event scatter plots of ΔI versus SD (Figures 43-44) . The DBSCAN algorithm was used to eliminate background noises. The previously trained quadratic SVM model (Figure 40) was used for event identification (Figures 3d, g and Figures 43-44) . Only prediction results with an appearance frequency >1%were further applied for sequence assembly (Figures 3e, h) . Specifically, peptide fragments generated by tryptic digestion were predicted to be SLR, NKR, GYIK and DQAR (Figure 3e) and those identified in the thermolysin hydrolysate are DQARNKRS, LRGY, LRG and IK. Notably, tyrosine (Y) , an amino acid, was also identified in the thermolysin hydrolysate (Figure 3h) , acknowledging the sensing generality offered by MspA-NTA-Ni. According to the sequence overlaps of these peptide fragments, the sequence of the reference peptide is determined to be DQARNKRSLRGYIK, consistent with its actual sequence (Figure 3i) .
[0371] The reference peptide (DQARNKRSLRGYIK) was also characterized by liquid chromatography-tandem mass spectrometry (LC-MS / MS) (Methods) . Limited backbone cleavage was detected, with fragment ions from peptide backbone cleavage either missing or appearing with very low intensity in the MS / MS spectrum (Figure 45) . This is attributed to the presence of multiple basic residues in the peptide sequence, with cleavage near these residues being restricted due to the sequestering of the protons at these sites (39, 50) . Similarly, incompletely hydrolysed peptides demonstrate similar characteristics to the reference peptide in MS / MS spectra, generating insufficient information for sequence analysis. However, the above demonstrated nanopore sequencing by hydrolysis strategy can eliminate the need for a collision-induced dissociation step and prevent issues arising from inadequate fragmentation.
[0372] Identification of mutations and PTMs
[0373] Protein mutations, which are closely associated with protein functions and relevant diseases, are widely observed in proteomics (51) . In principle, this nanopore technique has a resolution to detect mutations down to a single amino acid (Figure 4a) . To testify this, four peptides containing a single amino acid alteration in contrast to the reference peptide (DQARNKRSLRGYIK) , including D1F peptide (FQARNKRSLRGYIK) , K6M peptide (DQARNMRSLRGYIK) , I13L peptide (DQARNKRSLRGYLK) and K14deletion peptide (DQARNKRSLRGYI) , were synthesized. Trypsin was used for peptide hydrolysation (Methods) .
[0374] The hydrolysates acquired with above peptides were separately measured by MspA-NTA-Ni, each reporting a set of peptide results (Figures 46-49) . After noise reduction, the previously trained quadratic SVM model (Figure 28) was employed for event identification (Figure 50) and relevant results were demonstrated in Figure 4d, f, h, j. In the scatter plot, any missing of a certain peptide data cluster and the generation of a new peptide data cluster suggests the happening of mutation in comparison with the reference peptide (Figure 4b) . Assisted with the database, in corresponding event scatter plots, new peptide data clusters such as FQAR of D1F, NMR of K6M, GYLK of I13L peptide and GYI for K14deletion peptide, are clearly identified (Figure 4c, e, g, i, k) . Notably, the I13L peptide, with the isoleucine residue at position 13 substituted with its leucine isomer, could be unambiguously differentiated from the reference peptide (Figure 4b, h) . Notably, discrimination of Ile and Leu of peptide is still a challenge for mass spectrometry (MS) which requires multi-stage mass analysis and complicated data interpretation (52) .
[0375] Post-translational modifications (PTMs) of proteins play a crucial role in the modulation of their active state, localization, turnover and interactions with other proteins. Identification and localization of PTMs is thus essential for a comprehensive understanding of the function regulation of a proteome (53) . To demonstrate identification of PTMs by nanopore peptide sequencing, two peptides, respectively referred to as pS8 (DQARNKR (pS) LRGYIK) and acK14 (DQARNKRSLRGYI (acK) ) , were synthesized. They respectively demonstrate peptides containing phosphorylation or acetylation modifications. Their sequences differ from the reference peptide (DQARNKRSLRGYIK) only at the PTM site. Peptide fragments containing post-translational modifications, such as phosphorylated peptides (pS) LR, (pS) LRGYIK and NKR (pS) LR and acetylated peptides GYI (acK) , SLRGYI (acK) and SLRGYI (acK) , were also synthesized and measured by MspA-NTA-Ni (Figures 51-52 and Table 5) . Their corresponding nanopore signatures were also added to the peptide database for future reference (Figures 5a-b, Figures 53-54) .
[0376] Table 5. Core event parameters of events of peptides containing PTMs. The measurements were performed using MspA-NTA-Ni in a 1.5 M KCl buffer (1.5 M KCl, 10 mM CHES, pH 9) (Figures 51-52) . A voltage of + 100 mV was continually applied during the measurements. Three independent measurements (N=3) were performed for each condition to obtain the statistics.
[0377] Trypsin hydrolysis of pS8 and acK14 were respectively performed and their hydrolysates were separately measured by MspA-NTA-Ni, each reporting a set of distinct nanopore profile (Figures 55-56) . By using the updated peptide database (Figures 53-54) , all peptide fragments were clearly identified by machine learning (Figure 57) , according to which, the peptide fragments of pS8 were identified to be DQAR, NKR, (pS) LR, (pS) LRGYIK and GYIK (Figures 5c-e) . A partially cleaved fragment, (pS) LRGYIK, was also detected in the hydrolysate of pS8. It is thus speculated that the presence of a phosphate group may hinder the access of trypsin to the cleavage site (54) . For acK14, nanopore identification of its tryptic hydrolysate reports four peptide fragments including DQAR, NKR, SLR and GYI (acK) (Figure 5f-h) . By comparison with results of the reference strand (Figure 4b) , the PTM sites of pS8 and acK14 are clearly identified.
[0378] Direct Identification of bioactive peptides
[0379] Bioactive peptides are a category of peptides that exhibit positive impact on body functions and health conditions of human (55) . To show direct identification of bioactive peptides by MspA-NTA-Ni, aspartame, cyclo (His-Pro) , Arg-Gly-Asp (RGD) , reduced glutathione (GSH) , oxidized glutathione (GSSG) , Leu-enkephalin, angiotensin III, somatostatin-14, secretin, glucagon and adrenocorticotropic hormone (ACTH) were respectively measured by MspA-NTA-Ni (Figures 58-67, Table 1) . Here, aspartame is an artificial dipeptide sweetener widely used in food and beverages (56) . RGD peptide is the minimum unit of cell-adhesive domain in various adhesion proteins (57) . Leu-enkephalin (YGGFL) is a type of enkephalin, that plays a crucial role in modulating diverse neural functions and is closely linked to neurodegenerative conditions (58) . Cyclo (His-Pro) is an endogenous cyclic dipeptide derived from the cleavage of the hypothalamic thyrotropin releasing hormone (59) . Angiotensin III is a key active peptide in the renin-angiotensin system, contributing significantly to blood pressure regulation, cardiovascular functions, and fluid balance within the body (60) . Glutathione (GSH) is a ubiquitous tripeptide that plays a critical role in cellular resistance to oxidative stress. Oxidation of GSH can lead to the formation of GSSG, a disulfide-bonded dimer of two GSH molecules. The GSH / GSSG ratio serves as an essential marker of cellular toxicity and various human diseases (61) . Somatostatin-14, a cyclo-tetradeca-peptide containing a disulfide bond between the 3rd and the 14th cysteine residues, is a peptide hormone that have broad inhibitory effects on hormone secretion including growth hormone, insulin, and glucagon (62) . Secretin is a 27-amino acid peptide, impacting the function of various organ systems including heart, kidney, lung and brain (63) . Glucagon is a single-chain polypeptide containing twenty-nine amino acids, playing a vital role in the regulation of glucose homeostasis (64) . ACTH is a 39-amino acid peptide, regulating the diurnal secretion of glucocorticoids and the acute release of glucocorticoids during stress response (65) .
[0380] Briefly, aspartame, RGD, GSH, GSSG, Leu-enkephalin, somatostatin-14, and ACTH exhibited a single event type. However, cyclo (His-Pro) , angiotensin III, secretin and glucagon each report multiple event types. Based on results of three independent measurements, a high data consistency is confirmed (Table 6) . The corresponding machine learning model was also trained and an accuracy of 97.8%is achieved by the bagged trees model (Figure 68a-b) . To demonstrate simultaneous sensing of bioactive peptides, during single channel recording using MspA-NTA-Ni, aspartame, RGD, GSH, GSSG, cyclo (His-Pro) , leu-enkephalin and angiotensin III were simultaneously added to cis. Representative traces and relevant event scatter plot are shown in Figures 68c-g. All events were identified by the trained bagged trees model. Two isomeric peptides of RGD peptide, GDR and DRG, were also synthesized and measured by MspA-NTA-Ni (Figure 69) . Thorough discrimination of all three peptides are clearly demonstrated, further confirming the uniqueness of the corresponding peptide event signature and the resolution of MspA-NTA-Ni.
[0381] Table 6. Core event parameters of events acquired with bioactive peptide. The measurements were performed using MspA-NTA-Ni in a 1.5 M KCl buffer (1.5 M KCl, 10 mM CHES, pH 9) (Figures 58-67) . A voltage of + 100 mV was continually applied during the measurements. Three independent measurements (N=3) were performed for each condition to obtain the statistics.
[0382] Results of all above described peptides were also added to the peptide database, generating a comprehensive database containing a total of seventy-three types of analytes, including the initial thirty peptides (Figure 1b) , six peptides containing PTMs (Figures 5a-b) , twenty amino acids (Figure 30) , four amino acids containing PTMs (Figure 30) , eleven bioactive peptides (Figures 58-67) and two RGD isomeric peptides (Figure 69) . This database was then fed to machine learning for training (Figure 6 and Figure 70) . According to ten-fold cross-validation results, the quadratic SVM model has reported the highest validation accuracy of 97.4% (Table 7) . To evaluate the potential discriminatory power of this nanopore peptide sensor, an accuracy versus database size curve was generated (Figure 71) , according to which the model accuracy is consistently high. By extrapolating the curve, it is anticipated that a high model accuracy could be maintained when more peptide data is to be further added to the database.
[0383] Table 7. The accuracy of different models for classification of all peptides and amino acids investigated in this work. A total of 73 types of analytes, including twenty-four amino acids and forty-nine peptides were collected to form a database (Figure 6) . 300 events acquired with each analyte were collected to form a dataset. Each event has a known label. Five event features were extracted from each event to form a feature matrix. The feature matrix was further divided into a training set (80%) and a testing set (20%) , respectively for model training and testing. Ten-fold cross-validation was used to evaluate the model performance. The quadratic SVM model has reported the highest validation accuracy of 97.4%.
[0384] Nanopore identification of neoantigens
[0385] Neoantigens, which are peptides arising from tumor-specific somatic mutations, could be recognized by the host immune system (71, 72) . Given their exclusive expression in tumor cells, neoantigens serve as a promising target for precision immunotherapies (73) . To further demonstrate the resolution of MspA-NTA-Ni in the identification of neoantigens, two widely-recognized public neoantigens, BRAFV600E (sequence: LATEKSRWSG) (74) and KRASG12D (sequence: VVVGADGVGK) (75) were selected as models to study with. Their corresponding wild-type peptides, BRAFWT (sequence: LATVKSRWSG) and KRASWT (sequence: VVVGAGGVGK) were also synthesized, measured and compared. Experimentally, with an MspA-NTA-Ni inserted, four peptides were separately added to cis to initiate the measurements. Each peptide reports a unique and highly-consistent event feature (Figure 72) , sufficient for event discrimination. In two separate assays, both the neoantigens and their unmutated peptides were also simultaneously detected by the same nanopore, further confirming event discrimination (Figure 72e, g) , as also seen in corresponding event scatter plots (Figure 72f, h) . These findings suggest that neoantigens can be effectively detected and readily distinguished by resolution-enhanced nanopore.
[0386] Nanopore peptide sequencing assisted by Edman degradation
[0387] A proof-of-concept demonstration of peptide sequencing through the integration of Edman degradation with nanopore peptide sensing was demonstrated (Figure 73) . GGGGGG served as a model peptide, undergoing sequential N-terminal modification with phenylisothiocyanate (PITC) followed by acidic cleavage, thereby generating progressively truncated peptides (GGGGGG →GGGGG → GGGG → GGG → GG) (Figure 73a) . Distinctive event signatures were observed for each peptide, enabling clear discrimination of all intermediates (Figure 73b-c) . This approach establishes a framework for N-to-C terminal peptide sequencing through synchronized chemical degradation and nanopore detection, underscoring its potential for de novo sequencing applications.
[0388] Conclusion
[0389] We report the use of MspA-NTA-Ni for simultaneous sensing of amino acid and peptide. Though the sensing mechanisms for amino acid and peptide differ, they both require an immobilized nickel at the pore constriction for data generation, thus enabling simultaneous sensing of a total of 75 analytes, including 20 proteinogenic amino acids, 4 amino acids containing PTMs, 32 peptides, 6 peptides containing PTMs, 11 bioactive peptides and 2 neoantigens, confirming the significance of the strategy of transient immobilization of peptide N-terminus. By simultaneously considering multiple event features using machine learning, an overall 97.4%accuracy is reported. No labelling of amino acid or peptide is however needed and the sensing principle is suitable for both linear and cyclic peptides (Figures 59, 64) . A nanopore that has a wide dynamic range and high discrimination resolution for amino acids and peptides (1-39 aa) , which has not been previously reported, inspires a variety of uses in proteomics investigations. A peptide can be directly identified or hydrolysed to fragments for profiling. Direct peptide sensing is useful for targeted sensing of specific peptide types, even in the presence of interfering background analytes. It is also useful for the quality control of synthetic peptide or synthetic peptides including unnatural amino acids (UAA) . Peptide profiling, which is database independent, is useful for the characterization of complex peptide mixture or hydrolysate of proteins. As a proof of concept, peptide sequencing by hydrolysis is demonstrated, enabling accurate spotting of mutations, deletions and PTMs on the reference peptide with a single amino acid resolution. All three modes of nanopore peptide analysis are immediately useful in proteomic analysis. Though sequencing by hydrolysis is database dependent, results of model accuracy against database size, as generated by corresponding machine learning model, suggest a high capacity of MspA-NTA-Ni for simultaneous identification of many more types of peptides.
[0390] Materials
[0391] 1,2-diphytanoyl-sn-glycero-3-phosphocholine (DPhPC) was from Avanti Polar Lipids (USA) . Ethylene diamine tetraacetic acid (EDTA) , ammonium bicarbonate (NH4HCO3) , hexaglycine (GGGGGG) , trypsin (from bovine pancreas) , thermolysin (from Geobacillus stearothermophilus) , hexadecane, pentane and Genapol X-80 were from Sigma-Aldrich (USA) . E. coli strain BL21 (DE3) plysS was from Sangon Biotech (China) . Chelex 100 Resin was purchased from Bio-Rad (USA) . Potassium chloride (KCl) , potassium hydroxide (KOH) , calcium chloride (CaCl2) , sodium chloride (NaCl) , magnesium chloride (MgCl2) , nickel sulfate (NiSO4) , sodium dihydrogen phosphate (NaH2PO4) , 2- (N-cyclohexylamino) ethane sulfonic acid (CHES) , tyrosine disodium salt, aspartame and angiotensin III were from Aladdin (China) . Hydrochloric acid (HCl) and dibasic sodium phosphate (Na2HPO4) were from HUSHI (China) . Maleimido-C3-NTA was supplied by Dojindo Molecular Technologies (Japan) . Glycine (G) , acetylglycylglycine (N-Ac-GG) , Gly-His (GH) , Gly-Ile (GI) and secretin were from Bide Pharmatech Co., Ltd. (China) . Leu-enkephalin, pentaglycine (GGGGG) , Arg-Gly-Asp (RGD) , Gly-Leu (GL) , Gly-Pro (GP) and somatostatin-14 were from Macklin (China) . Cyclo (His-Pro) and oxidized glutathione (GSSG) were purchased from Shanghai Acmec Biochemical Co., Ltd (China) . Leucine aminopeptidase (LAP) , reduced glutathione (GSH) , Gly-Asp (GD) , glycylglycine (GG) , triglycine (GGG) , glucagon and adrenocorticotropic hormone (ACTH) were from Shanghai Yuanye Bio-technology (China) . Tetraglycine (GGGG) was from MERYER (China) .
[0392] Ile-Lys (IK) , Asp-Gln (DQ) , Gly-Arg (GR) , Ser-Leu-Arg (SLR) , Asn-Lys-Arg (NKR) , Asn-Met-Arg (NMR) , Leu-Arg-Gly (LRG) , Gly-Tyr-Ile (GYI) , Gly-Asp-Arg (GDR) , Asp-Arg-Gly (DRG) , Leu-Arg-Gly-Tyr (LRGY) , Gly-Tyr-Ile-Lys (GYIK) , Gly-Tyr-Leu-Lys (GYLK) , Asp-Gln-Ala-Arg (DQAR) , Asp-Gln-Ala-Arg-Asn-Lys-Arg (DQARNKR) , Asp-Gln-Ala-Arg-Asn-Lys-Arg-Ser (DQARNKRS) , Asp-Gln-Ala-Arg-Asn-Lys-Arg-Ser-Leu-Arg-Gly-Tyr-Ile-Lys (DQARNKRSLRGYIK) , Phe-Gln-Ala-Arg-Asn-Lys-Arg-Ser-Leu-Arg-Gly-Tyr-Ile-Lys (FQARNKRSLRGYIK) , Asp-Gln-Ala-Arg-Asn-Met-Arg-Ser-Leu-Arg-Gly-Tyr-Ile-Lys (DQARNMRSLRGYIK) , Asp-Gln-Ala-Arg-Asn-Lys-Arg-Ser-Leu-Arg-Gly-Tyr-Leu-Lys (DQARNKRSLRGYLK) , Asp-Gln-Ala-Arg-Asn-Lys-Arg-Ser-Leu-Arg-Gly-Tyr-Ile (DQARNKRSLRGYI) , Asp-Gln-Ala-Arg-Asn-Lys-Arg- (pSer) -Leu-Arg-Gly-Tyr-Ile-Lys (DQARNKR (pS) LRGYIK) and Asp-Gln-Ala-Arg-Asn-Lys-Arg-Ser-Leu-Arg-Gly-Tyr-Ile- (ac-Lys) (DQARNKRSLRGYI (acK) ) were synthesized by DGpeptides Co., Ltd. (China) . Asn-Lys (NK) , Arg-Gly (RG) , Gly-Gly-Arg (GGR) , Phe-Gln-Ala-Arg (FQAR) , Arg-Ser-Leu-Arg (RSLR) , Ala-Arg-Asn-Lys-Arg-Ser (ARNKRS) , (pSer) -Leu-Arg ( (pS) LR) , Asn-Lys-Arg- (pSer) -Leu-Arg (NKR (pS) LR) , (pSer) -Leu-Arg-Gly-Tyr-Ile-Lys ( (pS) LRGYIK) , Gly-Tyr-Ile- (ac-Lys) (GYI (acK) ) , Ser-Leu-Arg-Gly-Tyr-Ile- (ac-Lys) (SLRGYI (acK) ) and Asn-Lys-Arg-Ser-Leu-Arg-Gly-Tyr-Ile- (ac-Lys) (NKRSLRGYI (acK) ) were custom synthesized by Jietai Biotechnology Co., Ltd. (China) .
[0393] Methods
[0394] 1. Nanopore preparations.
[0395] (N90C) 1 (M2) 7 is a previously reported hetero-octameric MspA nanopore (66) , comprising of one N90C MspA-H6 (N90C) monomer and seven M2 MspA-D16H6 (M2) monomers. The N90C monomer contains a sole cysteine mutation at site 90, which is designed for later site-specific pore modification. The preparation of (N90C) 1 (M2) 7 is described below. Briefly, the genes of N90C and M2 were simultaneously inserted into a co-expression vector pETDuet-1 by GenScript (USA) . The plasmid was expressed in E. coli BL21 (DE3) pLysS competent cells. Subsequent purification was carried out by nickel affinity chromatography. The isolation of (N90C) 1 (M2) 7 from other hetero-octameric pore assemblies was performed by gel electrophoresis. The protein band corresponding to the desired (N90C) 1 (M2) 7 assembly was cut from the gel and the protein in the gel was recovered. The recovered protein is either directly used in nanopore measurements or stored at -80℃ for long-term storage.
[0396] To introduce a single nickel-nitrilotriacetic acid adapter (Ni-NTA) to the nanopore, the prepared (N90C) 1 (M2) 7 was thoroughly mixed with 20 mM maleimide-C3-NTA and 100 mM NiSO4 in a volume ratio of 1: 8: 1 (67) . The mixture was incubated for 6 h at 37℃. The resulting product, MspA-NTA-Ni, was immediately used in nanopore measurements or stored at -80℃ for long-term storage.
[0397] M2 MspA (68, 69) contains no cysteine and can’ t be modified by maleimide-C3-NTA. Thus, the homo-octameric M2 MspA is applied as a reference nanopore in this study. The gene coding for M2 MspA was introduced to the pET-30a (+) vector by GenScript (USA) (70) . The plasmid was expressed in E. coli BL21 (DE3) pLysS competent cells and purified by nickel affinity chromatography. The resulting product was either immediately used in nanopore measurements or stored at -80℃ for long-term storage.
[0398] 2. Electrolyte buffer preparations.
[0399] Unless otherwise stated, the KCl buffer (1.5 M KCl, 10 mM N-cyclohexyl-2-aminoethanesulfonic acid (CHES) , pH 9) was used as the electrolyte buffer for all nanopore measurements. The KCl buffer was prepared with Mill-Q water and treated with Chelex 100 resin to minimize interferences caused by other polyvalent metal ions. Subsequently, the buffer was filtered by a 0.2 μm membrane to remove the resin. Finally, the pH of the buffer was adjusted to 9, prior to use.
[0400] 3. Nanopore measurements.
[0401] The nanopore device consists of two chambers separated by a Teflon film with an orifice measuring 100 μm in the diameter. Prior to each use, the orifice was treated with 0.5% (v / v) hexadecane in pentane. Both chambers were then filled with 500 μl of electrolyte buffer. Generally, the electrically grounded chamber is defined as cis and the opposing chamber is defined as trans. A pair of Ag / AgCl electrodes, which are electrically connected to a patch clamp amplifier, were placed in both chambers, in contact with the buffer. Subsequently, a drop of DPhPC (5 mg / ml in pentane) was introduced to both chambers to assist the formation of a lipid membrane on the orifice. To initiate pore insertions, nanopores were then added to cis. Upon a single nanopore insertion, the buffer in cis was immediately exchanged with fresh buffer to prevent further pore insertions. To initiate nanopore measurements, the analytes were added to cis at desired concentrations.
[0402] To minimize interferences caused by environment noises, all nanopore measurements were performed in a Faraday cage mounted on a floating optical table. All electrophysiological measurements were carried out with an Axonpatch 200B patch-clamp amplifier paired with a Digidata 1550B digitizer. All single-channel recordings were sampled at 25 kHz and low-pass filtered with a corner frequency of 1 kHz. Unless otherwise specified, all measurements were performed in a 1.5 M KCl buffer (1.5 M KCl, 10 mM CHES, pH 9) at room temperature. A transmembrane voltage of +100 mV was continually applied during the measurements.
[0403] 4. Data analysis.
[0404] All nanopore events were detected using the “Single Channel Search” function of Clamfit 10.7 (Molecular Devices) . To minimize interferences of randomly appearing spiky events, all events with a dwell time less than 10 ms were ignored. Five event features, including ΔI, SD, toff, skewness (skew) and kurtosis (kurt) , were extracted from all events using a custom MATLAB code. Further analysis and plotting were performed by Origin 2023b (Origin Lab) .
[0405] Density-Based Spatial Clustering of Applications with Noise (DBSCAN) analysis was utilized for cluster analysis. The DBSCAN analysis was performed by Python. The epsilon and the min samples were set to appropriate values for different analyte class to capture the most clustered event populations in the event scatter plot of ΔI versus SD.
[0406] Machine learning was performed using the “Classification Learner” toolbox of MATLAB. A labeled dataset was created by collecting 300 events of each analyte type. To maximize the generality of the data, these 300 events were acquired from three independent measurements. Five event features (ΔI, SD, toff, skew, kurt) were extracted for all events to form a feature matrix. The feature matrix was then split to a training set (80%) and a testing set (20%) , respectively for model training and model testing. Seven classifiers, including Support Vector Machine (SVM) , Decision Trees, Ensemble, Bayes, Discriminant Analysis, K-Nearest Neighbor (KNN) and Neural Network were respectively trained by studying the event features of the training set. To minimize overfitting, ten-fold cross-validation was performed to assess the model performance, and the corresponding validation accuracy was evaluated. The testing set was used to test the model performance, and the corresponding testing accuracy was obtained. The best-performing model was screened based on the cross-validation accuracy. The best-performing model was then selected to predict unlabeled data. The corresponding confusion matrix result was generated by the best-performing model to display the classification details for each analyte class.
[0407] One-Class SVM is a supervised machine learning algorithm used for outlier analysis. Here, One-Class SVM was used to identify events that belong to previously trained event types from mixtures of analytes. The One-Class SVM algorithm in MATLAB was used. 300 events acquired with the corresponding analyte class were collected, from which five features from each event were extracted to build the feature matrix for training. The trained model was applied in the judgement of inlier or outlier events.
[0408] 5. Peptide digestion with different peptidases.
[0409] In separate measurements, leucine aminopeptidase (LAP) , trypsin and thermolysin were employed for peptide digestion. LAP was prepared in a 50 mM phosphate buffer (pH 8.0) , with a stock concentration of 48 mg / ml (>7 units / mg, Yuanye) . Trypsin was dissolved in a solution of 1 M HCl and 5 mM CaCl2 at a concentration of 10 mg / ml (~10,000 BAEE units / mg, Sigma-Aldrich) . Thermolysin was dissolved in a solution of 0.1 M NaCl and 5 mM CaCl2 at a concentration of 10 mg / ml (30-350 units / mg, Sigma-Aldrich) .
[0410] For LAP hydrolysis, the peptide was dissolved in a 50 mM phosphate, pH 8 buffer with a concentration of 10 mM. The hydrolysis reaction was initiated by mixing 5 μl LAP solution, 100 μl peptide solution and 5 μl MgCl2 solution, followed by incubation at 37℃ for 8 h in a block incubator. The reaction was terminated by heating the mixture at 90℃ for 5 min. Afterwards, the mixture was transferred to an ultracentrifuge tube with a 10 kDa molecular weight cut-off. Ultrafiltration was performed at 8000 rpm for 60 min at 4℃ to remove the enzyme from the digestion product. The filtrate was collected and stored at 4℃ for further nanopore measurements.
[0411] For trypsin and thermolysin hydrolysis, all substrate peptides were dissolved in a 50 mM NH4HCO3, 5 mM CaCl2 solution at a concentration of 2.5 mM (secretin, glucagon and ACTH) or 10 mM (the rest peptides) . To initiate the hydrolysis reaction, the peptide solution and the enzyme solution were mixed with an enzyme / peptide ratio of 1: 50 (w / w) and incubated at 37℃ for 8 h in a block incubator. After the hydrolysis reaction, trypsin was heat-deactivated at 90℃ for 5 min. In separate experiments, thermolysin was deactivated by the addition of 5 mM EDTA and heat treated at 40℃ for 5 min. Afterwards, the mixture was transformed to an ultracentrifuge tube with a 10 kDa molecular weight cut-off. Ultrafiltration was performed at 8000 rpm for 60 min at 4℃ to remove the enzyme. The filtrate was collected and stored at 4℃ for further nanopore measurements.
[0412] 6. LC-MS / MS analysis.
[0413] The LC-MS / MS analysis of peptide was performed by Biotech-Pack Scientific Co., Ltd (China) . The measurements were conducted using an Easy-nLC 1200 system coupled to an Orbitrap Exploris 480 mass spectrometer. Buffer A (0.1%formic acid in water) and buffer B (20%0.1%formic acid in water-80%acetonitrile) were employed as the mobile phases. The peptide was eluted in 60 minutes with a segmented linear gradient (from 4%to 35%buffer B for 54 min, from 35%to 99 %buffer B for 3 min and from 99 %to 99%buffer B for 3 min) at a flow rate of 600 nL / min. The ionization was carried out by electrospray ionization (ESI) at 270℃. High energy collision dissociation (HCD) was employed for MS / MS fragmentation and the collision energy was set as 27 V.
[0414] References
[0415] 1. J.W. Harper, E.J. Bennett, Proteome complexity and the forces that drive proteome imbalance. Nature 537, 328-338 (2016) .
[0416] 2. R. Aebersold, M. Mann, Mass-spectrometric exploration of proteome structure and function. Nature 537, 347-355 (2016) .
[0417] 3. I. Bludau, R. Aebersold, Proteomic and interactomic insights into the molecular basis of cell functional diversity. Nat Rev Mol Cell Biol 21, 327-340 (2020) .
[0418] 4. Y. Jiang et al., Proteomics identifies new therapeutic targets of early-stage hepatocellular carcinoma. Nature 567, 257-261 (2019) .
[0419] 5. L. Restrepo-Perez, C. Joo, C. Dekker, Paving the way to single-molecule protein sequencing. Nat Nanotechnol 13, 786-796 (2018) .
[0420] 6. P. Edman, E. L. G. Sillén, P.-O. Kinell, Method for determination of the amino acid sequence in peptides. Acta chem. scand 4, 283-293 (1950) .
[0421] 7. B.F. Cravatt, G.M. Simon, J.R. Yates, 3rd, The biological impact of mass-spectrometry-based proteomics. Nature 450, 991-1000 (2007) .
[0422] 8. A. Michalski, J. Cox, M. Mann, More than 100,000 detectable peptide species elute in single shotgun proteomics runs but the majority is inaccessible to data-dependent LC-MS / MS. J Proteome Res 10, 1785-1793 (2011) .
[0423] 9. L.D. Fricker, Limitations of Mass Spectrometry-Based Peptidomic Approaches. J Am Soc Mass Spectrom 26, 1981-1991 (2015) .
[0424] 10. T. Slechtova, M. Gilar, K. Kalikova, E. Tesarova, Insight into Trypsin Miscleavage: Comparison of Kinetic Constants of Problematic Peptide Sequences. Anal Chem 87, 7636-7643 (2015) .
[0425] 11. P. Giansanti, L. Tsiatsiani, T.Y. Low, A.J. Heck, Six alternative proteases for mass spectrometry-based proteomics beyond trypsin. Nat Protoc 11, 993-1006 (2016) .
[0426] 12. M.S. Kim, J. Zhong, A. Pandey, Common errors in mass spectrometry-based analysis of post-translational modifications. Proteomics 16, 700-714 (2016) .
[0427] 13. B.D. Reed et al., Real-time dynamic single-molecule protein sequencing on an integrated semiconductor device. Science 378, 186-192 (2022) .
[0428] 14. J. Swaminathan et al., Highly parallel single-molecule identification of proteins in zeptomole-scale mixtures. Nat Biotechnol 36, 1076-1082 (2018) .
[0429] 15. T. Ohshiro et al., Detection of post-translational modifications in single peptides using electron tunnelling currents. Nat Nanotechnol 9, 835-840 (2014) .
[0430] 16. Y. Zhao et al., Single-molecule spectroscopy of amino acids and peptides by recognition tunnelling. Nat Nanotechnol 9, 466-473 (2014) .
[0431] 17. M. Filius et al., Full-length single-molecule protein fingerprinting. Nat Nanotechnol 19, 652-659 (2024) .
[0432] 18. P. Shrestha et al., Single-molecule mechanical fingerprinting with DNA nanoswitch calipers. Nat Nanotechnol 16, 1362-1370 (2021) .
[0433] 19. H. Brinkerhoff, A. S. W. Kang, J. Liu, A. Aksimentiev, C. Dekker, Multiple rereads of single proteins at single–amino acid resolution using nanopores. Science 374, 1509-1513 (2021) .
[0434] 20. S. Yan et al., Single Molecule Ratcheting Motion of Peptides in a Mycobacterium smegmatis Porin A (MspA) Nanopore. Nano Lett 21, 6703-6710 (2021) .
[0435] 21. G.M. Cherf et al., Automated forward and reverse ratcheting of DNA in a nanopore at 5-A precision. Nat Biotechnol 30, 344-348 (2012) .
[0436] 22. E.A. Manrao et al., Reading DNA at single-nucleotide resolution with a mutant MspA nanopore and phi29 DNA polymerase. Nat Biotechnol 30, 349-353 (2012) .
[0437] 23. J.A. Alfaro et al., The emerging landscape of single-molecule protein sequencing technologies. Nat Methods 18, 604-617 (2021) .
[0438] 24. G. Huang, A. Voet, G. Maglia, FraC nanopores with adjustable diameter identify the mass of opposite-charge peptides with 44 dalton resolution. Nat Commun 10, 835 (2019) .
[0439] 25. L. Restrepo-Perez et al., Resolving Chemical Modifications to a Single Amino Acid within a Peptide Using a Biological Nanopore. ACS Nano 13, 13668-13676 (2019) .
[0440] 26. J. Jiang et al., Protein nanopore reveals the renin-angiotensin system crosstalk with single-amino-acid resolution. Nat Chem 15, 578-586 (2023) .
[0441] 27. F. Piguet et al., Identification of single amino acid differences in uniformly charged homopolymeric peptides with aerolysin nanopore. Nat Commun 9, 966 (2018) .
[0442] 28. R.C. Abraham Versloot et al., Seeing the Invisibles: Detection of Peptide Enantiomers, Diastereomers, and Isobaric Ring Formation in Lanthipeptides Using Nanopores. J Am Chem Soc 145, 18355-18365 (2023) .
[0443] 29. J. Nivala, D.B. Marks, M. Akeson, Unfoldase-mediated protein translocation through an alpha-hemolysin nanopore. Nat Biotechnol 31, 247-250 (2013) .
[0444] 30. P. Martin-Baniandres et al., Enzyme-less nanopore detection of post-translational modifications within long polypeptides. Nat Nanotechnol 18, 1335-1340 (2023) .
[0445] 31. L. Yu et al., Unidirectional single-file transport of full-length proteins through a nanopore. Nat Biotechnol 41, 1130-1139 (2023) .
[0446] 32. A. Sauciuc, B. Morozzo Della Rocca, M.J. Tadema, M. Chinappi, G. Maglia, Translocation of linearized full-length proteins through an engineered nanopore under opposing electrophoretic force. Nat Biotechnol, (2023) .
[0447] 33. S. Zhang et al., Bottom-up fabrication of a proteasome-nanopore that unravels and processes single proteins. Nat Chem 13, 1192-1199 (2021) .
[0448] 34. K. Wang et al., Unambiguous discrimination of all 20 proteinogenic amino acids and their modifications by nanopore. Nature Methods 21, 92-101 (2024) .
[0449] 35. M. Zhang et al., Real-time detection of 20 amino acids and discrimination of pathologically relevant peptides with functionalized nanopore. Nat Methods 21, 609-618 (2024) .
[0450] 36. H. Ouldali et al., Electrical recognition of the twenty proteinogenic amino acids using an aerolysin nanopore. Nat Biotechnol 38, 176-181 (2020) .
[0451] 37. Y. Zhang et al., Peptide sequencing based on host-guest interaction-assisted nanopore sensing. Nat Methods 21, 102-109 (2024) .
[0452] 38. W.J. Henzel, C. Watanabe, J. T. Stults, Protein identification: the origins of peptide mass fingerprinting. J Am Soc Mass Spectrom 14, 931-942 (2003) .
[0453] 39. J. Seidler, N. Zinn, M.E. Boehm, W.D. Lehmann, De novo sequencing of peptides by MS / MS. Proteomics 10, 634-649 (2010) .
[0454] 40. P. Sinitcyn et al., Global detection of human variants and isoforms by deep proteome sequencing. Nat Biotechnol 41, 1776-1786 (2023) .
[0455] 41. Y. Bian et al., Robust, reproducible and quantitative analysis of thousands of proteomes by micro-flow LC-MS / MS. Nat Commun 11, 157 (2020) .
[0456] 42. T. Kowalik‐Jankowska, H. Kozlowski, E. Farkas, I. Sóvágó, Nickel ion complexes of amino acids and peptides. Nickel and Its Surprising Impact in Nature 2, 63-107 (2007) .
[0457] 43. I. Sovago, K. Osz, Metal ion selectivity of oligopeptides. Dalton Trans, 3841-3854 (2006) .
[0458] 44. T.Z. Butler, M. Pavlenok, I.M. Derrington, M. Niederweis, J.H. Gundlach, Single-molecule DNA detection with an engineered MspA protein nanopore. Proceedings of the National Academy of Sciences 105, 20647-20652 (2008) .
[0459] 45. I.M. Derrington et al., Nanopore DNA sequencing with MspA. Proc Natl Acad Sci USA 107, 16060-16065 (2010) .
[0460] 46. M. D'Hondt et al., Related impurities in peptide medicines. J Pharm Biomed Anal 101, 2-30 (2014) .
[0461] 47. K. Khan, S.U. Rehman, K. Aziz, S. Fong, S. Sarasvady, in The Fifth International Conference on the Applications of Digital Information and Web Technologies (ICADIWT 2014) . (2014) , pp. 232-238.
[0462] 48. J. Rodriguez, N. Gupta, R.D. Smith, P.A. Pevzner, Does Trypsin Cut Before Proline? Journal of Proteome Research 7, 300-305 (2008) .
[0463] 49. O.A. Adekoya, I. Sylte, The Thermolysin Family (M4) of Enzymes: Therapeutic and Biotechnological Potential. Chemical Biology &Drug Design 73, 7-16 (2008) .
[0464] 50. L.M. Mikesh et al., The utility of ETD mass spectrometry in proteomic analysis. Biochim Biophys Acta 1764, 1811-1822 (2006) .
[0465] 51. M. Gao, H. Zhou, J. Skolnick, Insights into Disease-Associated Mutations in the Human Proteome through Protein Structural Analysis. Structure 23, 1362-1369 (2015) .
[0466] 52. Y. Xiao, M. M. Vecchi, D. Wen, Distinguishing between Leucine and Isoleucine by Integrated LC-MS Analysis Using an Orbitrap Fusion Mass Spectrometer. Anal Chem 88, 10757-10766 (2016) .
[0467] 53. M. Mann, O. N. Jensen, Proteomic analysis of post-translational modifications. Nature Biotechnology 21, 255-261 (2003) .
[0468] 54. J.A. Bubis, V. Gorshkov, M.V. Gorshkov, F. Kjeldsen, PhosphoShield: Improving Trypsin Digestion of Phosphoproteins by Shielding the Negatively Charged Phosphate Moiety. J Am Soc Mass Spectrom 31, 2053-2060 (2020) .
[0469] 55. D.D. Kitts, K. Weiler, Bioactive proteins and peptides from food sources. Applications of bioprocesses used in isolation and recovery. Current pharmaceutical design 9, 1309-1323 (2003) .
[0470] 56. A.K. Choudhary, E. Pretorius, Revisiting the safety of aspartame. Nutr Rev 75, 718-730 (2017) .
[0471] 57. M. Nieberler et al., Exploring the Role of RGD-Recognizing Integrins in Cancer. Cancers (Basel) 9, 116 (2017) .
[0472] 58. W.J. Meilandt et al., Enkephalin elevations contribute to neuronal and behavioral impairments in a transgenic mouse model of Alzheimer's disease. J Neurosci 28, 5007-5017 (2008) .
[0473] 59. A. Minelli et al., Cyclo (His-Pro) exerts anti-inflammatory effects by modulating NF-kappaB and Nrf2 signalling. Int J Biochem Cell Biol 44, 525-535 (2012) .
[0474] 60. V.G. Yugandhar, M.A. Clark, Angiotensin III: a physiological relevant peptide of the renin angiotensin system. Peptides 46, 26-32 (2013) .
[0475] 61. O. Zitka et al., Redox status expressed as GSH: GSSG ratio as a marker for oxidative stress in paediatric tumour patients. Oncol Lett 4, 1247-1253 (2012) .
[0476] 62. M. Gomes-Porras, J. Cardenas-Salas, C. Alvarez-Escola, Somatostatin Analogs in Clinical Practice: a Review. Int J Mol Sci 21, 1682 (2020) .
[0477] 63. S. Afroze et al., The physiological roles of secretin and its receptor. Ann Transl Med 1, 29 (2013) .
[0478] 64. B. Ahren, Glucagon--Early breakthroughs and recent discoveries. Peptides 67, 74-81 (2015) .
[0479] 65. A. Stevens, A. White, in Cellular Peptide Hormone Synthesis and Secretory Pathways, J.F. Rehfeld, J.R. Bundgaard, Eds. (Springer Berlin Heidelberg, Berlin, Heidelberg, 2010) , pp.121-135.
[0480] 66. S. Zhang et al., A Nanopore-Based Saccharide Sensor. Angew Chem Int Ed Engl 61, e202203769 (2022) .
[0481] 67. K. Wang et al., Unambiguous discrimination of all 20 proteinogenic amino acids and their modifications by nanopore. Nature Methods 21, 92-101 (2024) .
[0482] 68. T.Z. Butler, M. Pavlenok, I.M. Derrington, M. Niederweis, J.H. Gundlach, Single-molecule DNA detection with an engineered MspA protein nanopore. Proceedings of the National Academy of Sciences 105, 20647-20652 (2008) .
[0483] 69. I.M. Derrington et al., Nanopore DNA sequencing with MspA. Proc Natl Acad Sci USA 107, 16060-16065 (2010) .
[0484] 70. Y. Wang et al., Osmosis-Driven Motion-Type Modulation of Biological Nanopores forParallel Optical Nucleic Acid Sensing. ACS Appl Mater Interfaces 10, 7788-7797 (2018) .
[0485] 71 Peng, M. et al. Neoantigen vaccine: an emerging tumor immunotherapy. Mol Cancer 18, 128 (2019) .
[0486] 72 Schumacher, T.N. &Schreiber, R.D. Neoantigens in cancer immunotherapy. Science 348, 69-74 (2015) .
[0487] 73 Pearlman, A.H. et al. Targeting public neoantigens for cancer immunotherapy. Nat Cancer 2, 487-497 (2021) .
[0488] 74 Somasundaram, R. et al. Human leukocyte antigen-A2-restricted CTL responses to mutated BRAF peptides in melanoma patients. Cancer Res 66, 3287-3293 (2006) .
[0489] 75 Wang, Q. et al. Direct Detection and Quantification of Neoantigens. Cancer Immunol Res 7, 1748-1754 (2019) .
Claims
A method of characterizing one or more target analytes in a sample, comprising:(i) providing a nanopore incorporating at least one sensing moiety which is capable of interacting with a peptide, and causing the target analyte to contact the sensing moiety in a region that allows the target analyte to be detected based on a change of an ionic current across the nanopore; and(ii) measuring the ionic current as the target analyte migrates to provide a representative current pattern of the target analyte, wherein the representative current pattern indicates a characteristic of the target analyte;wherein the one or more target analytes comprise a peptide; wherein the sensing moiety is a metal ion.The method of claim 1, wherein the sensing moiety is Ni2+, Cu2+, Co2+, Zn2+, Cd2+, Ag+, Mn2+, Pb2+, Fe2+ or Fe3+; preferably, the sensing moiety is Ni2+.The method of claim 1 or 2, wherein the sensing moiety is attached to the nanopore by a ligand; preferably, the ligand is a metal ion chelating agent; more preferably, the ligand is nitrilotriacetic acid (NTA) .The method of any one of claims 1 to 3, wherein the nanopore is a protein nanopore comprising a reactive amino acid to which the sensing moiety is attached.The method of claim 4, wherein the reactive amino acid is selected from cysteine, methionine, histidine and lysine.The method of claims 4 or 5, wherein the protein nanopore is selected from alpha-hemolysin (α-HL) , Mycobacterium smegmatis porin A (MspA) , Aerolysin (Ael) , curli production assembly / transport component (CsgG) , outer membrane porin F (OmpF) , Cytolysin A (ClyA) , ferric hydroxamate uptake component A (FhuA) , Fragaceatoxin C (FraC) , Cytotoxin K (CytK) , Pleurotolysin A (PlyA) / Pleurotolysin B (PlyB) , Curli production assembly / transport component CsgG (CsgG) , Phi29 connector protein, and any variant thereof; preferably, the protein nanopore is a variant of MspA.The method of any one of claims 4 to 6, wherein the protein nanopore is a heterogeneous protein nanopore in which one or more but not all monomers incorporate the sensing moiety and the other monomers do not incorporate the sensing moiety.The method of any one of claims 4 to 7, wherein the protein nanopore is a variant of MspA which has the reactive amino acid residue located at a position selected from 83-111, preferably 90, 91, 92 and 93.The method of claim 8, wherein the protein nanopore has a mutation of N90C, N90M, N91M or N91C on one or more monomers compared to a wild-type MspA or a variant thereof; preferably, the variant is M2 MspA.The method of any one of claims 1 to 9, wherein the peptide comprises 2-50 amino acids in length.The method of any one of claims 1-10, wherein the peptide comprises a N atom in the N-terminal amino group and the N atom in the N-terminal amide bond moiety.The method of any one of claims 1-11, wherein the peptide comprises one or more amino acids selected from histidine, lysine, cysteine, methionine and proline.The method of any one of claims 1 to 12, wherein the one or more target analytes comprises one or more peptides and / or one or more amino acids, each of which optionally and independently and have a modification; preferably, the modification is a post-translational modification; preferably, the modification is acetylation, phosphorylation, methylation, or any combination thereof.The method of any one of claims 1 to 13, wherein the characteristic comprises the identity of the target analyte, the presence or absence of the target analyte, the structure of the target analyte, the conformation of the target analyte, the sequence of the target analyte, the modification of the target analyte, the confirmation of the target analyte, the fingerprint of the target analyte, or any combination thereof.The method of any one of claims 1 to 14, wherein the sample is a target polypeptide sample the one or more target analytes are comprised in the proteinogenic components in the target polypeptide sample, and the method comprises:(i) providing a nanopore incorporating at least one sensing moiety which is capable of interacting with a peptide, and causing each proteinogenic component of the target polypeptide sample to contact the sensing moiety in a region that allows each proteinogenic component of the target polypeptide sample to be detected based on a change of an ionic current across the nanopore;(ii) measuring the ionic current as each proteinogenic component of the target polypeptide sample migrates to provide a representative current pattern of the target polypeptide sample, and obtain a fingerprint of the target polypeptide sample; and(iii) comparing the fingerprint of the target polypeptide sample with a reference fingerprint of a reference polypeptide and detecting the purity of the target polypeptide sample.The method of claim 15, wherein if the fingerprint of the target polypeptide sample comprises one or more additional results that are not contained in the fingerprint of the reference polypeptide, it means that the target polypeptide sample contains impurities.The method of claim 15 or 16, wherein the target polypeptide sample is a sample obtained by solid-phase peptide synthesis method.A method of characterizing a target polypeptide, comprising:(a) fragmenting the target polypeptide to obtain a peptide fragment mixture;(b) characterizing the peptide fragment mixture using the method of any one of claims 1-14, to determine the characteristic of the peptide fragment mixture; and(c) determining the characteristic of the target polypeptide based on the characteristic of the peptide fragment mixture.The method of claim 18, wherein the target polypeptide has or hasn’t a modification; preferably, the modification is a post-translational modification; preferably, the modification is acetylation, phosphorylation, methylation, or any combination thereof.The method of claim 18 or 19, wherein the target polypeptide is fragmented by endopeptidase digestion or Edman degradation.The method of any one of claims 18-20, wherein the step (b) comprises causing each component of the peptide fragment mixture to contact the sensing moiety in a region that allows each component of the peptide fragment mixture to be detected based on a change of an ionic current across the nanopore, measuring an ionic current as each component of the peptide fragment mixture migrates to provide a representative current pattern of the peptide fragment mixture, comparing the representative current pattern of the peptide fragment mixture to reference current patterns in a reference database, and determine the characteristic of the peptide fragment mixture based on the comparison.The method of claim 21, wherein the method further comprises an assisting-assay on the target polypeptide comprising hydrolyzing the target polypeptide into an amino acid mixture using an exopeptidase, determining the identity of each amino acid in the mixture, selecting reference current patterns of amino acids and peptides containing the identified amino acid to form the reference database.The method of any one of claims 18-22, wherein the characteristic of the target polypeptide is the identity of the target polypeptide, the presence or absence of the target polypeptide, the amino acid sequence of the target polypeptide, the amino acid alteration in the target polypeptide, the modification in the target polypeptide, the fingerprint spectrum of the target polypeptide, the purity of the target polypeptide, or any combination thereof.The method of any one of claims 18-23, wherein the step (b) includes identifying the amino acid sequence of each peptide fragment in the peptide fragment mixture; and wherein the step (c) includes assembling the amino acid sequence of each peptide fragment in the peptide fragment mixture to the amino acid sequence of the target polypeptide.The method of claim 24, wherein the target polypeptide is fragmented by multiple fragmentation operations to obtain multiple peptide fragment mixtures, and the multiple peptide fragment mixtures are separately subjected to step (b) to identify the amino acid sequences of each peptide fragment in each peptide fragment mixture, and these sequences are subsequently subjected to step (c) to be assembled to the amino acid sequence of the target polypeptide.The method of claim 25, wherein the multiple fragmentation operations are multiple digestions by the same endopeptidase or by different endopeptidases respectively.The method of any one of claims 18-23, wherein the step (b) includes identifying the fingerprint of the target polypeptide; and wherein the step (c) includes comparing the fingerprint of the target polypeptide to a fingerprint of a reference polypeptide, and determining whether the target polypeptide is identical to or different from the reference polypeptide, or whether the target polypeptide has amino acid alteration and / or modifications relative to the reference polypeptide based on the comparison.The method of claim 27, wherein the fingerprint spectra of the reference polypeptide and the target polypeptide were obtained by performing the same fragmentation operation and the same nanopore detection operation on the reference peptide and the target peptide.A method of characterizing a target polypeptide, comprising:(1) preparing a target polypeptide sample including the target polypeptide;(2) preparing different shortened polypeptide samples, wherein each shortened polypeptide sample including a shortened target polypeptide by truncating the target polypeptide, and lengths of the shortened target polypeptides in different shortened polypeptide samples are different;(3) characterizing the different shortened polypeptides using the method of any one of claims 1-28, to determine the characteristic of the shortened target polypeptides separately; and(4) determining the characteristic of the target polypeptide based on the characteristic of the different shortened polypeptides.The method of claim 29, wherein the different shortened polypeptide samples are prepared by stepwise shortening the target polypeptide.The method of claim 30, wherein the different shortened polypeptide samples are prepared by stepwise shortening the target polypeptide from N-terminus or C-terminus of the target polypeptide.The method of claim 30, wherein the different shortened polypeptide samples are prepared by Edman degradation.The method of any one of claims 29-32, wherein the characteristic of the target polypeptide is the identity of the target polypeptide, the presence or absence of the target polypeptide, the amino acid sequence of the target polypeptide, the amino acid alteration in the target polypeptide, the modification in the target polypeptide, the fingerprint spectrum of the target polypeptide, the purity of the target polypeptide, or any combination thereof.The use of the method of any one of claims 1-33 for determining characteristic of a neoantigen.The method of claim 34, wherein the characteristic of the neoantigen is the identity of the neoantigen, the presence or absence of the neoantigen, the amino acid sequence of the neoantigen, the amino acid alteration in the neoantigen, the modification in the neoantigen, the fingerprint spectrum of the neoantigen, the purity of the neoantigen, or any combination thereof.
Citation Information
Patent Citations
Protein / polypeptide sequencing method adopting Aerolysin nanopores
CN112480204A
Protein nanopore for identifying an analyte
CN112997080A
Polypeptide component analysis method based on copper ion modified MspA nanopore
CN116165373A
Polypeptide mixture detection method based on nanopore channels
CN118197454A
Electrical double layer in nanopores for detection and identification of molecules and submolecular units
US20180202969A1