De novo sequencing of target polypeptides with nanopores

WO2026112073A1PCT designated stage Publication Date: 2026-05-28ILLUMINA INC
View PDF 10 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
ILLUMINA INC
Filing Date
2025-11-18
Publication Date
2026-05-28

AI Technical Summary

Technical Problem

Current nanopore-based technologies are insufficient for sensitive and accurate sequencing of proteins due to challenges in differentiating between amino acids with similar steric volumes and the neutral nature of protein backbones, which complicates unidirectional translocation and signal interpretation.

Method used

A hybrid Edman degradation nanopore approach is employed, where a target polypeptide is immobilized, N-terminal residues are labeled with a charged tag, and cleaved using Edman degradation, allowing nanopore capture and analysis of the modified residues, utilizing a sequenceable strand with cycle markers and arresting moieties to achieve single-amino acid accuracy.

Benefits of technology

Enables de novo protein sequencing with high accuracy by transferring sequential amino acid information to a nanopore, allowing detection of post-translational modifications and protein sequences in real-time at the single-molecule level.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2025055970_28052026_PF_FP_ABST
    Figure US2025055970_28052026_PF_FP_ABST
Patent Text Reader

Abstract

Some embodiments of the methods and compositions provided herein relate to nanopore sequencing of a target polypeptide. In some embodiments, a sequencing substrate is prepared by Edman degradation from a target polypeptide immobilized on a substrate. De novo sequencing is performed with the sequence substrate and a nanopore embedded in a membrane.
Need to check novelty before this filing date? Find Prior Art

Description

ILLINC.863WO / IP-2908-PCT PATENT DE NOVO SEQUENCING OF TARGET POLYPEPTIDES WITH NANOPORESRELATED APPLICATIONS

[0001] This application claims priority to U. S. Prov. App. No. 63 / 722455 filed November 19, 2024 entitled “DE NOVO SEQUENCING OF TARGET POLYPEPTIDES WITH NANOPORES” which is incorporated by reference in its entirety.FIELD OF THE INVENTION

[0002] Some embodiments of the methods and compositions provided herein relate to nanopore sequencing of a target polypeptide. In some embodiments, a sequencing substrate is prepared by Edman degradation from a target polypeptide immobilized on a substrate. De novo sequencing is performed with the sequence substrate and a nanopore embedded in a membrane.BACKGROUND OF THE INVENTION

[0003] Nanopore-based sensing is an emerging technology with great potential for the detection of diverse organic molecules, sequencing of nucleic acids, and single-molecule analyses of enzymatic reactions and protein folding. Conceptually, nanopore biosensing belongs to the so-called resistive-pulse methods. The use of biological nanopores for resistive- pulse detection of molecular analytes was made possible by the development of planar bilayer recording and the development of single-channel current measurements. Mueller, P. el al., “Reconstitution of cell membrane structure in vitro and its transformation into an excitable system.” Nature 1962, 194, 979 -980; Hladky, S. B. etal., “Discreteness of conductance change in bimolecular lipid membranes in the presence of certain antibiotics.” Nature 1970, 225, 451 - 453. Early studies used channels generated by alamethicin, an antibiotic peptide, Staphylococcus aureus a-hemolysin, and cholera toxins. Krasilnikov, O. V.tV al., “A simple method for the determination of the pore radius of ion channels in planar lipid bilayer membranes.” FEMS Microbiol. Immunol. 1992, 5, 93-100; Korchev, Y. E.et al., “Low conductance states of a single ion channel are not ’closed’.” J. Membr. Biol. 1995, 147, 233-239. Possibly the biggest push for the development of nanopore biosensing came after it was found that single-stranded DNA and RNA could be threaded through a nanopore, making it apotential technology for nucleic acid sequencing. Kasianowicz, J. J. etal., “Characterization of individual polynucleotide molecules using a membrane channel.” Proc. Natl. Acad. Sci. USA 1996, 93, 13770-13773.SUMMARY OF THE INVENTION

[0004] Some embodiments of the methods and compositions provided herein include a method for preparing a sequencing substrate, comprising: (a) immobilizing a target polypeptide on a substrate; (b) coupling an N-terminal residue of the target polypeptide to a reactive moiety of an Edman reagent; and (c) cleaving the N-terminal residue from the target polypeptide, thereby obtaining the sequencing substrate comprising the Edman reagent and the coupled residue.

[0005] In some embodiments, the target polypeptide is immobilized on the substrate via a C -terminal residue of the target polypeptide. In some embodiments, the target polypeptide is covalently immobilized on the substrate. In some embodiments, the target polypeptide is immobilized on the substrate via a cleavable linker.

[0006] In some embodiments, the reactive moiety comprises an isothiocyanate moiety or a phenyl isothiocyanate (PITC) moiety.

[0007] In some embodiments, the Edman reagent comprises a net negative charge. In some embodiments, the Edman reagent comprises a polymer. In some embodiments, the polymer is a homopolymer. In some embodiments, the polymer comprises a subunit selected from a nucleotide, an ammo acid, a polyethylene glycol (PEG), a phosphonamidite, or a peptide nucleic acid (PNA) subunit. In some embodiments, the polymer comprises a plurality of arginine residues.

[0008] In some embodiments, the Edman reagent comprises a cycle marker indicative of a position of a coupled residue in the target polypeptide. In some embodiments, the cycle marker is selected from a non-nucleosidic reporter moiety, a peptide reporter moiety, PEG-based reporter moiety, or a single-stranded DNA reporter moiety.

[0009] In some embodiments, the Edman reagent comprises an arresting moiety adapted to pause or inhibit translocation of the sequencing substrate through a nanopore.

[0010] In some embodiments, step (c) comprises contacting the N-terminal residue with an Edmanase or with conditions sufficient to cleave the Edman reagent and the coupled residue from the target polypeptide.

[0011] In some embodiments, the Edmanase is immobilized on the substrate.

[0012] In some embodiments, the Edman reagent is immobilized on the substrate via a cleavable linker. In some embodiments, the cleavable linker comprises a cysteine and selenocysteine group modified with a protecting group.

[0013] In some embodiments, the reactive moiety is located in the Edman reagent between arresting moieties,

[0014] Some embodiments also include repeating steps (b) and (c) with an additional Edman reagent, wherein the immobilized Edman reagent and the additional Edman reagent each comprise an adapter adapted to conjugate to one another.

[0015] Some embodiments also include repeating steps (b) and (c) for a plurality of cycles, wherein each additional Edman reagent comprises a different cycle marker,

[0016] Some embodiments also include conjugating the immobilized Edman reagent to the additional Edman reagent.

[0017] In some embodiments, each adapter comprises a nucleic acid. In some embodiments, the adapters of the immobilized Edman reagent and of the additional Edman reagent are complementary to one another.

[0018] In some embodiments, the conjugating comprises a polymerase reaction or a ligase reaction.

[0019] Some embodiments also include (d) cleaving the immobilized sequencing substrate from the substrate.

[0020] In some embodiments, the Edman reagent is in solution.

[0021] In some embodiments, the Edman reagent comprises a nucleic acid.

[0022] In some embodiments, a 5' end and a 3' end of the Edman reagent each comprise a molecular lock. In some embodiments, the molecular lock is adapted to inhibit the translocation of an end of the sequencing substrate through a nanopore. In some embodiments, the molecular lock is selected from an oligonucleotide capable of forming a hairpin structure; biotin; a biotin derivative; streptavidin; a streptavidin derivative; or a polypeptide derived from a fibronectin-binding protein.

[0023] In some embodiments, the Edman reagent comprises in a 5' to 3!order: (i) a 5' phosphorylated end, (ii) a 5' molecular lock, (iii) the reactive moiety, (iv) an end-pomt tick mark, (v) a cycle marker, (vi) a primer binding site, and (vii) a 3' molecular lock.

[0024] In some embodiments, the Edman reagent comprises in a 5' to 3' order: (i) a 5' phosphorylated end, (ii) a 5' molecular lock, (iii) a cycle marker, (iv) a primer binding site, (v) the reactive moiety, (vi) an arresting moiety,, and (7) a 3' molecular lock.

[0025] In some embodiments, the target polypeptide has a length greater than 5, 10, 15, 20, 25, 30, 50, 100, 200, 500, 5000, 30,000 consecutive amino acid residues or wherein the target polypeptide has a length in a range from 5 to 1000, 5 to 500, 5 to 200, or 5 to 100 consecutive amino acid residues.

[0026] In some embodiments, the substrate comprises a bead or a flow cell.

[0027] In some embodiments, the substrate comprises a plurality of beads, wherein each bead comprises a target polypeptide different from one another.

[0028] In some embodiments, substrate comprises a plurality of clusters of target polypeptides, wherein each cluster comprises the same target polypeptide and the clusters comprise a different target polypeptide from one another. In some embodiments, the plurality of clusters of target polypeptides are located on the flow cell.

[0029] In some embodiments, the substrate comprises a bead; step (b) comprises coupling the N-terminal residue of the target polypeptide to the reactive moiety of an initial Edman reagent, wherein the initial Edman reagent comprises a cycle marker, and step (b) further comprises linking to the bead the initial Edman reagent coupled to the N-terminal residue; step (c) further comprises obtaining a modified initial Edman reagent coupled to the N-terminal residue and linked to the bead; and repeating steps (b) to (c) one or more times with an additional Edman reagent comprising a cycle marker, wherein the cycle marker of the additional Edman reagent is different from the cycle marker of a prior cycle.

[0030] Some embodiments of the methods and compositions provided herein include a method for preparing a sequencing substrate on a bead, comprising: (a) immobilizing a target polypeptide on a bead; (b) coupling the N-terminal residue of the target polypeptide to a reactive moiety of an initial Edman reagent, wherein the initial Edman reagent comprises a cycle marker; and (c) linking the initial Edman reagent to the bead; (d) cleaving the N-terminal residue from the target polypeptide, thereby obtaining a modified initial Edman reagentcoupled to the N-termmal residue and linked to the bead; (e) repeating steps (b) to (d) with an additional Edman reagent comprising a cycle marker, wherein the cycle marker of the additional Edman reagent is different from the cycle marker of a prior cycle to prepare the sequencing substrate on a bead.

[0031] Some embodiments of the methods and compositions provided herein include a method for preparing a sequencing substrate, comprising: (a) obtaining a target polypeptide comprising a barcode, wherein the barcode is indicative of the target polypeptide: (b) coupling an N-terminal residue of the target polypeptide to a reactive moiety of an initial Edman reagent, wherein the initial Edman reagent comprises a cycle marker; (c) extending the cycle marker to obtain an extended cycle marker by (i) hybridizing the cycle marker to the barcode and extending the cycle marker to obtain the extended cycle marker, or (ii) linking the cycle marker to the barcode to obtain the extended cycle marker; (d) cleaving the N-terminal residue from the target polypeptide, thereby obtaining a modified initial Edman reagent coupled to the N-termmal residue and comprising the extended cycle marker; and (e) repeating steps (b) to (d) with an additional Edman reagent comprising a cycle marker, wherein the cycle marker of the additional Edman reagent is different from the cycle marker of a prior cycle, and wherein the linking comprises linking the cycle marker of the additional cycle marker to the cycle marker of the prior cycle to prepare the sequencing substrate.

[0032] In some embodiments, the barcode is linked to the C-terminal residue of the target polypeptide.

[0033] In some embodiments, the barcode and cycle marker each comprise a polymer. In some embodiments, the polymer comprises a nucleic acid, PEG-based moiety, or a polypeptide.

[0034] In some embodiments, the target polypeptide or the Edman reagent, or both are in solution.

[0035] In some embodiments, step (c) comprises extending the cycle marker by polymerase extension.

[0036] In some embodiments, step (c) comprises linking the cycle marker by ligation.

[0037] Some embodiments of the methods and compositions provided herein include a sequencing substrate prepared according to any one of the methods for preparing a sequencing substrate provided herein.

[0038] Some embodiments of the methods and compositions provided herein include a method for characterizing a target polypeptide, comprising: (i) obtaining any one of the sequencing substrates provided herein; (ii) obtaining a nanopore embedded in a membrane, wherein the membrane has a cis and trans surface; (iii) translocating the sequencing substrate through the nanopore; and (iv) measuring a signal indicative of a coupled residue of the sequencing substrate in the nanopore.

[0039] In some embodiments, step (iii) comprises applying over the membrane: a potential difference, and / or an electroosmotic force. In some embodiments, step (iii) further comprises cycling the sequencing substrate located in the nanopore in a repeated movement towards the trans surface and to the cis surface,

[0040] In some embodiments, step (iv) comprises measuring a repeated signal. In some embodiments, step (iv) comprises measuring a signal generated by a coupled residue of the sequencing substrate located in a read region of the nanopore. Some embodiments also include measuring a signal indicative of a cycle moiety of the sequencing substrate in the nanopore. Some embodiments also include identifying the coupled residue of the sequencing substrate m the nanopore.

[0041] Some embodiments also include repeating steps (iii) and (iv), wherein the translocating comprises moving the sequencing substrate through the nanopore such that an arresting moiety pauses the translocating and the signal is measured.

[0042] Some embodiments also include moving the arresting moiety through the nanopore.

[0043] Some embodiments also include applying a pulsed potential difference over the membrane.

[0044] Some embodiments also include repeating steps (i) to (iv). Some embodiments also include repeating steps (i) to (iv) for a plurality of cycles.

[0045] In some embodiments, the cycle marker of the sequence substrate is different in each cycle.

[0046] In some embodiments, the translocating comprises contacting the sequencing substrate with a polymerase.

[0047] In some embodiments, the membrane comprises a lipid bilayer or block copolymer.

[0048] In some embodiments, the nanopore comprises a protein nanopore. In some embodiments, the protein nanopore is selected from OmpF, OmpG, CsgG, MspA, a-HL, FhuA, AeL, FraC, Lys, >29p, ClyA, Ply AB; MspA or CsgG.

[0049] In some embodiments, step (ii) comprises obtaining a plurality7of the nanopores embedded in the membrane, wherein a nanopore substrate comprises the plurality of nanopores.

[0050] In some embodiments, step (iii) comprises translocating a plurality of sequencing substrates through the plurality of nanopores.

[0051] In some embodiments, step (iv) comprises measuring a plurality of signals indicative of a coupled residue of a sequencing substrate in each nanopore of the nanopores.

[0052] In some embodiments, the nanopore substrate comprises a flow cell. In some embodiments, the flow cell comprises a plurality of wells, wherein each well comprises a nanopore.

[0053] Some embodiments of the methods and compositions provided herein include a system for preparing a sequencing substrate, comprising: (a) a target polypeptide immobilized on a substrate; (b) an Edman reagent comprising a reactive moiety adapted to couple with an N-ternunal residue of the target polypeptide; and (c) a reagent adapted to cleave the N-terminal residue from the target polypeptide, thereby obtaining the sequencing substrate comprising the Edman reagent and the coupled residue.

[0054] In some embodiments, the target polypeptide is immobilized on the substrate via a C-terminal residue of the target polypeptide.

[0055] In some embodiments, the reactive moiety comprises an isothiocyanate moiety or a phenyl isothiocyanate (PITC) moiety7.

[0056] In some embodiments, the Edman reagent comprises a polymer comprising a subunit selected from a nucleotide, an amino acid, a polyethylene glycol (PEG), a phosphonamidite, or a peptide nucleic acid (PNA) subunit.

[0057] In some embodiments, the Edman reagent comprises a cycle marker indicative of a position of a coupled residue in the target polypeptide. In some embodiments, the cycle marker is selected from a non-nucleosidic reporter moiety, a peptide reporter moiety, PEG-based reporter moiety,, or a single-stranded DNA reporter moiety.

[0058] In some embodiments, the Edman reagent comprises one or more arresting moieties adapted to pause or inhibit translocation of the sequencing substrate through a nanopore. In some embodiments, the reactive moiety is located in the Edman reagent between arresting moieties.

[0059] In some embodiments, the reagent adapted to cleave the N-terminal residue from the target polypeptide comprises an Edmanase, or a reagent or conditions sufficient to cleave the Edman reagent and the coupled residue from the target polypeptide,

[0060] In some embodiments, the Edman reagent is immobilized on the substrate via a cleavable linker.

[0061] Some embodiments also include an additional Edman reagent. In some embodiments, the additional Edman reagent is adapted to conjugate with an initial Edman reagent. In some embodiments, the additional Edman reagent comprises a cycle marker different from the initial cycle marker.

[0062] In some embodiments, the Edman reagent comprises a nucleic acid.

[0063] In some embodiments, a 5' end and a 3' end of the Edman reagent each comprise a molecular lock, wherein the molecular lock is adapted to inhibit the translocation of an end of the sequencing substrate through a nanopore. In some embodiments, the molecular lock is selected from an oligonucleotide capable of forming a hairpin structure; biotin; a biotin derivative; streptavidin; a streptavidin derivative; or a polypeptide derived from a fibronectin-binding protein.

[0064] In some embodiments, the Edman reagent comprises in a 5' to 3' order: (i) a 5' phosphorylated end, (ii) a 5' molecular lock, (iii) the reactive moiety, (iv) an end-point tick mark, (v) a cycle marker, (vi) a primer binding site, and (vii) a 3' molecular lock; or (i) a 5' phosphorylated end, (ii) a 5' molecular lock, (iii) a cycle marker, (iv) a primer binding site, (v) the reactive moiety, (vi) an arresting moiety, and (7) a 3' molecular lock.

[0065] In some embodiments, the target polypeptide has a length greater than 5, 10, 15, 20, 25, 30, 50, 100, 200, 500, 5000, 30,000 consecutive amino acid residues or whereintlie target polypeptide has a length in a range from 5 to 1000, 5 to 500, 5 to 200, or 5 to 100 consecutive amino acid residues.

[0066] In some embodiments, the substrate comprises a bead or a flow cell. In some embodiments, the substrate comprises a plurality of beads, wherein each bead comprises a target polypeptide different from one another. In some embodiments, the substrate comprises a plurality of clusters of target polypeptides, wherein each cluster comprises the same target polypeptide and the clusters comprise a different target polypeptide from one another. In some embodiments, the plurality of clusters of target polypeptides is located on a flow cell,

[0067] In some embodiments, the substrate comprises a bead; an initial Edman reagent comprises a cycle marker, and an additional Edman reagent comprising a cycle marker, wherein the cycle marker of the additional Edman reagent is different from the cycle marker of the initial cycle.

[0068] Some embodiments of the methods and compositions provided herein include a system for preparing a sequencing substrate, comprising: (a) a target polypeptide immobilized on a bead, wherein the target moiety comprises a barcode indicative of the target polypeptide; (b) an initial Edman reagent comprising a reactive moiety adapted to couple with the N -terminal residue of the target polypeptide, and a cycle marker; and (c) a reagent adapted for cleaving the N-terminal residue from the target polypeptide; and (d) an additional Edman reagent comprising a cycle marker different from the cycle marker of the initial Edman reagent. In some embodiments, the barcode is linked to the C-terminal residue of the target polypeptide. In some embodiments, the barcode and cycle marker each comprise a polymer. In some embodiments, the polymer comprises a nucleic acid or a polypeptide. In some embodiments, the target polypeptide or the Edman reagent, or both are in solution. In some embodiments, at least a portion of the barcode and of the cycle marker are adapted to hybridize to one another, and the system further comprises a polymerase. Some embodiments also include a reagent adapted to link the barcode and the cycle marker to one another.

[0069] Some embodiments of the methods and compositions provided herein include a system for characterizing a target polypeptide, comprising: (a) any one of the sequencing substrates provided herein or any one of the systems for preparing a sequencing substrate provided herein; (b) a membrane having a cis and trans surface; and (c) a nanopore embedded in the membrane. Some embodiments also include a nanopore substrate comprisinga plurality of nanopores embedded in the membrane. In some embodiments, the nanopore substrate comprises a flow cell. In some embodiments, the flow cell comprises a plurality of wells, wherein each well comprises a nanopore.BRIEF DESCRIPTION OF THE DRAWINGS

[0070] FIG. 1 depicts an overview of the Edman degradation cycle in which a free N-termmal residue of a polypeptide reacts with phenyl isothiocyanate (PITC) to form a phenylthiocarbamyl (PTC) intermediate, the PTC intermediate is removed from the polypeptide via an Edmanase-catalyzed reaction to expose a new N-terminal residue of the polypeptide (NFE-peptide).

[0071] FIG, 2 depicts an example reaction scheme comparing conventional Edman degradation with a DNA-compatible BFi. EtcO -mediated protocol.

[0072] FIG, 3 depicts an example mechanism of Edmanase-assi ted Edman degradation.

[0073] FIG. 4 depicts an example scheme for enhancement of ammo acid cleavage efficiency using a bifunctional PITC reagent and a chimeric Edmanase. Panels (A) and (B) depict a peptide attached to a solid-phase substrate is modified with a bifunctional NTAA modifier, such as biotin-phenyl isothiocyanate (PITC). Panel (C) depicts a low affinity Edmanase which is recruited to biotin-PITC labeled NTAAs using a streptavidin-Edmanase chimeric protein. Panel (D) depicts the efficiency of Edmanase cleavage is greatly improved due to the increase in effective local concentration as a result of the biotin-strepavidin interaction. Panel (E) depicts the cleaved biotin-PITC labeled NTAA and associated streptavidin-Edmanase chimeric protein diffuse away after cleavage. A number of other bioconjugation recruitment strategies can also be employed. An azide modified PITC is commercially available (4- Azidophenyl isothiocyanate, Sigma), allowing a number of simple transformations of azide-PITC into other bioconjugates of PITC, such as biotin-PITC via a click chemistry reaction with alkyne-biotin.

[0074] FIG. 5 depicts a schematic of a workflow for preparing a sequence substrate (nanopore sequenceable strand) including (1) immobilizing a target polypeptide on a substrate. (2) end-labeling the target polypeptide with an Edman reagent, such as an Edman reagent capable of linking to a capture probe attached to the substrate via a cleavable linker. TheEdman reagent can include adapters for linking to the capture probe and an additional Edman reagent, a barcode (BC), a reactive moiety, such as a PITC moiety capable of coupling with an N-termmal residue of the target polypeptide, and arresting moieties capable of pausing translocation of the sequence substrate through a nanopore. (3) coupling the reactive moiety with the N-terminal residue of the target polypeptide. Repeating steps (2) and (3) with an additional Edman reagent. (4) Cleaving the sequence substrate from the surface, and sequencing the sequence substrate using a nanopore embedded in a membrane.

[0075] FIG 6 depicts embodiments for attaching a capture probe via a cleavable linker to a substrate, or to a linker which attaches the target polypeptide to the substrate.

[0076] FIG 7 depicts embodiments of adaptors.

[0077] FIG, 8A depicts schematics of using a polyarginine peptide and an aerolysm nanopore

[0078] FIG, 8B depicts a graph for detecting various different residues,

[0079] FIG. 9 depicts graphs for various moieties useful for cycle markers including non-nucleosidic reporter, peptide reporter, and single stranded DNA reporter.

[0080] FIG. 10A depicts a schematic for templated selenocystine-selenoester and native chemical ligations.

[0081] FIG. 10B depicts graphs for templated selenocystine-selenoester and native chemical ligations.

[0082] FIG. 11 depicts a schematic workflow for de novo sequencing via an on- surface Edman degradation protein preparation and nanopore analysis.

[0083] FIG. 12. depicts a schematic workflow for de novo sequencing via an on- surface Edman degradation protein preparation and nanopore analysis.

[0084] FIG. 13 depicts a schematic workflow for de novo sequencing via an on- surface Edman degradation protein preparation in which a polypeptide of interest (POI), such as a target polypeptide, is immobilized on a bead.

[0085] FIG. 14 depicts a schematic workflow for de novo sequencing via an in¬ solution Edman degradation protein preparation in which a barcode is linked to the C-terminus of a target polypeptide.

[0086] FIG. 15 depicts a schematic workflow for de novo sequencing via an in¬ solution Edman degradation protein preparation in which a barcode is linked to the C-terminus of a target polypeptide.

[0087] FIG. 16A depicts a scheme for a chemical ligation strategy comprising an RNA-based reductive amination.

[0088] FIG. 16B depicts a scheme for a chemical ligation strategy comprising a templated “click” ligation.

[0089] FIG. 16C depicts a scheme for a chemical ligation strategy comprising a non-templated “click” ligation.

[0090] FIG. 17A depicts a scheme for a selective chemical ligation strategy comprising an enzyme-assisted chemical ligation which includes macrocyclization.

[0091] FIG. 17B depicts a scheme for a selective chemical ligation strategy comprising an enhanced sortase-ligation strategy.

[0092] FIG. 17C depicts a scheme for a selective chemical ligation strategy comprising a sortase-A-mediated strategy.

[0093] FIG. 17D depicts a scheme for a selective autoligation strategy that generates natural backbone mimics usable by enzymes.

[0094] FIG. 18A depicts an embodiment of a nanopore embedded in a membrane in which a sequencing substrate comprising a nucleic acid comprises hairpin structures at each end.

[0095] FIG. 18B depicts an embodiment in which a nanopore embedded in a membrane in which a sequencing substrate comprising a nucleic acid provides a first signal indicative of a residue of the sequencing substrate as a polymerase binds the sequencing substrate, and a second signal in the absence of the polymerase.DETAILED DESCRIPTION

[0096] Some embodiments of the methods and compositions provided herein relate to nanopore sequencing of a target polypeptide. In some embodiments, a sequencing substrate is prepared by Edman degradation from a target polypeptide immobilized on a substrate. De novo sequencing is performed with the sequence substrate and a nanopore embedded in a membrane.

[0097] Nanopore biosensing is based on naturally occurring protein pores. In a typical experiment, the pores are embedded in a lipid bilayer, which separates two chambers, cis and trans, filled w’ith an electrolyte solution. An applied voltage causes ions to move through the pore and create an electrical field. An analyte can be captured and transported across the pore by different mechanisms. Chinappi, M. et al., “Analytical model for particle capture in nanopores elucidates competition among electrophoresis, electroosmosis, and dielectrophoresis.” ACS Nano 2020, 14, 15816-15828. For charged analytes, electrophoresis may be the dominant form of transport, carrying the analyte toward the electrode of opposite polarity. Carson, S. et al., “Challenges in DNA motion control and sequence readout using nanopore devices.” Nanotechnology 2015, 26, 074004. For neutral or less-charged molecules, electroosmotic flow may be the dominant force directing capture and / or transport. Piguet, F et al., “Electroosmosis through a-hemolysin that depends on alkali cation type, ” J, Phys. Chem. Let. 2014, 5, 4362-4367; Gu, L. Q. el al., “Electroosmotic enhancement of the binding of a neutral molecule to a transmembrane pore.” Proc. Natl. Acad. Sci, USA 2003, 100, 15498- 15503. Furthermore, an analyte of appropriate size can enter the pore and alter the ionic current.

[0098] An analyte may change the ionic current by (i) producing a change in the electric field within the pore or through (ii) volume exclusion and binding of ions to the traversing analyte, which reduces the ionic current. Bezrukov, S. M. et al., “Current noise reveals protonation kinetics and number of ionizable sites in an open protein ion channel.” Phys. Rev. Lett. 1993, 70, 2352-2355; Remer, J. E. et al., “Theory for polymer analysis using nanopore- based single-molecule mass spectrometry.” Proc. Natl. Acad. Sci. USA 2010, 107, 12080-12085. Importantly, because numerous ions accompany the passage of a single analyte, a large electrical amplification occurs during a single-molecule translocation. Gurnev, P. A. et al., “Channel-forming bacterial toxins in biosensing and macromolecule delivery.” Toxins 2014, 6, 2483-2540. In a simplified scenario, the duration and amplitude of the current alteration is unique to each analyte or monomeric unit, in the case of polymer analytes. Nanopore- based sensing can employ both naturally occurring, protein pores and manufactured, solid-state nanopores. Xue, L. et al., “Solid-state nanopore sensors.” Nat. Rev. Mater. 2020, 5, 931-951. Biological nanopores, although not as robust as solid-state nanopores, offer better reproducibility due to their defined channel sizes. Wang, S. et al., Engineering of protein nanopores for sequencing, chemical or protein sensing and disease diagnosis. Curr. Opin.Biotechnol. 2018, 51, 80-89. In addition, lipid bilayers exhibit much lower electrical noise than the materials used to manufacture solid-state nanopores. Fragasso, A. et al., “Comparing current noise in biological and solid-state nanopores.” ACS Nano 2020, 14, 1338-1349.

[0099] For sensing purposes, nanopore technology generally exploits two distinct concepts: (i) direct sensing, in which changes in the current arise as the analyte traverses and directly interacts with the nanopore; and (ii) indirect sensing, in which translocation of an adapter molecule that specifically interacts with the analyte is monitored). Ayub, M. et al., “Engineered transmembrane pores,” Curr. Opin. Chem. Biol, 2016, 34, 117-126; Reynaud, L. et al., “Sensing with Nanopores and Aptamers: A Way Forward,” Sensor 2020, 20, 4495.

[0100] As used herein “peptide” and “polypeptide” can be used interchangeably and include a polymer comprising more than one amino acid residues linked via a peptide bond.

[0101] Next generation sequencing has revolutionized genomics but genetic information alone cannot predict protein abundance, post translational modifications or protein-driven biological processes. Proteomics on the other hand offers a direct insight into protein function, regulation and interactions. Combining proteomics with genomics, transcriptomics and metabolomics will bridge the genotype-phenotype gap and shed more information on novel biomarkers for disease diagnosis and precision medicine.

[0102] Mass spectroscopy has traditionally been the workhorse for high throughput proteomics profiling. The process begins with the digestion of protein samples followed by the introduction of peptide fragments to a mass spectrometer with quadrupole selectivity and high resolution accurate mass capability. Bottom-up proteomics identifies proteins by the sum of their parts but since the connectivity between peptides is lost, single peptides may be mapped to multiple entries in the protein database. This is especially prevalent with databases of higher eukaryotes due to the presence of related protein family members, alternative splice forms and partial sequences.

[0103] Nanopore-mediated protein sequencing enables the direct, real-time analysis of proteins at the single molecule level. This approach utilizes nanoscale pores to thread individual proteins and detect ionic current changes that occur as each amino acid passes. The protein sequence, including post-translational modifications, are then reconstructed from the unique current signatures generated. Given multiple amino acids residein the read head concurrently, complex data analysis and nanopore signal interpretation are typically associated with direct strand sequencing. None of the currently available biological nanopores are sensitive enough to differentiate between amino acids with similar steric volumes. In addition, unlike DNA, protein backbones are neutral and amino acid side chains are not uniformly charged, making translocation of proteins in a unidirectional and controlled manner difficult.

[0104] Edman degradation is a chemical reaction that enables the release of N-terminal ammo acids, one residue at a time, from a protein / peptide. First described in 1950, the 3 -step process begins with the selective modification of the N-termmal amine with phenyl isothiocyanate (PITC) in aqueous pyridine (FIG. 1). In the second step, the phenylthiocarbamoyl (PTC) intermediate is treated with neat trifluoroacetic acid. This yields a cyclic 2-anilino-5(4)-thiozolinone (ATZ)-modified amino acid and a truncated protein / peptide intermediate that is available for derivisation in a subsequent Edman reaction cycle. In the final step, the cleaved amino acid rearranges into a more stable phenylthiohydantoin (PTH) form in the presence of 20% aqueous TFA. This entire process is repeated in a cyclical manner until all or a defined portion of the protein / peptide sequence has been identified.

[0105] Since conventional Edman degradation employs harsh acidic conditions to initiate the cleavage step, applications involving DN / X may be at risk of depurination. To obviate this, in some embodiments, dA and dG may be substituted with 7-deazapurine deoxynucleotides (c7dA and c7dG) and N-termmal PITC-modified amino acids cleaved using boron trifluoride etherate as a milder alternative (FIG. 2). ATZ ammo acids are subsequently transformed to the more stable PTC amino acids under alkaline reducing conditions. Mild cleavage may also be achieved using a 1:1 mixture of TEAA: acetonitrile at 75 °C for 10 min (See e.g., Tetrahedron Lett. 1985, 26, 4375).

[0106] In more embodiments, Edmanases may be engineered to catalyse the cleavage step in aqueous buffer at neutral pH. Edmanase is an engineered cysteine protease with a C25G active site mutation that selectively acts on PITC-modified peptides See e.g.. Protein Sci. 2015, 24, 571). The reported catalytic efficiency, however, varies from amino acid to amino acid, with proline cleavage being the least efficient. An example mechanism of Edmanase-assisted Edman degradation and the associated catalytic efficiency for variousamino acid substrates is depicted in FIG. 3. Example kinetic parameters of Edmanase for PTC- Xaa substrates are listed in TABLE 1.TABLE 1Substrate kcat (s-1) KM (pM) kcat / KM (s-1M-1) PTC-Ala-AMC 0.55 (±0.013) 21.3 (±2.7) 2.6 x 104PTC-Asp-AMC 3.6 (±0.41) 124.5 (±35.0) 2.9 x 104PTC-Phe-AMC 0.47 (±0.060) 122.8 (±29.8) 3.8 x 103PTC-Met-AMC 0.54 (±0.083) 271.8 (±67.6) 2.0 x 103PTC-Pro-AMC 0.0014 (±0.0011) | 252.0 (±184.1) 5.7 x 101|PTC-Arg-AMC 0.087 (±0.017) 167.8 (±43.8) 5.2 x 102

[0107] To enhance catalytic efficiency, in some embodiments, peptides may be modified with a bifunctional biotin-PITC label followed by cleavage with a streptavidin-Edmanase chimeric protein. FIG. 4 depicts a schematic for enhancement of amino acid cleavage efficiency using a bifunctional PITC reagent and a chimeric Edmanase (See e.g., WO 2019089851 Al; US 20210355483). This effectively increases the local concentration as well as the affinity of the enzyme.De novo protein sequencing by transferring sequential amino acid information to a nanopore sequenceable strand

[0108] Some embodiments provided herein include a hybrid Edman degradation nanopore approach for de novo protein sequencing. Briefly, an N-terminal ammo acid residue of a target polypeptide is first labeled with a charged tag. The modified residue is then cleaved from the target polypeptide in a controlled manner via Edman degradation, either chemically or enzymatically. The presence of the charged tag increases the probability of nanopore capture of the cleaved residue and promotes unidirectional translocation of the cleaved residue through a nanopore. The cleaved residue is then identified based on specific changes in the ionic current readings before the entire process is repeated in a cyclic manner. In this set up, post-translational modifications can be detectable with up to single amino acid accuracy.

[0109] Some embodiments include a protein sequencing method comprising: (1) immobilization of a target polypeptide; (2) selective labeling of an N-terminal amine of aterminal residue of the target polypeptide with an Edman reagent comprising a reactive moiety which is covalently attached to a nanopore sequenceable strand; (3) cleavage of the N-terminal residue from the target polypeptide to release a modified Edman reagent comprising the sequenceable strand; and (4) nanopore capture and analysis of the modifed Edman reagent.

[0110] In some embodiments, the target polypeptide is first immobilised onto a surface containing a capture probe. Using an Edman degradation approach, the protein sequence is transferred to one or more nanopore sequenceable strands that can be concatemerized. Amino acid readings can then be achieved using a ratchet assay (FIG. 5).

[0111] (1) Immobilisation. A C-terminal end of the target polypeptide is modified with an appropriate conjugation handle to enable surface immobilization. In some embodiments, the target polypeptide can be digested into fragments using LysK or other appropriate protease, a side chain of a fragment can be selectively tagged and immobilized on the surface. A capture probe can include a cleavable linker that allows strand release in a later step and may be directly attached to the surface or to the tether connecting the protem / peptide to the surface (FIG. 6).

[0112] (2) Chemoselective end labeling. The N-terminal amine of the target polypeptide reacts exclusively with a reactive moiety of a nanopore sequenceable strand. The reactive moiety can include PITC under used under slightly basic conditions (pH 8). After labelling, a washing step is conducted to remove excess unreacted oligonucleotides in the medium.

[0113] The nanopore sequenceable strand can include: (a) an adapter; (b) an arresting construct (ARC); (c) a tether; and (d) a cycle number barcode.

[0114] (a) Adapter. To promote the concatenation of nanopore sequenceable strands, the adapters at both termini are designed to be non-complementary. Minimally 2 adapter sets can be used to differentiate between odd and even Edman cycles (FIG 7).

[0115] (b) Arresting construct (ARC). Two arresting constructs can be used. The first arresting construct positions the cycle number barcode in a nanopore read head. After a voltage pulse is applied, the second arresting construct would then position the cleaved amino acid in the read head for identification.

[0116] (c) Tether. The tether may be peptide-, PNA- or DNA-based and may contain a blend of PEG, amino acid, PNA monomer and / or phosphoramidite units. Uniformnegative charges are positioned along the backbone of the tether to ensure unidirectional translocation through the pore. The tether region adjacent to PITC is expected to be confined in the sensing region of the pore along with a cleaved residue. In order to achieve distinct current blocks with each residue and reduce signal complexity, the tether can be made up of a string of homooligomeric PEG monomers, nucleotide units or amino acids units. Polyarginines in particular have previously been shown to allow the unambiguous distinction of 13 of the 20 proteinogenic amino acids using the aerolysin pore (FIG. 8A, FIG. 8B).

[0117] (d) Cycle number barcode. A barcode can be inserted between the first arresting construct and the adapter to indicate the cycle number of sequencing, such as a corresponding ammo acid position in the target polypeptide. Barcodes may be introduced in a cyclic manner, where a single barcode, such as ‘1’, can represent multiple cycles, such as 1st, 11th, 21st, etc. A set of 6-10 reporter moieties with non-overlapping signals have been identified (FIG 9). Specifically, non-nucleosidic reporter moieties can be utilized in conjunction with ssDNA reporters and a DNA-based tether; whereas peptide reporter moieties can be used with a peptide tether.

[0118] (3) Ligation and Edman cleavage. Ligation of the adapter to the capture probe can be achieved chemically or with the aid of ligases or polymerases (FIG. 10A and FIG.10B). Native chemical ligation between the N-terminal cysteine of one peptide and the C- termmal thioester of its reactive partner has been shown to proceed significantly faster in a templated format than the corresponding non- temp lated transformation (See e.g., J. Am. Chem. Soc. 2004, 124, 9970). An analogous diselenide-selenoester ligation has also been shown to have fast kinetics (See e.g., Chem. Sci. 2018, 9, 896). To prevent bimolecular reaction between the capture probe and excess nanopore sequenceable strands in solution, the cysteine and selenocysteine group may be modified with a protecting group that is unveiled only after the PITC modification had taken place. More example embodiments for ligation are provided herein.

[0119] Subsequently, the PITC-modified terminal acid is cleaved. PNA-based adapters may be cleaved using the enzymatic or chemical strategies described in the previous section. To minimize the risk of depurination, DNA-based adapters are cleaved using Edmanases or mild Edman cleavage conditions. Protein engineering may be required to effect efficient cleavage of a PITC moiety linked to additional functional groups. Additionally, theEdmanase may be tethered to the surface to increase the effective local concentration and PITC binding affinity.

[0120] (4) Strand release and sequencing. The concatenated nanopore sequenceable strand is cleaved from the surface and captured on the pore. A voltage pulse is applied until the first arresting construct suspends the cycle number barcode in the read head. Once the cycle number is decoded, another voltage pulse is applied such that the corresponding cleaved amino acid can be read. This process is repeated until the entire protein sequence is determined, one amino acid at a time,

[0121] In some embodiments, a modified oligonucleotide tag is used as the mam reactant for de novo sequencing (FIG. 11). Arranged from 5' to 3', the oligonucleotide tag has: (1 ) a 5' phos, (2) a 5' locking group, (3) a PITC functional handle covalently linked close to the 5' lock, (4) an end-point tick-mark sequence, (5) a unique molecular identifier (UMI) sequence, (6) a primer-binding sequence, and finally (7) a 3' lock. Upon addition onto the immobilized target polypeptide with a free N-terminus, the PITC handle would react with the N-terminal residue to form a PITC-AA group, which covalently links the target polypeptide with the oligonucleotide tag. Subsequently, excess non-linked oligonucleotide tags can be removed from the surface, leaving behind only the peptide-oligo conjugates attached on the surface. The PITC-AA group is then liberated with the oligonucleotide tag upon chemical or enzymatic cleavage. The resultant molecule is then flowed into the chip carrying a nanopore system for subsequent capture and sequencing via a nanopore assay such as the ‘sequencing in a cis well’ depicted in FIG. 11 (right-most portion), and disclosed in certain embodiments in US 20230090867 which is incorporated by reference in its entirety. In some such embodiments, a single-stranded target polynucleotide to be sequenced is disposed through the aperture of a nanopore. A duplex is formed with a portion of the target polynucleotide, on a first side of the nanopore. A force then is applied that lodges the duplex within the aperture of the nanopore, and an electrical measurement is made. The particular value of that measurement is based on the particular complementary bases that are located at the 3' end of the duplex, and the particular sequence of bases that are in a single-stranded portion of the target polynucleotide, within the aperture of the nanopore. Accordingly, the particular measured value provides information from which the sequence of the bases in the target polynucleotide may be determined. A force then is applied that dislodges the duplex from within the apertureof the nanopore so that the duplex may be extended by a nucleotide, and the measurement repeated. The repeated measurements provide further information from which the sequence of the bases in the target polynucleotide may be determined.

[0122] In another embodiment, a modified oligonucleotide tag is used as the main reactant for de novo sequencing (FIG. 12). Arranged from 5' to 3', the oligonucleotide tag has: (1) a 5' phos, (2) a 5' locking group, (3) a unique molecular identifier (UMI) sequence, (4) a primer-binding sequence, (5) a PITC functional handle, (6) an arresting construct (ARC), and finally (7) a 3' lock. Upon addition onto the immobilized target polypeptide with a free N-terminus, the PITC handle would react with the N-terminal residue to form a PITC-AA group, which covalently links the target polypeptide with the oligonucleotide tag. Subsequently, the excess non-linked oligonucleotide tags are removed from the surface, leaving behind only the peptide-oligo conjugates attached on the surface. The PITC-AA group is then liberated with the oligonucleotide tag upon chemical or enzymatic cleavage. The resultant molecule is then flowed into the chip carrying a nanopore system for subsequent capture and sequencing via a dual Ratchet and the nanopore assay such as the ‘sequencing in a cis well’ depicted in FIG. 11 and FIG. 12 (right-most portion). The Ratchet read is performed first to read the PITC-AA molecular signature by pulsing to arrest next to the ARC handle, and the sequencing read is performed last to read the UMI sequence.

[0123] Another example embodiment is depicted in FIG. 13 which relates to a workflow for de novo sequencing via an on-surface Edman degradation protein preparation in which a polypeptide of interest (POI), such as a target polypeptide, is immobilized on a bead. Some such embodiment can include massive parallel processing of target polypeptides. Some such embodiments can include protein fingerprinting, for example, by determining certain amino acid residues or types of amino acid residues, such as a ratio of such amino acid residues or types, and determining the target polypeptide from such a determination. In the workflow, an Edman reagent comprising a phosphate group (P), a PITC reactive moiety, and a Step-ID such as a cycle marker, is coupled to the N-terminal residue of the target polypeptide. In some embodiments, the Edman reagent can include one more of a 5'-phosphate, a primer binding region and / or arresting construct / moiety. The coupled / modified Edman reagent is ligated to the bead via a linker, and is cleaved from the target polypeptide. The coupling, ligating and cleaving steps are repeated to obtain a sequencing substrate for analysis with a nanopore. Somesuch embodiments include spatial separation of target polypeptides with barcoded beads or of target polypeptides in clusters on a surface. In some embodiments, a POI, such as a target polypeptide is immobilized on a bead. The N-terminal residue of the target polypeptide is modified by contacting the N-terminal residue with a first Edman reagent in which the first Edman reagent comprises a reactive moiety such as a PITC moiety and a first step-ID such as a cycle marker. In some embodiments, the reactive moiety couples with the N-terminal residue of the target polypeptide to obtain a first coupled Edman reagent. The first coupled Edman reagent is ligated to the bead via a linker, and the first coupled Edman reagent is cleaved from the target polypeptide. The steps of modifying, ligating and cleaving can be repeated. For example, the N-terminal residue of the target polypeptide is modified with a second Edman reagent comprising a second step ID; the second coupled Edman reagent is ligated to the bead via a linker, and the second coupled Edman reagent is cleaved from the target polypeptide. In some such embodiments, a bead is obtained in which the bead comprises a plurality of different coupled Edman reagents, each coupled Edman reagent comprising a coupled residue from the target polypeptide and a step ID in which the step ID is indicative of a cycle in the repeated method. In some embodiments, the coupled Edman reagents of the bead can be cleaved from the bead characterized via a nanopore embedded in a membrane. In some embodiments, a plurality of beads in which each bead comprises a different target polypeptide from another bead of the plurality of beads can be modified, ligated and cleaved to prepare a plurality of beads in which each bead comprises a plurality of different coupled Edman reagents, each coupled Edman reagent comprising a coupled residue from the target polypeptide and a step ID in which the step ID indicative of a cycle in the repeated method.

[0124] Another example embodiment is depicted in FIG. 14 which relates to a workflow for de novo sequencing via an in-solution Edman degradation protein preparation in which a barcode is linked to the C-terminus of a target polypeptide. Some such embodiment can include massive parallel processing of target polypeptides. Some such embodiments can include protein fingerprinting, for example, by determining certain amino acid residues or types of amino acid residues, such as a ratio of such amino acid residues or types, and determining the target polypeptide from such a determination. In the workflow, the barcode can include a region for hybridizing with a portion of a step-ID such as cycle marker, and encode a sequence indicative of the target polypeptide. An Edman reagent comprising aphosphate group (P), a PITC reactive moiety, and a Step-ID such as a cycle marker, is coupled to the N-terminal residue of the target polypeptide. In some embodiments, the Edman reagent can include one more of a 5 '-phosphate, a primer binding region and / or arresting construct / moiety. A portion of the cycle marker and a portion of the barcode can hybridize to one another. The cycle marker can be extended, thereby copying the barcode or complement thereof. The coupled / modified Edman reagent is cleaved from the target polypeptide. The coupling, hybridizing and extending steps are repeated. Some such embodiments include processing a target polypeptide in solution. In some embodiments, a peptide-ID such as a barcode is linked to the C -terminal residue of a target polypeptide. The barcode is indicative of the target polypeptide. In some embodiments, the barcode comprises a nucleic acid. The N-terminal is modified by contacting the N-terminal residue with a first Edman reagent in which the first Edman reagent comprises a reactive moiety, such as a PITC moiety and a first step-ID such as a cycle marker. In some embodiments, the cycle marker comprises a nucleic acid. In some such embodiments, the reactive moiety couples with the N-terminal residue of the target polypeptide to obtain a first coupled Edman reagent. In some embodiments, the cycle marker is extended by copying a portion of the barcode. For example, the cycle marker hybridizes to a portion of the barcode and is extended by polymerase extension. The coupled Edman reagent comprising the extended cycle marker is cleaved from the target polypeptide. The steps of coupling an Edman reagent to the N-terminal residue of the target polypeptide, extending the cycle marker, and cleaving the coupled Edman reagent comprising an extended cycle marker can be repeated in which the cycle marker is different from a cycle marker of a prior cycle.

[0125] Another example embodiment is depicted in FIG. 15 which relates to a workflow for de novo sequencing via an in-solution Edman degradation protein preparation in which a barcode is linked to the C-terminus of a target polypeptide. Some such embodiment can include massive parallel processing of target polypeptides. Some such embodiment can include de novo sequencing of target polypeptides. In the workflow, the barcode can encode a sequence indicative of the target polypeptide. An Edman reagent comprising a phosphate group (P), a PITC reactive moiety, and a Step-ID such as a cycle marker, is coupled to the N-terminal residue of the target polypeptide. In some embodiments, the Edman reagent can include one more of a 5 ’-phosphate, a primer binding region and / or arresting construct / moiety.The cycle marker is linked to an end of the barcode. The coupled / modified Edman reagent is cleaved from the N-terminal of the target polypeptide. The coupling, linking and cleaving steps are repeated with additional Edman reagents. Some such embodiments include processing a target polypeptide in solution, in which a peptide-ID of a target polypeptide is extended with a coupled Edman reagent. In some embodiments, a peptide-ID, such as a barcode, is linked to the C-terminal residue of a target polypeptide. The N-terminal is modified by contacting the N-terminal residue with a first Edman reagent in which the first Edman reagent comprises a reactive moiety such as a PITC moiety and a first step-ID such as a cycle marker. In some embodiments, the cycle marker comprises a nucleic acid. In some such embodiments, the reactive moiety couples with the N-terminal residue of the target polypeptide to obtain a first coupled Edman reagent. In some embodiments, the cycle marker is extended by linking, coupling or ligating to the barcode of the target polypeptide. The N-terminal residue of the coupled Edman reagent is cleaved from the target polypeptide, such that the coupled Edman reagent remains linked to the barcode of the target polypeptide. The steps of coupling an Edman reagent to the N-terminal residue of the target polypeptide, extending the cycle marker, and cleaving the coupled Edman reagent comprising an extended cycle marker can be repeated in which the cycle marker is different from a cycle marker of a prior cycle.Certain embodiments for ligation

[0126] Some of the methods and compositions provided herein include chemically ligating moieties to one another, such as barcodes, cycle markers of Edman reagents, and / or coupled Edman reagents. Example embodiments useful with the methods and compositions provided herein include the templated selenocystine-selenoester and native chemical ligation depicted in FIG. 10A. More examples include RNA-based reductive amination depicted in 16A; templated ‘click’ ligations depicted in FIG. 16B which can be selective for matching complementary sequence; and non-templated ‘click’ ligations depicted in FIG. 16C, which can be useful for linking non-complementary sequences for up to 5 kb oligonucleotides.

[0127] Some embodiments include enzyme-assisted chemical ligation which can include macrocyclization, and may not include oligonucleotides (FIG. 17A). See e.g., W. Liu, et. al. Nat. Chem. 2025, 25. Some embodiments include an enhanced sortase-ligation strategy which can include a C-terminus to N-terminus ligation, irreversible charge-guided capture andefficient and traceless sortase A-mediated transpeptidati on (FIG. 17B). Some embodiments include a pseudokinase application in which there can be covalent linkage upon hybridization, and can include other enzyme-based ligation, such as sortase (FIG. 17C). Wang C., etal., Org. Lett. 2025, 27, 42, 11854-11858. Some embodiments include autoligation that can generate natural backbone mimics usable by enzymes such as polymerases (FIG. 17D). K. Yamaoka, et. al. ChemBioChem. 2021, 22, 3273. Some embodiments include reagent-free DNA autoligation strategies based on intramolecular cross-activation between 3'-phosphorothioate (PS) and 5'-dinitrobenzenesulfonamide (DNBSA) on a splint DNA. Some such embodiments can yield a P3'→N5' phosphoramidate linkage under near-physiological conditions with ligation proceeding with over 80% yield at 37 °C and pH 8 without external reagents, H. Yokoe, et. al., Commun. Chem. 2025, 232, 8.Certain methods for preparing a sequencing substrate

[0128] Some embodiments of the methods and compositions provided herein include a method for preparing a sequencing substrate for a nanopore. The sequencing substrate can be prepared from a target polypeptide using an Edman reagent. Some embodiments can include coupling the N-terminal residue of a target polypeptide with an Edman reagent; cleaving the N-terminal residue from the target polypeptide to obtain a modified Edman reagent; and characterizing the modified Edman reagent with a nanopore. In some embodiments, a sequencing substrate can include a modified Edman reagent, or more than one modified Edman reagents linked together. In some such embodiments comprising linked modified Edman reagents, each modified Edman reagent can include a different cycle marker.

[0129] Some embodiments of methods for preparing a sequencing substrate can include: (a) immobilizing a target polypeptide on a substrate, such as a solid substrate, such as a planar substrate or bead; (b) coupling an N-terminal residue of the target polypeptide to a reactive moiety of an Edman reagent; and (c) cleaving the N-terminal residue from the target polypeptide, thereby obtaining the sequencing substrate comprising the Edman reagent and the coupled residue.

[0130] In some embodiments, the target polypeptide is immobilized on the substrate via a C-terminal residue of the target polypeptide. In some embodiments, the targetpolypeptide is covalently immobilized on the substrate. In some embodiments, the target polypeptide is immobilized on the substrate via a cleavable linker.

[0131] In some embodiments, the reactive moiety comprises an isothiocyanate moiety. In some embodiments the reactive moiety comprises a phenyl isothiocyanate (PITC) moiety.

[0132] In some embodiments, the Edman reagent comprises a net negative charge.

[0133] In some embodiments, the Edman reagent comprises a polymer. In some embodiments, the polymer is a homopolymer. In some embodiments, the polymer comprises a subunit selected from a nucleotide, an amino acid, a polyethylene glycol (PEG), a phosphonamidite, or a peptide nucleic acid (PNA) subunit. In some embodiments, the polymer comprises a plurality of arginine residues,

[0134] In some embodiments, the Edman reagent comprises a cycle marker indicative of a position of a coupled residue in the target polypeptide. In some embodiments, the cycle marker is selected from a non-nucleosidic reporter moiety, a peptide reporter moiety, PEG-based reporter moiety, or a single-stranded DNA reporter moiety.

[0135] In some embodiments, the Edman reagent comprises an arresting moiety adapted to pause or inhibit translocation of the sequencing substrate through a nanopore.

[0136] In some embodiments, step (c) comprises contacting the N-terminal residue with an Edmanase or with one or more reagents adapted to cleave the Edman reagent coupled to N-terminal residue from the target polypeptide. In some embodiments, the Edmanase is immobilized on the substrate. In some embodiments, the Edmanase is in solution.

[0137] In some embodiments, the Edmanase is immobilized on the substrate. In some embodiments, the Edman reagent is immobilized on the substrate via a cleavable linker. In some embodiments, the cleavable linker comprises a cysteine and selenocysteine group modified with a protecting group. In some embodiments, the reactive moiety is located in the Edman reagent between arresting moieties. Some embodiments also include repeating steps (b) and (c) with an additional Edman reagent, wherein the immobilized Edman reagent and the additional Edman reagent each comprise an adapter adapted to conjugate to one another. Some embodiments also include repeating steps (b) and (c) for a plurality of cycles, wherein each additional Edman reagent comprises a different cycle marker. Some embodiments also include conjugating the immobilized Edman reagent to the additional Edman reagent. In someembodiments, each adapter comprises a nucleic acid. In some embodiments, the adapters of the immobilized Edman reagent and of the additional Edman reagent are complementary to one another. In some embodiments, the conjugating comprises a polymerase reaction or a ligase reaction. Some embodiments also include step (d) cleaving the immobilized sequencing substrate from the substrate.

[0138] In some embodiments, the Edman reagent is in solution. In some embodiments, the Edman reagent comprises a nucleic acid. In some embodiments, a 5' end and a 3' end of the Edman reagent each comprise a molecular lock. In some embodiments, the molecular lock is adapted to inhibit the translocation of an end of the sequencing substrate through a nanopore. In some embodiments, the molecular lock is selected from an oligonucleotide capable of forming a hairpin structure; biotin; a biotin derivative; streptavidin; a streptavidin derivative; or a polypeptide derived from a fibronectin-binding protein. Example embodiments of polypeptide derived from a fibronectin-binding protein are disclosed in Zakeri, B, el al., (2012) PNAS 109 (12) E690-E697.

[0139] In some embodiments, the Edman reagent comprises in a 5' to 3' order: (i) a 5' phosphorylated end, (ii) a 5' molecular lock, (iii) the reactive moiety, (i v) an end-point tick mark, (v) a cycle marker, (vi) a primer binding site, and (vii) a 3' molecular lock.

[0140] In some embodiments, the Edman reagent comprises in a 5' to 3' order: (i) a 5' phosphorylated end, (ii) a 5' molecular lock, (iii) a cycle marker, (iv) a primer binding site, (v) the reactive moiety, (vi) an arresting moiety, and (7) a 3' molecular lock.

[0141] In some embodiments, the target polypeptide has a length greater than 5, 10, 15, 20, 25, 30, 50, 100, 200, 500, 5000, 30,000 consecutive amino acid residues. In some embodiments the target polypeptide has a length in a range from 5 to 1000, 5 to 500, 5 to 200, or 5 to 100 consecutive amino acid residues.

[0142] In some embodiments, the substrate comprises a flow cell. In some embodiments, the substrate comprises a bead. In some such embodiments, the substrate comprises a plurality of beads, wherein each bead comprises a target polypeptide different from one another.

[0143] In some embodiments, the substrate comprises a bead; step (b) comprises coupling the N-terminal residue of the target polypeptide to the reactive moiety of an initial Edman reagent, wherein the initial Edman reagent comprises a cycle marker, and step (b) alsocomprises linking to the bead the initial Edman reagent coupled to the N-terminal residue; step (c) also comprises obtaining a modified initial Edman reagent coupled to the N-terminal residue and linked to the bead; and the method also includes repeating steps (b) to (c) one or more times with an additional Edman reagent comprising a cycle marker, wherein the cycle marker of the additional Edman reagent is different from the cycle marker of a prior cycle.

[0144] Some embodiments include a method for preparing a sequencing substrate, comprising: (a) immobilizing a target polypeptide on a bead; (b) coupling the N-terminal residue of the target polypeptide to a reactive moiety of an initial Edman reagent, wherein the initial Edman reagent comprises a cycle marker; and (c) linking the initial Edman reagent to the bead; (d) cleaving the N-terminal residue from the target polypeptide, thereby obtaining a modified initial Edman reagent coupled to the N-terminal residue and linked to the bead; (e) repeating steps (b) to (d) with an additional Edman reagent comprising a cycle marker, wherein the cycle marker of the additional Edman reagent is different from the cycle marker of a prior cycle,

[0145] Some embodiments include a method for preparing a sequencing substrate, comprising: (a) obtaining a target polypeptide comprising a barcode, wherein the barcode is indicative of the target polypeptide; (b) coupling an N-terminal residue of the target polypeptide to a reactive moiety of an initial Edman reagent, wherein the ini tial Edman reagent comprises a cycle marker; (c) extending the cycle marker to obtain an extended cycle marker by (i) hybridizing the cycle marker to the barcode and extending the cycle marker to obtain the extended cycle marker, or (ii) linking the cycle marker to the barcode to obtain the extended cycle marker; (d) cleaving the N-terminal residue from the target polypeptide, thereby obtaining a modified initial Edman reagent coupled to the N-terminal residue and comprising the extended cycle marker; and (e) repeating steps (b) to (d) with an additional Edman reagent comprising a cycle marker, wherein the cycle marker of the additional Edman reagent is different from the cycle marker of a prior cycle, and wherein the linking comprises linking the cycle marker of the additional cycle marker to the cycle marker of the prior cycle. In some embodiments, the barcode is linked to the C-terminal residue of the target polypeptide. In some embodiments, the barcode and cycle marker each comprise a polymer. In some embodiments, the polymer comprises a nucleic acid, PEG-based moiety,, or a polypeptide. In some embodiments, the target polypeptide or the Edman reagent, or both are in solution. In someembodiments, step (c) comprises extending the cycle marker by polymerase extension. In some embodiments, step (c) comprises linking the cycle marker by ligation.Certain methods for characterizing a target polypeptide

[0146] Some embodiments of the methods and compositions provided herein include a method for characterizing a target polypeptide, comprising: (a) preparing a sequencing substrate according to any one of the methods provided herein for preparing a sequencing substrate; (b) obtaining a nanopore embedded in a membrane, wherein the membrane has a cis and trans surface; (c) translocating the sequencing substrate through the nanopore; and (d) measuring a signal indicative of a coupled residue of the sequencing substrate in the nanopore.

[0147] In some embodiments, step (c) comprises applying over the membrane: (i) a potential difference, and / or (ii) an electroosmotic force.

[0148] In some embodiments, step (c) further comprises cycling the sequencing substrate located in the nanopore in a repeated movement towards the trans surface and to the cis surface. In some embodiments, step (d) comprises measuring a repeated signal. In some embodiments, step (d) comprises measuring a signal generated by a coupled residue of the sequencing substrate located in a read region of the nanopore. Some embodimen ts also include measuring a signal indicative of a cycle moiety of the sequencing substrate in the nanopore.

[0149] Some embodiments also include identifying the coupled residue of the sequencing substrate in the nanopore.

[0150] Some embodiments also include repeating steps (c) and (d), wherein the translocating comprises moving the sequencing substrate through the such that an arresting moiety pauses the translocating and the signal is measured. Some embodiments also include moving the arresting moiety through the nanopore. Some embodiments also include applying a pulsed potential difference over the membrane.

[0151] Some embodiments also include repeating steps (a) to (d). Some embodiments also include repeating steps (a) to (d) for a plurality of cycles. In some embodiments, the cycle marker of the sequence substrate is different in each cycle.

[0152] In some embodiments, the translocating comprising contacting the sequencing substrate with a polymerase.

[0153] In some embodiments, the membrane comprises a lipid bilayer or block copolymer. In some embodiments, the membrane comprises a lipid bilayer or a block copolymer. Each molecule of a block copolymer may include one or more hydrophilic blocks and one or more hydrophobic blocks. The hydrophilic blocks may form outer surfaces of the barrier and the hydrophobic blocks may be located within the barrier. The hydrophobic blocks may include a polymer selected from the group consisting of poly(dimethylsiloxane) (PDMS), polybutadiene (PBd), polyisoprene, polymyrcene, polychloroprene, hydrogenated polydiene, fluorinated polyethylene, polypeptide, and poly(isobutylene) (PIB). See e.g., U. S.20230312856 which is incorporated by reference in its entirety.

[0154] In some embodiments, the nanopore comprises a protein nanopore. In some embodiments, the protein nanopore is selected from OmpF, OmpG, CsgG, MspA, a -HL, FhuA, AeL, FraC, Lys, φ29p, ClyA, Ply AB. In some embodiments the protein nanopore is MspA or CsgG.

[0155] Some embodiments comprise a plurality of nanopores. Some such embodiments can include a nanopore substrate comprising the plurality of nanopores. In some embodiments, the nanopore substrate comprises a flow cell. In some embodiments, the flow cell comprises a plurality of wells, each well comprising a nanopore.

[0156] In some embodiments, a sequencing substrate can be characterized with a nanopore, in which the sequencing substrate comprises a nucleic acid. In some such embodiments, the nucleic acid can comprise hairpin secondary structures, thereby pausing the translocation of the sequencing substrate through the nanopore and increasing the duration of a signal (FIG. 18A). In some embodiments, the nucleic acid can comprise a molecular lock, the Some embodiments of the methods and compositions provided herein can include aspects of the disclosure of U. S. 2025 / 0215487 and WO 2023 / 049682 which are each incorporated by reference in its entirety. In some embodiments, a polymerase can bind to a sequencing substrate and while the polymerase is bound, a first type of signal indicative of a residue coupled to the sequencing substrate can be measured. The polymerase can be removed, and a second type of signal can be measured indicative of a sequence, such as a cycle marker or extended barcode (FIG. 18B). Some embodiments of the methods and compositions provided herein can include aspects of the disclosure of (1) US 63 / 890,914, filed September 30, 2025, “NANOPORE SEQUENCING TECHNIQUES”; (2) US 63 / 890,788, filed September 30,2025, “IDENTIFYING AMINO ACIDS IN POLYPEPTIDES USING NANOPORES”; and (3) US 63 / 783,492, filed April 4, 2025, “DUPLEX DNA AND PEPTIDE SEQUENCING” which are each incorporated by reference in its entirety.

[0157] Certain aspects of methods and compositions useful in detecting, measuring and determining signals from a nanopore are disclosed in Chan Cao et al., Sci. Adv. 6, eabc2661 (2020). DOI: 10.1126 / sciadv.abc2661; and Zheng-Li, H, et al., (2018) Anal. Chem. 90:4268-4272 which are each incorporated by reference in its entirety.Certain compositions, systems and kits

[0158] Some embodiments of the methods and compositions provided herein include a sequencing substrate prepared according to any one of the methods for preparing a sequencing substrate provided herein. Some embodiments of the methods and compositions provided herein include systems for characterizing a target polypeptide, comprising: (a) a sequencing substrate prepared according to any one of the methods for preparing a sequencing substrate provided herein; (b) a membrane having a cis and trans surface; and (c) a nanopore embedded in the membrane. Some embodiments of the methods and compositions provided herein include systems for characterizing a target polypeptide, comprising: (a) a target polypeptide immobilized on a surface; (b) any one of the Edman reagents provided herein; (c) a membrane having a cis and trans surface; and (d) a nanopore embedded in the membrane. In some embodiments, the membrane comprises a lipid bilayer or block copolymer. In some embodiments, the nanopore comprises a protein nanopore. In some embodiments, the protein nanopore is selected from OmpF, OmpG, CsgG, MspA, a-HL, FhuA, AeL, FraC, Lys, φ29p, ClyA, PlyAB. In some embodiments the protein nanopore is MspA or CsgG. Some embodiments comprise a plurality of nanopores. Some such embodiments can include a nanopore substrate comprising the plurality of nanopores. In some embodiments, the nanopore substrate comprises a flow cell. In some embodiments, the flow cell comprises a plurality of wells, each well comprising a nanopore.

[0159] Some embodiments of the methods and compositions provided herein include systems and kits for preparing a sequencing substrate. Some such embodiments include: (a) a target polypeptide immobilized on a substrate, wherein the target polypeptide comprises an N-terminal residue; (b) an Edman reagent comprising a reactive moiety,, whereinthe reactive moiety is adapted to couple the Edman reagent with the N-terminal residue; and (c) an Edmanase or reagents adapted to cleave the Edman reagent coupled to N-terminal residue from the target polypeptide. In some embodiments, the Edmanase is immobilized on the substrate. In some embodiments, the Edmanase is in solution.

[0160] In some embodiments, the target polypeptide is immobilized on the substrate via a C-terminal residue of the target polypeptide. In some embodiments, the target polypeptide is covalently immobilized on the substrate. In some embodiments, the target polypeptide is immobilized on the substrate via a cleavable linker.

[0161] In some embodiments, the reactive moiety comprises an isothiocyanate moiety or a phenyl isothiocyanate (PITC) moiety.

[0162] In some embodiments, the Edman reagent comprises a net negative charge. In some embodiments, the Edman reagent comprises a polymer. In some embodiments, the polymer is a homopolymer. In some embodiments, the polymer comprises a subunit selected from a nucleotide, an amino acid, a polyethylene glycol (PEG), a phosphonamidite, or a peptide nucleic acid (PNA) subunit. In some embodiments, the polymer comprises a plurality of arginine residues.

[0163] In some embodiments, the Edman reagent comprises a cycle marker indicative of a position of a coupled residue in the target polypeptide. In some embodiments, the cycle marker is selected from a non-nucleosidic reporter moiety, a peptide reporter moiety, PEG-based reporter moiety, or a single-stranded DNA reporter moiety.

[0164] In some embodiments, the Edman reagent comprises an arresting moiety adapted to pause or inhibit translocation of the sequencing substrate through a nanopore. In some embodiments, the reactive moiety is located in the Edman reagent between two arresting moieties.

[0165] In some embodiments, the Edman reagent is immobilized on the substrate via a cleavable linker. In some embodiments, the cleavable linker comprises a cysteine and selenocysteine group modified with a protecting group. Some embodiments also include an additional Edman reagent, wherein the immobilized Edman reagent and the additional Edman reagent each comprise an adapter adapted to conjugate to one another. In some embodiments, the additional Edman reagent and the immobilized Edman reagent comprise a cycle marker different from one another. In some embodiments, each adapter comprises a nucleic acid. Insome embodiments, the adapters of the immobilized Edman reagent and of the additional Edman reagent are complementary to one another. In some embodiments, the adapters are adapted to be conjugated to one another via a polymerase reaction or a ligase reaction. Some embodiments also include a reagent adapted to cleave the immobilized sequencing substrate from the substrate.

[0166] In some embodiments, the Edman reagent is in solution.

[0167] In some embodiments, the Edman reagent comprises a nucleic acid. In some embodiments, a 5' end and a 3' end of the Edman reagent each comprise a molecular lock. In some embodiments, the molecular lock is adapted to inhibit the translocation of an end of the sequencing substrate through a nanopore. In some embodiments, the molecular lock is selected from an oligonucleotide capable of forming a hairpin structure; biotin; a biotin derivative; streptavidin; a streptavidin derivative; or a polypeptide derived from a fibronectin-binding protein,

[0168] In some embodiments, the Edman reagent comprises in a 5' to 3' order: (i) a 5' phosphorylated end, (ii) a 5' molecular lock, (iii) the reactive moiety, (iv) an end-point tick mark, (v) a cycle marker, (vi) a primer binding site, and (vii) a 3' molecular lock. In some embodiments, the Edman reagent comprises in a 5' to 3' order: (i) a 5' phosphorylated end, (ii) a 5' molecular lock, (iii) a cycle marker, (iv) a primer binding site, (v) the reactive moiety, (vi) an arresting moiety, and (7) a 3' molecular lock.

[0169] In some embodiments, the target polypeptide has a length greater than 5, 10, 15, 20, 25, 30, 50, 100, 200, 500, 5000, 30,000 consecutive amino acid residues or wherein the target polypeptide has a length in a range from 5 to 1000, 5 to 500, 5 to 200, or 5 to 100 consecutive amino acid residues.

[0170] In some embodiments, the substrate comprises a flow cell, or a bead.

[0171] Some embodiments also include a nanopore embedded in a membrane, wherein the membrane has a cis and trans surface. In some embodiments, the membrane comprises a lipid bilayer or block copolymer. In some embodiments, the nanopore comprises a protein nanopore. In some embodiments, the protein nanopore is selected from OmpF, OmpG, CsgG, MspA, a-HL, FhuA, AeL, FraC, Lys, (|>29p, ClyA, Ply AB; MspA or CsgG.

[0172] In some embodiments, the substrate comprises a bead; an initial Edman reagent comprises a cycle marker; and an additional Edman reagent comprising a cycle marker,wherein the cycle marker of the additional Edman reagent is different from the cycle marker of the initial cycle.

[0173] Some embodiments include a system for preparing a sequencing substrate, comprising: (a) immobilizing a target polypeptide immobilized on a bead; (b) an initial Edman reagent comprising a reactive moiety adapted to couple with the N-terminal residue of the target polypeptide, and a cycle marker; (d) a reagent adapted for cleaving the N-terminal residue from the target polypeptide; and (e) an additional Edman reagent comprising a cycle marker different from the cycle marker of the initial Edman reagent. In some embodiments, the barcode is linked to the C-terminal residue of the target polypeptide. In some embodiments, the barcode and cycle marker each comprise a polymer. In some embodiments, the polymer comprises a nucleic acid, a PEG-based moiety, or a polypeptide. In some embodiments, the target polypeptide or the Edman reagent, or both are in solution. In some embodiments, at least a portion of the barcode and of the cycle marker are adapted to hybridize to one another, and the system further comprises a polymerase. Some embodiments also include a reagent adapted to link the barcode and the cycle marker to one another.

[0174] Some embodiments include a system for characterizing a target polypeptide, comprising: (a) a sequencing substrate provided herein, or any one of the systems for preparing a sequencing substrate provided herein; (b) a membrane having a cis and trans surface; and (c) a nanopore embedded in the membrane. Some embodiments also include a nanopore substrate comprising a plurality of the nanopore embedded in the membrane. In some embodiments, the nanopore substrate comprises a flow cell. In some embodiments, the flow cell comprises a plurality of wells, wherein each well comprises a nanopore.

[0175] Some embodiments also include a detector adapted to measure a signal indicative of a coupled residue of the sequencing substrate in the nanopore. In some embodiments, the detector is adapted to measure a signal generated by a coupled residue of the sequencing substrate located in a read region of the nanopore. In some embodiments, the detector is adapted to measure a signal indicative of a cycle moiety of the sequencing substrate in the nanopore. Some embodiments also include a device adapted to apply over the membrane: (i) a potential difference, and / or (ii) an electroosmotic force. In some embodiments, the detector can measure a signal for a characteristic of a polymer of a sequencing substrate, such as a nucleotide residue or other subunit. In some embodiments, the a signal for a characteristicof a polymer of a sequencing substrate can be used to identify a target polypeptide or portion thereof, and / or a position of a coupled residue of the sequencing substrate within a target polypeptide.

[0176] The term “comprising” as used herein is synonymous with “including,” “containing,” or “characterized by,” and is inclusive or open-ended and does not exclude additional, unrecited elements or method steps.

[0177] The above description discloses several methods and materials of the present invention. This invention is susceptible to modifications in the methods and materials, as well as alterations in the fabrication methods and equipment. Such modifications will become apparent to those skilled in the art from a consideration of this disclosure or practice of the invention disclosed herein. Consequently, it is not intended that this invention be limited to the specific embodiments disclosed herein, but that it cover all modifications and alternati ves coming within the true scope and spirit of the invention.

[0178] All references cited herein, including but not limited to published and unpublished applications, patents, and literature references, are incorporated herein by reference in their entirety and are hereby made a part of this specification. To the extent publications and patents or patent applications incorporated by reference contradict the disclosure contained in the specification, the specification is intended to supersede and / or take precedence over any such contradictory material.

Claims

WHAT IS CLAIMED IS:

1. A method for preparing a sequencing substrate, comprising:(a) immobilizing a target polypeptide on a substrate;(b) coupling an N-terminal residue of the target polypeptide to a reactive moiety of an Edman reagent; and(c) cleaving the N-terminal residue from the target polypeptide, thereby obtaining the sequencing substrate comprising the Edman reagent and the coupled residue.

2. The method of claim 1, wherein the target polypeptide is immobilized on the substrate via a C-terminal residue of the target polypeptide.

3. The method of claim 1 or 2, wherein the target polypeptide is covalently immobilized on the substrate,4. The method of any one of claims 1-3, wherein the target polypeptide is immobilized on the substrate via a cleavable linker.

5. The method of any one of claims 1-4, wherein the reactive moiety comprises an isothiocyanate moiety or a phenyl isothiocyanate (PITC) moiety.

6. The method of any one of claims 1-5, wherein the Edman reagent comprises a net negative charge.

7. The method of any one of claims 1-6, wherein the Edman reagent comprises a polymer.

8. The method of claim 7, wherein the polymer is a homopolymer.

9. The method of claim 7 or 8, wherein the polymer comprises a subunit selected from a nucleotide, an amino acid, a polyethylene glycol (PEG), a phosphonamidite, or a peptide nucleic acid (PNA) subunit.

10. The method of any one of claims 7-9, wherein the polymer comprises a plurality of arginine residues.

11. The method of any one of claims 1-10, wherein the Edman reagent comprises a cycle marker indicative of a position of a coupled residue in the target polypeptide.

12. The method of claim 12, wherein the cycle marker is selected from a non-nucleosidic reporter moiety, a peptide reporter moiety, PEG-based reporter moiety, or a single¬ stranded DNA reporter moiety.

13. The method of any one of claims 1-12, wherein the Edman reagent comprises an arresting moiety adapted to pause or inhibit translocation of the sequencing substrate through a nanopore.

14. The method of any one of claims 1-13, wherein step (c) comprises contacting the N-terminal residue with an Edmanase or with conditions sufficient to cleave the Edman reagent and the coupled residue from the target polypeptide.

15. The method of claim 14, wherein the Edmanase is immobilized on the substrate.

16. The method of any one of claims 1-15, wherein the Edman reagent is immobilized on the substrate via a cleavable linker.

17. The method of claim 16, wherein the clea vable linker comprises a cysteine and selenocysteine group modified with a protecting group.

18. The method of claim 16 or 17, wherein the reactive moiety is located in the Edman reagent between arresting moieties.

19. The method of any one of claims 16-18, further comprising repeating steps (b) and (c) with an additional Edman reagent, wherein the immobilized Edman reagent and the additional Edman reagent each comprise an adapter adapted to conjugate to one another.

20. The method of claim 19, further comprising repeating steps (b) and (c) for a plurality of cycles, wherein each additional Edman reagent comprises a different cycle marker.

21. The method of claim 19 or 20, further comprising conjugating the immobilized Edman reagent to the additional Edman reagent.

22. The method of any one of claims 19-21, wherein each adapter comprises a nucleic acid.

23. The method of claim 22, wherein the adapters of the immobilized Edman reagent and of the additional Edman reagent are complementary to one another.

24. The method of any one of claims 21-23, wherein the conjugating comprises a polymerase reaction or a ligase reaction.

25. The method of any one of claims 16-24, further comprising (d) cleaving the immobilized sequencing substrate from the substrate.

26. The method of any one of claims 1-15, wherein the Edman reagent is in solution.

27. The method of claim 26, wherein the Edman reagent comprises a nucleic acid.

28. The method of claim 27, wherein a 5' end and a 3' end of the Edman reagent each comprise a molecular lock.

29. The method of claim 28, wherein the molecular lock is adapted to inhibit the translocation of an end of the sequencing substrate through a nanopore.

30. The method of claim 28 or 29, wherein the molecular lock is selected from an oligonucleotide capable of forming a hairpin structure; biotin; a biotin derivative; streptavidin; a streptavidin derivative; or a polypeptide derived from a fibronectin-binding protein.

31. The method of any one of claims 26-30, wherein the Edman reagent comprises in a 5' to 3' order: (i) a 5' phosphorylated end, (li) a 5' molecular lock, (iii) the reactive moiety, (iv) an end-point tick mark, (v) a cycle marker, (vi) a primer binding site, and (vii) a 3' molecular lock.

32. The method of any one of claims 26-30, wherein the Edman reagent comprises in a 5' to 3' order: (i) a 5' phosphorylated end, (ii) a 5' molecular lock, (iii) a cycle marker, (iv) a primer binding site, (v) the reactive moiety, (vi) an arresting moiety, and (7) a 3' molecular lock.

33. The method of any one of claims 1-32, wherein the target polypeptide has a length greater than 5, 10, 15, 20, 25, 30, 50, 100, 200, 500, 5000, 30,000 consecutive amino acid residues or wherein the target polypeptide has a length in a range from 5 to 1000, 5 to 500, 5 to 200, or 5 to 100 consecutive amino acid residues.

34. The method of any one of claims 1-33, wherein the substrate comprises a bead or a flow cell.

35. The method of any one of claims 1-34, wherein the substrate comprises a plurality of beads, wherein each bead comprises a target polypeptide different from one another.

36. The method of any one of claims 1-34, wherein substrate comprises a plurality of clusters of target polypeptides, wherein each cluster comprises the same target polypeptide and the clusters comprise a different target polypeptide from one another.

37. The method of any one of claims 1-35, wherein:the substrate comprises a bead;step (b) comprises coupling the N-terminal residue of the target polypeptide to the reactive moiety of an initial Edman reagent, wherein the initial Edman reagent comprises a cycle marker, and step (b) further comprises linking to the bead the initial Edman reagent coupled to the N-terminal residue;step (c) further comprises obtaining a modified initial Edman reagent coupled to the N-terminal residue and linked to the bead; andfurther comprising repeating steps (b) to (c) one or more times with an additional Edman reagent comprising a cycle marker, wherein the cycle marker of the additional Edman reagent is different from the cycle marker of a prior cycle.

38. A method for preparing a sequencing substrate on a bead, comprising:(a) immobilizing a target polypeptide on a bead;(b) coupling the N-terminal residue of the target polypeptide to a reactive moiety of an initial Edman reagent, wherein the initial Edman reagent comprises a cycle marker; and(c) linking the initial Edman reagent to the bead;(d) cleaving the N-terminal residue from the target polypeptide, thereby obtaining a modified initial Edman reagent coupled to the N-terminal residue and linked to the bead;(e) repeating steps (b) to (d) with an additional Edman reagent comprising a cycle marker, wherein the cycle marker of the additional Edman reagent is different from the cycle marker of a prior cycle to prepare the sequencing substrate on a bead.

39. A method for preparing a sequencing substrate, comprising:(a) obtaining a target polypeptide comprising a barcode, wherein the barcode is indicative of the target polypeptide;(b) coupling an N-terminal residue of the target polypeptide to a reactive moiety of an initial Edman reagent, wherein the initial Edman reagent comprises a cycle marker;(c) extending the cycle marker to obtain an extended cycle marker by (i) hybridizing the cycle marker to the barcode and extending the cycle marker to obtainthe extended cycle marker, or (ii) linking the cycle marker to the barcode to obtain the extended cycle marker;(d) cleaving the N-terminal residue from the target polypeptide, thereby obtaining a modified initial Edman reagent coupled to the N-terminal residue and comprising the extended cycle marker; and(e) repeating steps (b) to (d) with an additional Edman reagent comprising a cycle marker, wherein the cycle marker of the additional Edman reagent is different from the cycle marker of a prior cycle, and wherein the linking comprises linking the cycle marker of the additional cycle marker to the cycle marker of the prior cycle to prepare the sequencing substrate.

40. The method of claim 38 or 39, wherein the barcode is linked to the C-terminal residue of the target polypeptide,41. The method of any one of claims 38-40, wherein the barcode and cycle marker each comprise a polymer,42. The method of claim 41, wherein the polymer comprises a nucleic acid, a PEG-based moiety or a polypeptide.

43. The method of any one of claims 38-42, wherein the target polypeptide or the Edman reagent, or both are in solution.

44. The method of any one of claims 39-43, wherein step (c) comprises extending the cycle marker by polymerase extension.

45. The method of any one of claims 39-43, wherein step (c) comprises linking the cycle marker by ligation.

46. A sequencing substrate prepared according to the method of any one of claims 1-45.

47. A method for characterizing a target polypeptide, comprising:(i) obtaining the sequencing substrate of claim 46;(ii) obtaining a nanopore embedded in a membrane, wherein the membrane has a cis and trans surface;(iii) translocating the sequencing substrate through the nanopore; and(iv) measuring a signal indicative of a coupled residue of the sequencing substrate in the nanopore.

48. The method of claim 47, wherein step (iii) comprises applying over the membrane: a potential difference, and / or an electroosmotic force.

49. The method of claim 47 or 48, wherein step (iii) further comprises cycling the sequencing substrate located in the nanopore in a repeated movement towards the trans surface and to the cis surface.

50. The method of claim 49, wherein step (iv) comprises measuring a repeated signal.

51. The method of any one of claims 47-50, wherein step (iv) comprises measuring a signal generated by a coupled residue of the sequencing substrate located in a read region of the nanopore.

52. The method of any one of claims 47-51, further comprising measuring a signal indicative of a cycle moiety of the sequencing substrate in the nanopore.

53. The method of any one of claims 47-52, further comprising identifying the coupled residue of the sequencing substrate in the nanopore.

54. The method of any one of claims 47-53, further comprising repeating steps (iii) and (iv), wherein the translocating comprises moving the sequencing substrate through the nanopore such that an arresting moiety pauses the translocating and the signal is measured.

55. The method of claim 54, further comprising moving the arresting moiety through the nanopore.

56. The method of claim 55, further comprising applying a pulsed potential difference over the membrane.

57. The method of any one of claims 47-56, further comprising repeating steps (i) to (iv).

58. The method of claim 57, further comprising repeating steps (i) to (iv) for a plurality of cycles.

59. The method of claim 57 or 58, wherein the cycle marker of the sequence substrate is different in each cycle.

60. The method of any one of claims 47-59, wherein the translocating comprises contacting the sequencing substrate with a polymerase.

61. The method of any one of claims 47-60, wherein the membrane comprises a lipid bilayer or block copolymer.

62. The method of any one of claims 47-61, wherein the nanopore comprises a protein nanopore.

63. The method of claim 62, wherein the protein nanopore is selected from OmpF, OmpG, CsgG, MspA, a-HL, FhuA, AeL, FraC, Lys, (|)29p, ClyA, Ply AB; MspA or CsgG.

64. The method of any one of claims 47-63, wherein step (ii) comprises obtaining a plurality of the nanopores embedded in the membrane, wherein a nanopore substrate comprises the plurality of nanopores65. The method of claim 64, wherein step (iii) comprises translocating a plurality of sequencing substrates through the plurality of nanopores.

66. The method of claim 65, wherein step (iv) comprises measuring a plurality of signals indicative of a coupled residue of a sequencing substrate in each nanopore of the nanopores,67. The method of any one of claims 64-66, wherein the nanopore substrate comprises a flow cell.

68. The method of claim 67, wherein the flow cell comprises a plurality of wells, wherein each well comprises a nanopore.

69. A system for preparing a sequencing substrate, comprising:(a) a target polypeptide immobilized on a substrate;(b) an Edman reagent comprising a reactive moiety adapted to couple with an N-terminal residue of the target polypeptide; and(c) a reagent adapted to cleave the N-terminal residue from the target polypeptide, thereby obtaining the sequencing substrate comprising the Edman reagent and the coupled residue.

70. The system of claim 69, wherein the target polypeptide is immobilized on the substrate via a C-terminal residue of the target polypeptide.

71. The system of claim 69 or 70, wherein the reactive moiety comprises an isothiocyanate moiety, or a phenyl isothiocyanate (PITC) moiety.

72. The system of any one of claims 69-71, wherein the Edman reagent comprises a polymer comprising a subunit selected from a nucleotide, an amino acid, a polyethylene glycol (PEG), a phosphonamidite, or a peptide nucleic acid (PNA) subunit.

73. The system of any one of claims 69-72, wherein the Edman reagent comprises a cycle marker indicative of a position of a coupled residue in the target polypeptide.

74. The system of claim 73, wherein the cycle marker is selected from a non-nucleosidic reporter moiety, a peptide reporter moiety, PEG-based reporter moiety, or a single¬ stranded DNA reporter moiety.

75. The system of any one of claims 69-74, wherein the Edman reagent comprises one or more arresting moieties adapted to pause or inhibit translocation of the sequencing substrate through a nanopore.

76. The system of claim 75, wherein the reactive moiety is located in the Edman reagent between arresting moieties.

77. The system of any one of claims 69-76, wherein the reagent adapted to cleave the N-terminal residue from the target polypeptide comprises an Edmanase, or a reagent or conditions sufficient to cleave the Edman reagent and the coupled residue from the target polypeptide.

78. The system of any one of claims 69-77, wherein the Edman reagent is immobilized on the substrate via a cleavable linker.

79. The system of claim 78, further comprising an additional Edman reagent.

80. The system of claim 79, wherein the additional Edman reagent is adapted to conjugate with an initial Edman reagent.

81. The system of claim 79 or 80, wherein the additional Edman reagent comprises a cycle marker different from the initial cycle marker.

82. The system of any one of claims 69-81, wherein the Edman reagent comprises a nucleic acid.

83. The system of claim 82, wherein a 5' end and a 3' end of the Edman reagent each comprise a molecular lock, wherein the molecular lock is adapted to inhibit the translocation of an end of the sequencing substrate through a nanopore.

84. The system of claim 83, wherein the molecular lock is selected from an oligonucleotide capable of forming a hairpin structure; biotin; a biotin derivative; streptavidin; a streptavidin derivative; or a polypeptide derived from a fibronectin-binding protein.

85. The system of any one of claims 69-84, wherein the Edman reagent comprises in a 5' to 3' order:(i) a 5' phosphorylated end, (ii) a 5' molecular lock, (iii) the reactive moiety, (iv) an end-point tick mark, (v) a cycle marker, (vi) a primer binding site, and (vii) a 3' molecular lock; or(i) a 5' phosphorylated end, (ii) a 5' molecular lock, (iii) a cycle marker, (iv) a primer binding site, (v) the reactive moiety, (vi) an arresting moiety, and (7) a 3' molecular lock.

86. The system of any one of claims 69-85, wherein the target polypeptide has a length greater than 5, 10, 15, 20, 25, 30, 50, 100, 200, 500, 5000, 30,000 consecutive amino acid residues or wherein the target polypeptide has a length in a range from 5 to 1000, 5 to 500, 5 to 200, or 5 to 100 consecutive amino acid residues.

87. The system of any one of claims 69-83, wherein the substrate comprises a bead or a flow cell.

88. The system of claim 87, wherein the substrate comprises a plurality of beads, wherein each bead comprises a target polypeptide different from one another,89. The system of claim 87, wherein the substrate comprises a plurality of clusters of target polypeptides, wherein each cluster comprises the same target polypeptide and the clusters comprise a different target polypeptide from one another.

90. The system of any one of claims 69-87, wherein:the substrate comprises a bead;an initial Edman reagent comprises a cycle marker, andan additional Edman reagent comprising a cycle marker, wherein the cycle marker of the additional Edman reagent is different from the cycle marker of the initial cycle.

91. A system for preparing a sequencing substrate, comprising:(a) a target polypeptide immobilized on a bead, wherein the target moiety comprises a barcode indicative of the target polypeptide;(b) an initial Edman reagent comprising a reactive moiety adapted to couple with the N-terminal residue of the target polypeptide, and a cycle marker; and(c) a reagent adapted for cleaving the N-terminal residue from the target polypeptide; and(d) an additional Edman reagent comprising a cycle marker different from the cycle marker of the initial Edman reagent.

92. The system of claim 91, wherein the barcode is linked to the C-terminal residue of the target polypeptide.

93. The system of claim 91 or 92, wherein the barcode and cycle marker each comprise a polymer.

94. The system of claim 93, wherein the polymer comprises a nucleic acid, a PEG-based moiety, or a polypeptide.

95. The system of any one of claims 90-94, wherein the target polypeptide or the Edman reagent, or both are in solution.

96. The system of any one of claims 90-95, wherein at least a portion of the barcode and of the cycle marker are adapted to hybridize to one another, and the system further comprises a polymerase,97. The system of any one of claims 90-96, further comprising a reagent adapted to link the barcode and the cycle marker to one another.

98. A system for characterizing a target polypeptide, comprising:(a) the sequencing substrate of claim 46 or the system of any one of claims 69- 97;(b) a membrane having a cis and trans surface; and(c) a nanopore embedded in the membrane.

99. The system of claim 98, further comprising a nanopore substrate comprising a plurality of nanopores embedded in the membrane.

100. The system of claim 99, wherein the nanopore substrate comprises a flow cell.

101. The system of claim 100, wherein the flow cell comprises a plurality of wells, wherein each well comprises a nanopore.

Citation Information

Patent Citations

  • Methods and kits using nucleic acid encoding and / or label

    US20210355483A1

  • Sequencing polynucleotides using nanopores

    US20230090867A1

  • Nanopore devices including barriers using diblock or triblock copolymers, and methods of making the same

    US20230312856A1

  • Nanopore sequencing

    US20250215487A1

  • Concealed door positioning mechanism

    US5075923A