method

By employing proteinases to selectively cleave polypeptides and attach sequencing adapters, the method facilitates efficient nanopore characterization of oligopeptides, addressing the inefficiencies of existing techniques and enabling accurate analysis of polypeptides.

WO2026052817A1PCT designated stage Publication Date: 2026-03-12OXFORD NANOPORE TECH LTD
View PDF 54 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2025-09-05
Publication Date
2026-03-12

AI Technical Summary

Technical Problem

Existing methods for characterizing polypeptides, such as mass spectrometry and Edman degradation, are inefficient and unsuitable for single molecule analysis, particularly for long polypeptides with secondary and tertiary structures, and random shearing methods produce oligopeptides with disparate properties that are challenging to process for nanopore characterization.

Method used

A method involving the use of proteinases to selectively cleave polypeptides at specific sites, generating oligopeptides with consistent terminal chemistry, which are then attached to sequencing adapters for characterization using a nanopore.

Benefits of technology

Enables efficient and reliable characterization of oligopeptides by nanopore sensing, providing accurate measurements of identity, length, composition, and modifications, overcoming the limitations of existing techniques.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure IMGF000002_0001
    Figure IMGF000002_0001
  • Figure IMGF000002_0002
    Figure IMGF000002_0002
  • Figure IMGF000009_0001
    Figure IMGF000009_0001
Patent Text Reader

Abstract

Provided herein is a method of preparing an oligopeptide-adapter adduct for characterisation of said oligopeptide using a nanopore. Also provided herein are related systems, kits and analysis methods.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] METHOD

[0002] Cross reference to related

[0003] This application claims priority from United Kingdom Patent Application No.

[0004] 2413130.2 filed on 6 September 2024, the entire contents of which are hereby incorporated by reference.

[0005] Field

[0006] The present disclosure relates to methods of preparing an oligopeptide-adapter adduct for characterisation of said oligopeptide using a nanopore. Also provided are related systems, kits and analysis methods.

[0007] The characterisation of biological molecules is of increasing importance in biomedical and biotechnological applications. For example, sequencing of nucleic acids allows the study of genomes and the proteins they encode and, for example, allows correlation between nucleic acid mutations and observable phenomena such as disease indications. Nucleic acid sequencing can be used in evolutionary biology to study the relationship between organisms. Metagenomics involves identifying organisms present in samples, for example microbes in a microbiome, with nucleic acid sequencing allowing the identification of such organisms.

[0008] Whilst techniques to characterise (e.g. sequence) polynucleotides have been extensively developed, techniques to characterise polypeptides are less advanced, despite being of very significant biotechnological importance. For example, knowledge of a protein sequence can allow structure-activity relationships to be established and has implications in rational drug development strategies for developing ligands for specific receptors. Identification of post-translational modifications is also key to understanding the functional properties of many proteins. For example, typically 30-50% of protein species are phosphorylated in eukaryotes. Some proteins may have multiple phosphorylation sites, serving to activate or inactivate a protein, promote its degradation, or modulate interactions with protein partners.

[0009] Known methods of characterising polypeptides include mass spectrometry and Edman degradation.

[0010] Protein mass spectrometry involves characterising whole proteins or fragments thereof in an ionised form. Known methods of protein mass spectrometry include electrospray ionisation (ESI) and matrix-assisted laser desorption / ionisation (MALDI). Mass spectrometry has some benefits, but results obtained can be affected by the presence of contaminants and it can be difficult to process fragile molecules without their fragmentation. Moreover, mass spectrometry is not a single molecule technique and provides only bulk information about the sample interrogated. Mass spectrometry is unsuitable for characterising differences within a population of polypeptide samples and is unwieldy when seeking to distinguish neighbouring residues.

[0011] Edman degradation is an alternative to mass spectrometry which allows the residue- by-residue sequencing of polypeptides. Edman degradation sequences polypeptides by sequentially cleaving the N-terminal amino acid and then characterising the individually cleaved residues using chromatography or electrophoresis. However, Edman sequencing is slow, involves the use of costly reagents, and like mass spectrometry is not a single molecule technique.

[0012] As such, there remains a pressing need for new techniques to characterise polypeptides, especially at the single molecule level. Single molecule techniques for characterising biomolecules such as polynucleotides have proven to be particularly attractive due to their high fidelity and avoidance of amplification bias.

[0013] One attractive method of single molecule characterization of biomolecules such as polypeptides is nanopore sensing. Nanopore sensing is an approach to analyte detection and characterization that relies on the observation of individual binding or interaction events between the analyte molecules and an ion conducting channel. Nanopore sensors can be created by placing a single pore of nanometre dimensions in an electrically insulating membrane and measuring voltage-driven ion currents through the pore in the presence of analyte molecules. The presence of an analyte inside or near the nanopore will alter the ionic flow through the pore, resulting in altered ionic or electric currents being measured over the channel. The identity of an analyte is revealed through its distinctive current signature, notably the duration and extent of current blocks and the variance of current levels during its interaction time with the pore. Nanopore sensing has the potential to allow rapid and cheap polypeptide characterisation.

[0014] Nanopore sensing and characterisation of polypeptides has been proposed in the art. For example, WO 2013 / 123379 discloses the use of an NTP-driven protein processing unfoldase enzyme to process a protein to be translocated through a nanopore. WO 2021 / 111125 discloses methods in which a target polypeptide is conjugated to a polynucleotide to form a single-stranded polypeptide-polynucleotide conjugate, with the conjugate being moved through a nanopore using a polynucleotide-handling protein. WO 2021 / 133168 discloses protein and polypeptide fingerprinting and sequencing by nanopore translocation of polypeptide-oligonucleotide complexes. PCT / GB2023 / 052838 discloses methods of characterising a target polypeptide as it moves with relation to a nanopore. Each of these documents is incorporated by reference in their entireties.

[0015] These methods have provided useful techniques for characterising polypeptides using nanopores. However, some challenges remain. In particular, long polypeptides can have significant secondary and tertiary structure which can hamper characterisation methods. Random shearing of long polypeptides (e.g. by exposing the peptides to chemical reagents, e.g. by exposing the polypeptide to a change in pH, or to a chemical reagent) can generate random shorter oligopeptides which avoid or reduce such structures. However, it can be challenging to process such oligopeptides in appropriate manner for nanopore characterisation due at least in part to disparate properties of the resultant oligopeptides. Accordingly, there remains a need for further methods.

[0016] The inventors have recognised that proteinases can be used to process polypeptides and, if selective for a cleavage site in the polypeptide, will generate oligopeptide fragments which have consistent terminal chemistry suitable for attachment to a sequencing adapter to facilitate characterisation of the oligopeptide using a nanopore.

[0017] Accordingly, the disclosure relates to methods of preparing an oligopeptide-adapter adduct. The oligopeptide-adapter adduct comprises an oligopeptide and a sequencing adapter, which may be a polynucleotide or a polypeptide sequencing adapter. The method involves contacting a polypeptide (e.g. a protein) with a proteinase which is selective for a cleavage site in the polypeptide. Suitable proteinases are described in more detail herein. Contacting the polypeptide with the proteinase results in the proteinase selectively cleaving the polypeptide at the cleavage site, thereby forming one or more oligopeptides. Each oligopeptide thereby formed has a first end and a second end. The first end terminates in an amino acid motif comprising a first reactive amino acid and the method involves selectively attaching to each first reactive amino acid a first sequencing adapter. Accordingly the method results in generating an oligopeptide-adapter adduct from each oligopeptide. The oligopeptide-adapter adduct can be contacted with a nanopore for characterisation of at least the oligopeptide portion of the oligopeptide-adapter adduct. For example, the oligopeptide-adapter adduct can be contacted with a nanopore under conditions such that the oligopeptide moves with respect to the nanopore, and one or more measurements characteristic of the oligopeptide can be taken. The methods may allow properties of the oligopeptide to be determined, such as its identity, length, composition, amino acid sequence and / or whether and how the polypeptide is modified.

[0018] Accordingly, provided herein is a method of preparing an oligopeptide-adapter adduct for characterisation of said oligopeptide using a nanopore, the method comprising i) contacting a polypeptide comprising said oligopeptide with a proteinase selective for a cleavage site, under conditions such that the proteinase selectively cleaves the polypeptide thereby forming one or more oligopeptides, wherein each oligopeptide comprises a first end terminating in a first amino acid motif comprising a first reactive amino acid, and a second end terminating in a second amino acid; and ii) selectively attaching to each first reactive amino acid a first sequencing adapter, thereby forming an oligopeptide-adapter adduct.

[0019] In some embodiments the proteinase selectively cleaves the polypeptide at each occurrence of said cleavage site in the polypeptide.

[0020] In some embodiments, for each oligopeptide formed in step (i), the first end is the C -terminus of the oligopeptide and the second end is the N-terminus of the oligopeptide. In some embodiments the second end comprises a N-terminal amine group. In some embodiments step (ii) comprises reacting the N-terminal amine group with an aminereactive moiety.

[0021] In some embodiments, step (ii) comprises reacting the oligopeptide with the first sequencing adapter under conditions such that the first sequencing adapter reacts with the first reactive amino acid. In some embodiments, step (ii) comprises functionalising the first reactive amino acid by reacting the oligopeptide with a compound comprising a oligo- reactive functional group and an adapter-reactive functional group under conditions such that the oligo-reactive functional group reacts with the first reactive amino acid. In some embodiments, the method comprises reacting the first sequencing adapter with the adapter- reactive functional group.

[0022] In some embodiments the cleavage site comprises lysine or arginine.

[0023] In some embodiments the cleavage site, the first amino acid motif and the first reactive amino acid each comprise or consist of lysine. In some embodiments the proteinase is LysC or a or functional analog, fragment or variant thereof. In some embodiments the cleavage site, the first amino acid motif and the first reactive amino acid each comprise or consist of arginine. In some embodiments the proteinase is ArgC or a or functional analog, fragment or variant.

[0024] In some embodiments reacting the N-terminal amine group with an amine-reactive moiety comprises immobilizing the oligopeptide on a substrate. In some embodiments the substrate is an amine-reactive substrate comprising a reactive carbonyl group. In some embodiments the reactive carbonyl group is an aldehyde group. In some embodiments the amine-reactive substrate is functionalised with 2-PCA or an amine-binding derivative thereof.

[0025] In some embodiments the method comprises cleaving or eluting the oligopeptide from the substrate. In some embodiments said cleaving or eluting generates a second reactive functional group comprised in said oligopeptide capable of reacting with a complementary functional group on a second sequencing adapter.

[0026] In some embodiments reacting the N-terminal amine group with an amine-reactive moiety comprises reacting the oligopeptide with a nucleotide or peptide handle comprising a purification tag, under conditions such that the handle is attached to the N-terminus of the oligopeptide; thereby generating a fusion adduct. In some embodiments the method comprises contacting the oligopeptide with a ligase and a peptide handle comprising a purification tag, under conditions such that the ligase ligates the peptide handle to the N- terminus of the oligopeptide. In some embodiments the method comprises contacting the purification tag with a complementary substrate under conditions such that the fusion adduct is immobilised on the substrate.

[0027] In some embodiments the method comprises cleaving or eluting a portion of the fusion adduct comprising the oligopeptide from the substrate. In some embodiments the method comprises attaching a second sequencing adapter to the portion. In some embodiments the method comprises attaching a second sequencing adapter to the portion after the portion has been cleaved from the immobilised fusion adduct.

[0028] In some embodiments the handle comprises a second reactive functional group capable of reacting with a complementary functional group on the second sequencing adapter. In some embodiments the method comprises cleaving the portion from the immobilised fusion adduct thereby generating a second reactive functional group capable of reacting with a complementary functional group on the second sequencing adapter. In some embodiments a second sequencing adapter is comprised in the handle.

[0029] In some embodiments the oligopeptide is attached to the first sequencing adapter.

[0030] In some embodiments reacting the N-terminal amine group with an amine-reactive moiety comprises attaching a first sequencing adapter to the first reactive amino acid and to the N-terminal amine group. In some embodiments the method comprises cleaving the amino acid comprising the reacted N-terminal amine group from the oligopeptide-adapter adduct. In some embodiments cleaving the reacted N-terminal amine group from the oligopeptide-adapter adduct generates a second reactive functional group capable of reacting with a complementary functional group on a second sequencing adapter.

[0031] In some embodiments, the method comprises attaching the second sequencing adapter to the second reactive functional group. In some embodiments the second reactive functional group is an N-terminal amine group. In some embodiments the second reactive functional group is orthogonal to the optionally-functionalised first reactive amino acid.

[0032] In some embodiments the method comprises reacting one or more amine groups with one or more amine-reactive moieties comprising a 2-PCA group for reaction with said amine groups. In some embodiments reacting said one or more amine groups with said one or more amine-reactive moieties comprises reacting the one or more amine groups with the 2-PCA moieties and one or more maleimides.

[0033] In some embodiments attaching the first sequencing adapter and a second sequencing adapter to the oligopeptide generates an adduct of form:

[0034] S -N~O~C. F wherein F is the first sequencing adapter,W’O’Cis the oligopeptide wherein N- represents the N-terminal end of the oligopeptide and -C represents the C-terminal end of the oligopeptide, and S is the second sequencing adapter.

[0035] In some embodiments, the first sequencing adapter is a polynucleotide sequencing adapter. In some embodiments, the second sequencing adapter comprises a polynucleotide and / or a polypeptide. In some embodiments the second sequencing adapter is a polynucleotide sequencing adapter. In some embodiments the first sequencing adapter is different to the second sequencing adapter.

[0036] In some embodiments, the method comprises loading a motor protein onto the first sequencing adapter. In some embodiments the method comprises loading a motor protein onto the second sequencing adapter if present. In some embodiments, the oligopeptide comprises from about 2 to about 50 amino acids. In some embodiments the oligopeptide comprises from about 5 to about 30 amino acids.

[0037] Also provided is a method of characterising a polypeptide, comprising carrying out a method according to any one of the preceding claims, and contacting the oligopeptide with a nanopore under conditions such that the oligopeptide moves with respect to the nanopore; and

[0038] - taking one or more measurements characteristic of the oligopeptide as the oligopeptide moves with respect to the nanopore; thereby characterising the oligopeptide.

[0039] In some embodiments, the oligopeptide is comprised in an adduct of form:

[0040] S -N~O~C. F wherein F is the first sequencing adapter,W’O’Cis the oligopeptide wherein N- represents the N-terminal end of the oligopeptide and -C represents the C-terminal end of the oligopeptide, and S is a second sequencing adapter.

[0041] Also provided is a kit for modifying a polypeptide, comprising a first sequencing adapter capable of selectively reacting with an optionally- functionalised first reactive amino acid at the C-terminus of an oligopeptide generated by contacting a polypeptide with a proteinase; and a second sequencing adapter capable of reacting with a second reactive functional group at the N-terminus of said oligopeptide.

[0042] In some embodiments, the kit comprises one or more of: a proteinase capable of selectively cleaving a polypeptide at a cleavage site thereby forming one or more oligopeptides each comprising a first end terminating in a first amino acid motif comprising a first reactive amino acid, and a second end terminating in a second amino acid; a compound comprising an oligo-reactive functional group and an adapter-reactive functional group, wherein the compound is capable of reacting with the first reactive amino acid thereby functionalising the first reactive amino acid; and a peptide handle capable of being ligated onto the N-terminus of said oligopeptide; a ligase capable of ligating a peptide handle onto the N-terminus of said oligopeptide; a motor protein capable of controlling the movement of the oligopeptide with respect to a nanopore. Also provided is a computer program comprising instructions capable of execution by a computer system which are configured, on execution, to cause the computer system to carry out a method as described herein. Also provided is a computer storage medium storing a computer program as described herein. Also provided is a computer system configured to perform a method as described herein.

[0043] Figure 1 shows an exemplary non-limiting embodiment of the disclosed methods as described in more detail herein.

[0044] Figure 2 shows an exemplary non-limiting embodiment of the disclosed methods as described in more detail herein.

[0045] Figure 3 shows an exemplary non-limiting embodiment of the disclosed methods as described in more detail herein.

[0046] Figure 4 shows an exemplary non-limiting embodiment of the disclosed methods as described in more detail herein.

[0047] Figure 5 shows an exemplary non-limiting embodiment of the disclosed methods as described in more detail herein.

[0048] Figure 6 shows exemplary non-limiting examples of attachments between an oligopeptide and a sequencing adapter.

[0049] Figure 7 shows an exemplary non-limiting embodiment of the disclosed methods as described in more detail herein.

[0050] Figure 8 shows current traces showing current (pA) vs time (s) produced by example DNA-peptide conjugates.

[0051] Figure 9 shows current traces showing current (pA) vs time (s) produced by example DNA-peptide conjugates.

[0052] Figure 10 shows current traces showing current (pA) vs time (s) produced by example DNA-peptide conjugates.

[0053] Figure 11 shows an exemplary non-limiting embodiment of the disclosed methods as described in more detail herein.

[0054] Figure 12 shows examples of current traces showing current (pA) vs time (s) from Example 4.

[0055] Detailed The present invention will be described with respect to particular embodiments and with reference to certain drawings but the invention is not limited thereto but only by the claims. Any reference signs in the claims shall not be construed as limiting the scope. Of course, it is to be understood that not necessarily all aspects or advantages may be achieved in accordance with any particular embodiment of the invention. Thus, for example those skilled in the art will recognize that the invention may be embodied or carried out in a manner that achieves or optimizes one advantage or group of advantages as taught herein without necessarily achieving other aspects or advantages as may be taught or suggested herein.

[0056] The invention, both as to organization and method of operation, together with features and advantages thereof, may best be understood by reference to the following detailed description when read in conjunction with the accompanying drawings. The aspects and advantages of the invention will be apparent from and elucidated with reference to the embodiment(s) described hereinafter. Reference throughout this specification to "one embodiment" or "an embodiment" means that a particular feature, structure or characteristic described in connection with the embodiment is included in at least one embodiment of the present invention. Thus, appearances of the phrases "in one embodiment" or "in an embodiment" in various places throughout this specification are not necessarily all referring to the same embodiment, but may. Similarly, it should be appreciated that in the description of exemplary embodiments of the invention, various features of the invention are sometimes grouped together in a single embodiment, figure, or description thereof for the purpose of streamlining the disclosure and aiding in the understanding of one or more of the various inventive aspects. This method of disclosure, however, is not to be interpreted as reflecting an intention that the claimed invention requires more features than are expressly recited in each claim. Rather, as the following claims reflect, inventive aspects lie in less than all features of a single foregoing disclosed embodiment.

[0057] It should be appreciated that “embodiments” of the disclosure can be specifically combined together unless the context indicates otherwise. The specific combinations of all disclosed embodiments (unless implied otherwise by the context) are further disclosed embodiments of the claimed invention.

[0058] In addition as used in this specification and the appended claims, the singular forms “a”, “an”, and “the” include plural referents unless the content clearly dictates otherwise. Thus, for example, reference to “a polynucleotide” includes two or more polynucleotides, reference to “a motor protein” includes two or more such proteins, reference to “a helicase” includes two or more helicases, reference to “a monomer” refers to two or more monomers, reference to “a pore” includes two or more pores and the like.

[0059] All publications, patents and patent applications cited herein, whether supra or infra, are hereby incorporated by reference in their entirety.

[0060] Definitions

[0061] Where an indefinite or definite article is used when referring to a singular noun e.g. "a" or "an", "the", this includes a plural of that noun unless something else is specifically stated. Where the term "comprising" is used in the present description and claims, it does not exclude other elements or steps. Furthermore, the terms first, second, third and the like in the description and in the claims, are used for distinguishing between similar elements and not necessarily for describing a sequential or chronological order. It is to be understood that the terms so used are interchangeable under appropriate circumstances and that the embodiments of the invention described herein are capable of operation in other sequences than described or illustrated herein. The following terms or definitions are provided solely to aid in the understanding of the invention. Unless specifically defined herein, all terms used herein have the same meaning as they would to one skilled in the art of the present invention. Practitioners are particularly directed to Sambrook et al., Molecular Cloning: A Laboratory Manual, 4thed., Cold Spring Harbor Press, Plainsview, New York (2012); and Ausubel et al., Current Protocols in Molecular Biology (Supplement 114), John Wiley & Sons, New York (2016), for definitions and terms of the art. The definitions provided herein should not be construed to have a scope less than understood by a person of ordinary skill in the art.

[0062] "About" as used herein when referring to a measurable value such as an amount, a temporal duration, and the like, is meant to encompass variations of ± 20 % or ± 10 %, more preferably ± 5 %, even more preferably ± 1 %, and still more preferably ± 0.1 % from the specified value, as such variations are appropriate to perform the disclosed methods.

[0063] “Nucleotide sequence”, “DNA sequence” or “nucleic acid molecule(s)” as used herein refers to a polymeric form of nucleotides of any length, either ribonucleotides or deoxyribonucleotides. This term refers only to the primary structure of the molecule. Thus, this term includes double- and single-stranded DNA, and RNA. The term “nucleic acid” as used herein, is a single or double stranded covalently-linked sequence of nucleotides in which the 3' and 5' ends on each nucleotide are joined by phosphodiester bonds. The polynucleotide may be made up of deoxyribonucleotide bases or ribonucleotide bases. Nucleic acids may be manufactured synthetically in vitro or isolated from natural sources. Nucleic acids may further include modified DNA or RNA, for example DNA or RNA that has been methylated, or RNA that has been subject to post-transcriptional modification, for example 5 ’-capping with 7-methyl guanosine, 3 ’-processing such as cleavage and polyadenylation, and splicing. Nucleic acids may also include synthetic nucleic acids (XNA), such as hexitol nucleic acid (HNA), cyclohexene nucleic acid (CeNA), threose nucleic acid (TNA), glycerol nucleic acid (GNA), locked nucleic acid (LNA) and peptide nucleic acid (PNA). Sizes of nucleic acids, also referred to herein as “polynucleotides” are typically expressed as the number of base pairs (bp) for double stranded polynucleotides, or in the case of single stranded polynucleotides as the number of nucleotides (nt). One thousand bp or nt equal a kilobase (kb). Polynucleotides of less than around 40 nucleotides in length are typically called “oligonucleotides” and may comprise primers for use in manipulation of DNA such as via polymerase chain reaction (PCR).

[0064] The term “amino acid” in the context of the present disclosure is used in its broadest sense and is meant to include organic compounds containing amine (NH2) and carboxyl (COOH) functional groups, along with a side chain (e.g., a R group) specific to each amino acid. In some embodiments, the amino acids refer to naturally occurring L a- amino acids or residues. The commonly used one and three letter abbreviations for naturally occurring amino acids are used herein: A=Ala; C=Cys; D=Asp; E=Glu; F=Phe; G=Gly; H=His; I=Ile; K=Lys; L=Leu; M=Met; N=Asn; P=Pro; Q=Gln; R=Arg; S=Ser; T=Thr; V=Val; W=Trp; and Y=Tyr (Lehninger, A. L., (1975) Biochemistry, 2d ed., pp. 71-92, Worth Publishers, New York). The general term “amino acid” further includes D- amino acids, retro-inverso amino acids as well as chemically modified amino acids such as amino acid analogues, naturally occurring amino acids that are not usually incorporated into proteins such as norleucine, and chemically synthesised compounds having properties known in the art to be characteristic of an amino acid, such as P-amino acids. For example, analogues or mimetics of phenylalanine or proline, which allow the same conformational restriction of the peptide compounds as do natural Phe or Pro, are included within the definition of amino acid. Such analogues and mimetics are referred to herein as "functional equivalents" of the respective amino acid. Other examples of amino acids are listed by Roberts and Vellaccio, The Peptides: Analysis, Synthesis, Biology, Gross and Meiehofer, eds., Vol. 5 p. 341, Academic Press, Inc., N.Y. 1983, which is incorporated herein by reference.

[0065] The terms “polypeptide”, and “peptide” are interchangeably used herein to refer to a polymer of amino acid residues and to variants and synthetic analogues of the same. Thus, these terms apply to amino acid polymers in which one or more amino acid residues is a synthetic non-naturally occurring amino acid, such as a chemical analogue of a corresponding naturally occurring amino acid, as well as to naturally-occurring amino acid polymers. Polypeptides can also undergo maturation or post-translational modification processes that may include, but are not limited to: glycosylation, proteolytic cleavage, lipidization, signal peptide cleavage, propeptide cleavage, phosphorylation, and such like. A peptide can be made using recombinant techniques, e.g., through the expression of a recombinant or synthetic polynucleotide. A recombinantly produced peptide is typically substantially free of culture medium, e.g., culture medium represents less than about 20 %, more typically less than about 10 %, and most typically less than about 5 % of the volume of the protein preparation.

[0066] The term “protein” is used to describe a folded polypeptide having a secondary, tertiary, or quaternary structure. The protein may be composed of a single polypeptide, or may comprise multiple polypeptides that are assembled to form a multimer. The multimer may be a homooligomer, or a heterooligmer. The protein may be a naturally occurring, or wild type protein, or a modified, or non-naturally, occurring protein. The protein may, for example, differ from a wild type protein by the addition, substitution or deletion of one or more amino acids.

[0067] A “variant” of a protein encompass peptides, oligopeptides, polypeptides, proteins and enzymes having amino acid substitutions, deletions and / or insertions relative to the unmodified or wild-type protein in question and having similar biological and functional activity as the unmodified protein from which they are derived. The term "amino acid identity" as used herein refers to the extent that sequences are identical on an amino acid- by-amino acid basis over a window of comparison. Thus, a "percentage of sequence identity" is calculated by comparing two optimally aligned sequences over the window of comparison, determining the number of positions at which the identical amino acid residue (e.g., Ala, Pro, Ser, Thr, Gly, Vai, Leu, He, Phe, Tyr, Trp, Lys, Arg, His, Asp, Glu, Asn, Gin, Cys and Met) occurs in both sequences to yield the number of matched positions, dividing the number of matched positions by the total number of positions in the window of comparison (i.e., the window size), and multiplying the result by 100 to yield the percentage of sequence identity.

[0068] For all aspects and embodiments of the present invention, a “variant” has at least 40%, 50%, 60%, 70%, 80%, 90%, 95% or 99% complete sequence identity to the amino acid sequence of the corresponding wild-type protein. Sequence identity can also be to a fragment or portion of the full length polynucleotide or polypeptide. Hence, a sequence may have only 50 % overall sequence identity with a full length reference sequence, but a sequence of a particular region, domain or subunit could share 80 %, 90 %, or as much as 99 % sequence identity with the reference sequence.

[0069] The term “wild-type” refers to a gene or gene product isolated from a naturally occurring source. A wild-type gene is that which is most frequently observed in a population and is thus arbitrarily designed the “normal” or “wild-type” form of the gene. In contrast, the term “modified”, “mutant” or “variant” refers to a gene or gene product that displays modifications in sequence (e.g., substitutions, truncations, or insertions), post- translational modifications and / or functional properties (e.g., altered characteristics) when compared to the wild-type gene or gene product. It is noted that naturally occurring mutants can be isolated; these are identified by the fact that they have altered characteristics when compared to the wild-type gene or gene product. Methods for introducing or substituting naturally-occurring amino acids are well known in the art. For instance, methionine (M) may be substituted with arginine (R) by replacing the codon for methionine (ATG) with a codon for arginine (CGT) at the relevant position in a polynucleotide encoding the mutant monomer. Methods for introducing or substituting non-naturally-occurring amino acids are also well known in the art. For instance, non- naturally-occurring amino acids may be introduced by including synthetic aminoacyl- tRNAs in the IVTT system used to express the mutant monomer. Alternatively, they may be introduced by expressing the mutant monomer in E. coli that are auxotrophic for specific amino acids in the presence of synthetic (i.e. non-naturally-occurring) analogues of those specific amino acids. They may also be produced by native chemical ligation if the mutant monomer is produced using partial peptide synthesis. Conservative substitutions replace amino acids with other amino acids of similar chemical structure, similar chemical properties or similar side-chain volume. The amino acids introduced may have similar polarity, hydrophilicity, hydrophobicity, basicity, acidity, neutrality or charge to the amino acids they replace. Alternatively, the conservative substitution may introduce another amino acid that is aromatic or aliphatic in the place of a pre-existing aromatic or aliphatic amino acid. Conservative amino acid changes are well-known in the art and may be selected in accordance with the properties of the 20 main amino acids as defined in Table 1 below. Where amino acids have similar polarity, this can also be determined by reference to the hydropathy scale for amino acid side chains in Table 2.

[0070] Table 1 - Chemical properties of amino acids

[0071] Table 2 - Hydropathy scale

[0072] Side Chain Hydropathy

[0073] He 4.5 Vai 4.2 Leu 3.8 Phe 2.8 Cys 2.5 Met 1.9 Ala 1.8 Gly -0.4 Thr -0.7 Ser -0.8 Trp -0.9 Tyr -1.3 Pro -1.6 His -3.2 Glu -3.5 Gin -3.5 Asp -3.5 Asn -3.5

[0074] Lys -3.9 Arg -4.5 A mutant or modified protein, monomer or peptide can also be chemically modified in any way and at any site. A mutant or modified monomer or peptide may be chemically modified by attachment of a molecule to one or more cysteines (cysteine linkage), attachment of a molecule to one or more lysines, attachment of a molecule to one or more non-natural amino acids, enzyme modification of an epitope or modification of a terminus. Suitable methods for carrying out such modifications are well-known in the art. The mutant of modified protein, monomer or peptide may be chemically modified by the attachment of any molecule. For instance, the mutant of modified protein, monomer or peptide may be chemically modified by attachment of a dye or a fluorophore.

[0075] Disclosed Methods

[0076] Provided herein is method of preparing an oligopeptide-adapter adduct for characterisation of said oligopeptide using a nanopore, the method comprising i) contacting a polypeptide comprising said oligopeptide with a proteinase selective for a cleavage site, under conditions such that the proteinase selectively cleaves the polypeptide thereby forming one or more oligopeptides, wherein each oligopeptide comprises a first end terminating in a first amino acid motif comprising a first reactive amino acid, and a second end terminating in a second amino acid; and ii) selectively attaching to each first reactive amino acid a first sequencing adapter, thereby forming an oligopeptide-adapter adduct.

[0077] The disclosed methods take advantage of the fact that a polypeptide (e.g. a protein) can be selectively fragmented into oligopeptides using proteinases. A proteinase which is selective for a recognition site in a polypeptide will selectively fragment the polypeptide at each occurrence of the recognition site. For long polypeptides (e.g. as described herein in more detail), the statistical frequency of a given recognition site can be determined. Accordingly, by selection of an appropriate proteinase, fragmentation of a polypeptide results in oligopeptides having a statistically determined length, rather than being of random length.

[0078] A further advantage is that proteinases can be selected or designed to cleave polypeptides thereby leaving an amino acid motif at one end of the resulting oligopeptides which can comprise a reactive amino acid. This is described in more detail herein. For example, in some embodiments the proteinase cleaves the polypeptide at a reactive amino acid and therefore generates oligopeptide fragments which comprise the reactive amino acid at a terminus, such as at the C-terminus. Because the proteinase is selective for the cleavage site, each oligopeptide fragment generated by proteinase activity on the polypeptide has a common terminal amino acid motif, which can comprise a common reactive amino acid. The reactive amino acid can be exploited for attachment of a sequencing adapter such as a polynucleotide or polypeptide sequencing adapter as described in more detail herein.

[0079] Accordingly, the disclosed methods have dual advantage. The oligopeptide fragments that are generated by the action of the proteinase on the polypeptide have a statistically determined average length which can be determined by selection of an appropriate proteinase, rather than a random length. The proteinase can be chosen or designed to cleave the polypeptide leaving a common reactive amino acid at a terminus of the oligopeptide. Because each oligopeptide fragment is generated by action of the proteinase each oligopeptide has the same terminal amino acid chemistry and thus is amenable to attachment of a sequencing adapter.

[0080] Accordingly, the disclosed methods offer a beneficial sample preparation method for generating oligopeptide-adapter adducts which are suitable for characterisation e.g. using a nanopore.

[0081] The methods disclosed herein thus involve the processing of a polypeptide. Any suitable polypeptide can be used in the disclosed methods. Exemplary polypeptides are described in more detail herein.

[0082] The methods involve contacting the polypeptide with a proteinase selective for a cleavage site. Any suitable proteinase can be used in the disclosed methods. Exemplary proteinases are described in more detail herein.

[0083] The action of the proteinase on the polypeptide generates oligopeptides each comprising a first end terminating in a first amino acid motif comprising a first reactive amino acid, and a second end terminating in a second amino acid. This is described in more detail herein.

[0084] The disclosed methods further involve selectively attaching to each first reactive amino acid a first sequencing adapter. Any suitable sequencing adapter can be used. Exemplary sequencing adapters are described in more detail herein. Some disclosed methods involve functionalising the first reactive amino acid prior to attachment of the sequencing adapter. Exemplary methods are described in more detail herein.

[0085] In some disclosed methods, the activity of the proteinase on the polypeptide generates oligopeptides which comprise an N-terminal amine group. Some methods involve reacting the N-terminal amine group with an amine-reactive moiety. Exemplary amine-reactive moieties are described in more detail herein.

[0086] The methods result in the preparation of an oligopeptide-adapter adduct for characterisation of the oligopeptide in the oligopeptide-adapter adduct using a nanopore. The oligopeptide-adapter adduct is not limited for use in such methods, however. The oligopeptide-adapter adduct is suitable for any appropriate use and may have independent utility. When the oligopeptide-adapter adduct is to be characterised, it can be characterised in any suitable method. Most generally, the oligopeptide-adapter adduct can be characterised by contacting it with a suitable detector. As further described herein, a nanopore is provided as an exemplary detector which can be used in the disclosed methods. Thus, whilst embodiments described herein refer to characterisation of an oligopeptide-adapter adduct using a nanopore, the methods provided herein are also amenable to other detectors including (i) a zero-mode waveguide, (ii) a field-effect transistor, optionally a nanowire field-effect transistor; (iii) an AFM tip; (iv) a nanotube, optionally a carbon nanotube and (v) a nanopore. The disclosed methods are particularly amenable to methods in which an oligopeptide-adapter adduct is moved through a detector or through a structure containing a detector, e.g. a well in a detector chip.

[0087] Some disclosed methods involve taking one or more measurements characteristic of the oligopeptide as the oligopeptide in the oligopeptide-adapter adduct moves with respect to a nanopore. Some suitable measurements are described in more detail herein.

[0088] These and further details of the disclosed methods are described in more detail herein.

[0089] Polypeptide

[0090] The disclosed methods comprise contacting a polypeptide with a proteinase selective for a cleavage site. Any suitable polypeptide can be used.

[0091] As used herein, the term polypeptide refers to a peptide, polypeptide or protein capable of being fragmented by a proteinase to yield an oligopeptide fragment. The term polypeptide and peptide, polypeptide or protein can be used interchangeably unless implied otherwise by the context.

[0092] In some embodiments the polypeptide is an unmodified protein or a portion thereof, or a naturally occurring polypeptide or a portion thereof.

[0093] In some embodiments the polypeptide is a modified protein or a portion thereof.

[0094] In some embodiments the polypeptide is secreted from cells. Alternatively, the polypeptide can be produced inside cells such that it must be extracted from cells for use in the disclosed methods. The polypeptide may comprise the products of cellular expression of a plasmid, e.g. a plasmid used in cloning of proteins in accordance with the methods described in Sambrook et al., Molecular Cloning: A Laboratory Manual, 4thed., Cold Spring Harbor Press, Plainsview, New York (2012); and Ausubel et al., Current Protocols in Molecular Biology (Supplement 114), John Wiley & Sons, New York (2016).

[0095] In some embodiments, for example, the expression is expression in a bacterial cell, a yeast cell, an insect cell or a mammalian cell; or may be a cell free expression method such as a translation system selected from rabbit reticulocyte lysate, wheat germ extract, and E. coli cell-free systems (available commercially, such as from the PURExpress® systems available from New England Biolabs (Ipswich, MA, USA)). In some embodiments the expression is from the genomic DNA of an organism.

[0096] The polypeptide may be obtained from or extracted from any organism or microorganism. The polypeptide may be obtained from a human or animal, e.g. from urine, lymph, saliva, mucus, seminal fluid or amniotic fluid, or from whole blood, plasma or serum. The polypeptide may be obtained from a plant e.g. a cereal, legume, fruit or vegetable. The polypeptide may be obtained from bacteria, protozoa, algae or fungi.

[0097] The polypeptide can be provided as an impure mixture of one or more polypeptides and one or more impurities. Impurities may comprise truncated forms of the polypeptide which are distinct from the intended polypeptide for use in the disclosed methods. For example, the polypeptide may be a full length protein and impurities may comprise fractions of the protein. Impurities may also comprise proteins other than the polypeptide e.g. which may be co-purified from a cell culture or obtained from a sample.

[0098] A polypeptide may comprise any combination of any amino acids, amino acid analogs and modified amino acids (i.e. amino acid derivatives). Amino acids (and derivatives, analogs etc) in the polypeptide can be distinguished by their physical size and charge.

[0099] The amino acids / derivatives / analogs can be naturally occurring or artificial. In some embodiments the polypeptide may comprise any naturally occurring amino acid. Twenty amino acids are encoded by the universal genetic code. These are alanine (A), arginine (R), asparagine (N), aspartic acid (D), cysteine (C), glutamic acid / glutamate (E), glutamine (Q), glycine (G), histidine (H), isoleucine (I), leucine (L), lysine (K), methionine (M), phenylalanine (F), proline (P), serine (S), threonine (T), tryptophan (W), tyrosine (Y) and valine (V). Other naturally occurring amino acids include selenocysteine and pyrrolysine.

[0100] In some embodiments the polypeptide is modified. In some embodiments the polypeptide is modified for detection using the disclosed methods. In some embodiments the disclosed methods are for characterising modifications in the polypeptide.

[0101] In some embodiments one or more of the amino acids / derivatives / analogs in the polypeptide is modified. In some embodiments one or more of the amino acids / derivatives / analogs in the polypeptide is post-translationally modified. As such, the methods disclosed herein can be used to detect the presence, absence, number of positions of post-translational modifications in a polypeptide. The disclosed methods can be used to characterise the extent to which a polypeptide has been post-translationally modified.

[0102] Any one or more post-translational modifications may be present in the polypeptide. Typical post-translational modifications include modification with a hydrophobic group, modification with a cofactor, addition of a chemical group, glycation (the non-enzymatic attachment of a sugar), biotinylation and pegylation. Post-translational modifications can also be non-natural, such that they are chemical modifications done in the laboratory for biotechnological or biomedical purposes. This can allow monitoring the levels of the laboratory made polypeptide in contrast to the natural counterparts.

[0103] Examples of post-translational modification with a hydrophobic group include myristoylation, attachment of myristate, a Ci4 saturated acid; palmitoylation, attachment of palmitate, a Ci6 saturated acid; isoprenylation or prenylation, the attachment of an isoprenoid group; famesylation, the attachment of a farnesol group; geranylgeranylation, the attachment of a geranylgeraniol group; and glypiation, and glycosylphosphatidylinositol (GPI) anchor formation via an amide bond.

[0104] Examples of post-translational modification with a cofactor include lipoylation, attachment of a lipoate (Cs) functional group; flavination, attachment of a flavin moiety (e.g. flavin mononucleotide (FMN) or flavin adenine dinucleotide (FAD)); attachment of heme C, for instance via a thioether bond with cysteine; phosphopantetheinylation, the attachment of a 4'-phosphopantetheinyl group; and retinylidene Schiff base formation. Examples of post-translational modification by addition of a chemical group include acylation, e.g. O-acylation (esters), N-acylation (amides) or S-acylation (thioesters); acetylation, the attachment of an acetyl group for instance to the N-terminus or to lysine; formylation; alkylation, the addition of an alkyl group, such as methyl or ethyl; methylation, the addition of a methyl group for instance to lysine or arginine; amidation; butyrylation; gamma-carboxylation; glycosylation, the enzymatic attachment of a glycosyl group for instance to arginine, asparagine, cysteine, hydroxylysine, serine, threonine, tyrosine or tryptophan; polysialylation, the attachment of polysialic acid; malonylation; hydroxylation; iodination; bromination; citrulination; nucleotide addition, the attachment of any nucleotide such as any of those discussed above, ADP ribosylation; oxidation; phosphorylation, the attachment of a phosphate group for instance to serine, threonine or tyrosine (O-linked) or histidine (N-linked); adenylyl ati on, the attachment of an adenylyl moiety for instance to tyrosine (O-linked) or to histidine or lysine (N-linked); propionylation; pyroglutamate formation; S-glutathionylation; Sumoylation; S- nitrosylation; succinylation, the attachment of a succinyl group for instance to lysine; sei enoyl ati on, the incorporation of selenium; and ubiquitinilation, the addition of ubiquitin subunits (N-linked).

[0105] It is within the scope of the methods provided herein that the polypeptide is labelled with a molecular label. A molecular label may be a modification to the polypeptide which promotes the detection of the polypeptide in the methods provided herein. For example the label may be a modification to the polypeptide which alters the signal obtained as a conjugate comprising the polypeptide (e.g. an oligopeptide-adapter adduct) is characterised. For example, the label may interfere with a flux of ions through the nanopore. In such a manner, the label may improve the sensitivity of the methods.

[0106] In some embodiments the polypeptide contains one or more cross-linked sections, e.g. C-C bridges. In some embodiments the polypeptide is not cross-linked prior to being characterised using the disclosed methods.

[0107] In some embodiments the polypeptide comprises sulphide-containing amino acids and thus has the potential to form disulphide bonds. Typically, in such embodiments, the polypeptide is reduced using a reagent such as DTT (Dithiothreitol) or TCEP (tris(2- carboxyethyl)phosphine) prior to being characterised using the disclosed methods.

[0108] In some embodiments the polypeptide is a full length protein or naturally occurring polypeptide. The polypeptide can be a polypeptide of any suitable length. In some embodiments the polypeptide has a length of from about 50 to about 40,000 peptide units. In some embodiments the polypeptide has a length of from about 50 to about 35,000 peptide units. In some embodiments the polypeptide has a length of from about 50 to about 20,000 peptide units. In some embodiments the polypeptide has a length of from about 50 to about 10,000 peptide units. In some embodiments the polypeptide has a length of from about 100 to about 5,000 peptide units, for example from about 200 to about 1000 peptide units, e.g. from about 300 to about 500 peptide units, such as from about 200 to about 400 peptide units.

[0109] The polypeptide is fragmented in the disclosed methods to form one or more oligopeptides. As described herein, an oligopeptide generated in the disclosed methods typically has a length of from about 2 to about 50 peptide units, such as from about 5 to about 30 peptide units. Accordingly, the polypeptide has a length greater than the length of the oligopeptide.

[0110] Any number of polypeptides can be used in the disclosed methods. For instance, the method may comprise processing 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 50, 100, 200, 500, 1000, 2000 or more polypeptides or about 10, about 50, about 100, about 200, about 500, about 1000, or about 2000 polypeptides. If two or more polypeptides are used, they may be different polypeptides or two or more instances of the same polypeptide.

[0111] In some embodiments the polypeptide to be processed in the disclosed methods is present in a sample. In some embodiments the sample comprises a plurality of different polypeptides. In some embodiments the plurality may comprise at least 2, at least 3, at least 4, at least 5, at least 10, at least 20, at least 50, at least 100, at least 500, at least 1000, at least 5000, at least 10000, at least 100000, at least 1000000 or more polypeptides.

[0112] Proteinase

[0113] The disclosed methods comprise contacting a polypeptide with a proteinase selective for a cleavage site, under conditions such that the proteinase selectively cleaves the polypeptide thereby forming one or more oligopeptides.

[0114] The term contacting is used herein in its broadest sense to refer to subjecting the polypeptide to the proteinase under conditions such that the proteinase is capable of exerting its polypeptide-cleavage activity. Thus, the polypeptide is contacted with the proteinase (also known as a protease) under conditions that the enzyme is catalytically competent to act on the polypeptide to enzymatically cleave the polypeptide at a suitable cleavage site, also known as a recognition site.

[0115] Any suitable proteinase can be used. Suitable proteinases are described in detail in Handbook of Proteolytic Enzymes, Rawlings and Salvesen (eds), Elsevier 2013, the entire contents of which are hereby incorporated by reference.

[0116] In some embodiments the proteinase is an endo-protease. In some embodiments the proteinase is an exo-proteinase.

[0117] As those skilled in the art will appreciate, choice of a suitable proteinase is an operational parameter of the disclosed methods which can be made by the user of the method according to the polypeptide to be processed and the desired properties of the oligopeptide-adapter adduct generated in the disclosed methods.

[0118] As explained above, the proteinase is selective for a cleavage site, also known as a recognition site. This is described in more detail herein. In brief, however, and by way of non-limiting illustration, the statistical distribution of a given recognition site in a polypeptide comprising a random or pseudo-random arrangement of amino acids (e.g. in a full length protein or other polypeptide as described herein) can be calculated based on the length and sequence of cleavage site and the composition of the polypeptide. For example, it has been observed that lysine residues typically occur roughly every 10-15 amino acids in a random- or pseudo-random polypeptide sequence. If the cleavage site is a lysine, then statistically the protease will cleave the polypeptide every 10-15 amino acids and the average length of the oligopeptide will thus be 10-15 amino acids.

[0119] Accordingly, the proteinase can be chosen or selected to yield oligopeptides of a desired length (e.g. for characterisation in a specific method).

[0120] Furthermore, the selective action of a proteinase at a cleavage site will result in oligopeptides which comprise a uniform terminal amino acid motif. By way of nonlimiting example, the enteropeptidases from E. coli and S. cerevisiae selectively cleave the amino acid sequences DDDDK (SEQ ID NO: 27) immediately C-terminal to the lysine reside. Similarly, thrombin selectively cleaves the sequence LVPRGS (SEQ ID NO: 28) between the arginine (R) and glycine (G) residues.

[0121] Accordingly, the proteinase can be chosen or selected according to the properties of the cleavage sites for which it is specific. Thus, a proteinase can be chosen or selected to cleave a polypeptide and thereby generate oligopeptides which comprise a terminal amino acid motif comprising a first reactive amino acid. With reference to the examples above, action of thrombin on a protein will result in oligopeptides which comprise a terminal amino acid motif LVPR which comprises the reactive amino acid arginine. Thus, a proteinase can be chosen or selected according to the desired properties of the resulting oligopeptides.

[0122] The proteinase selectively cleaves the polypeptide at a cleavage site in the polypeptide. In some embodiments the proteinase selectively cleaves the polypeptide at at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98% or at least 99% of the cleavage sites in the polypeptide. In some embodiments the proteinase selectively cleaves the polypeptide at substantially all or all of the cleavage sites in the polypeptide.

[0123] In some embodiments of the disclosed methods, the cleavage site comprises lysine or arginine. In some embodiments of the disclosed methods, the cleavage site comprises lysine. In some embodiments of the disclosed methods, the cleavage site comprises arginine. In some embodiments the cleavage site consists of lysine or arginine. In some embodiments the cleavage site consists of lysine. In some embodiments, the cleavage site consists of arginine.

[0124] In some embodiments a proteinase that is selective for lysine or arginine is selected from LysC (cleavage site: K), LysN (cleavage site: K), ArgC (cleavage site: R), ArgN (cleavage site: R), enterokinase (cleavage site: [D / E][D / E][D / E]K; SEQ ID NO: 29), and trypsin (cleavage site: [K / R][not-P]; WKP or MRP).

[0125] In some embodiments the cleavage site, the first amino acid motif and the first reactive amino acid each comprise or consist of lysine. Accordingly, in some embodiments the proteinase acts on the polypeptide to cleave at lysine residues in the polypeptide, and results in oligopeptides which comprise a terminal amino acid motif which comprises lysine, which constitutes a first reactive amino acid. Lysine is a reactive amino acid due to the amine group on the side chain.

[0126] In some embodiments the proteinase is LysC or LysN or a or functional analog, fragment or variant thereof. In some embodiments the proteinase is LysC or a or functional analog, fragment or variant thereof.

[0127] LysC (lysyl endopeptidase; E.C. 3.4.21.50) is a lysine-specific proteinase which cleaves polypeptides C-terminal to lysine residues. LysC and equivalent enzymes may be isolated from Achromobacter lyticus, Lysobacter enzymogenes and Pseudomonas aeruginosa. LysC is a serine protease that hydrolyzes specifically at the carboxyl side of lysines. LysC typically retains proteolytic activity under strong protein denaturing conditions such as 8M urea, which can be used to improve digestion of proteolytically resistant proteins. LysC typically has optimal activity in the range of pH 7.0 — 9.0. LysC is available from e.g. Promega and New England Biolabs.

[0128] In some embodiments the cleavage site, the first amino acid motif and the first reactive amino acid each comprise or consist of arginine. Accordingly, in some embodiments the proteinase acts on the polypeptide to cleave at arginine residues in the polypeptide, and results in oligopeptides which comprise a terminal amino acid motif which comprises arginine, which constitutes a first reactive amino acid. Arginine is a reactive amino acid due to the guanidine group on the side chain.

[0129] In some embodiments the proteinase is ArgC or ArgN or a or functional analog, fragment or variant thereof. In some embodiments the proteinase is ArgC or a or functional analog, fragment or variant thereof.

[0130] ArgC (clostripain; E.C. 3.4.21.35) is a arginine-specific proteinase which cleaves polypeptides C-terminal to arginine residues. ArgC and equivalent enzymes may be isolated from Clostridium histolyticum. ArgC is a serine protease that hydrolyzes specifically at the carboxyl side of arginines. ArgC typically has optimal activity in the range of pH 7.0 — 9.0, such as from about 7.6 to 7.9. ArgC is available from e.g. Promega.

[0131] Cleavage of the polypeptide

[0132] In the disclosed methods, contacting the polypeptide with the proteinase results in the proteinase selectively cleaving the polypeptide thereby forming one or more oligopeptides, wherein each oligopeptide comprises a first end terminating in a first amino acid motif comprising a first reactive amino acid, and a second end terminating in a second amino acid.

[0133] It will be apparent that the polypeptide therefore comprises the oligopeptides. The oligopeptides are portions of the polypeptide and are liberated from the polypeptide by the action of the proteinase. Properties of the polypeptide discussed herein therefore typically correspond to properties of the oligopeptide.

[0134] In some embodiments the or each oligopeptide independently has a length (e.g. an average length, e.g. a mean length) of from about 2 to about 50 amino acids. In some embodiments the or each oligopeptide independently has a length of from about 5 to about 30 amino acids, such as from about 6 to about 28 amino acids, e.g. from about 8 to about 25 amino acids, e.g. from about 10 to about 20 amino acids, e.g. about 10, 11, 12, 13, 14, 15, 16, 17, 18, 19 or 20 amino acids. In some embodiments at least about 50%, e.g. at least about 60%, e.g. at least about 70%, e.g. at least about 80%, e.g. at least about 90%, e.g. at least about 95%, 96%, 97%, 98% or 99% of the oligopeptides generated by contacting the polypeptide with the protease have a length of from about 2 to about 50 amino acids, such as from about 5 to about 30 amino acids, such as from about 6 to about 28 amino acids, e.g. from about 8 to about 25 amino acids, e.g. from about 10 to about 20 amino acids, e.g. about 10, 11, 12, 13, 14, 15, 16, 17, 18, 19 or 20 amino acids.

[0135] First and second ends of the oligopeptide

[0136] In some embodiments the first end is the C-terminus of the oligopeptide and the second end is the N-terminus of the polypeptide. In some embodiments the first end is the N-terminus of the oligopeptide and the second end is the C-terminus of the polypeptide.

[0137] In some embodiments the first end is the C-terminus of the oligopeptide and the second end is the N-terminus of the polypeptide, and the second end comprises an N- terminal amine group. The N-terminal amine group is typically the amine group N- terminal to the u-carbon of the N-terminal amino acid residue, wherein the u-carbon has the side chain of the amino acid attached thereto. In some embodiments the amine group is -NH2. In some embodiments the amine group is protonated and is -NTfC.

[0138] In some embodiments the disclosed methods comprise reacting the N-terminal group with an amine-reactive moiety. This is described in more detail herein.

[0139] The disclosed methods comprise reacting the oligopeptide with a first sequencing adapter. Sequencing adapters are described in more detail herein.

[0140] In some embodiments the method comprises reacting the oligopeptide with the first sequencing adapter under conditions such that the first sequencing adapter reacts with the first reactive amino acid. Thus, in some embodiments the method comprises reacting the oligopeptide with the first sequencing adapter under conditions such that the first sequencing adapter reacts with the first reactive amino acid comprised in the first amino acid motif generated by action of the proteinase on the polypeptide.

[0141] In some embodiments the first sequencing adapter is capable of reacting with the first reactive amino acid by reacting with a functional group present in the first reactive amino acid.

[0142] For example, in some embodiments the first reactive amino acid is lysine or arginine and the first sequencing adapter comprises an amine- or guanidine-reactive group capable of reacting with the side chain of the first reactive amino acid. In some embodiments the first reactive amino acid is lysine and the first sequencing adapter comprises an amine-reactive functional group capable of reacting with the side chain of the lysine. In some embodiments the first reactive amino acid is arginine and the first sequencing adapter comprises an guanidine-reactive functional group capable of reacting with the side chain of the arginine.

[0143] In some embodiments the method comprises functionalising the first reactive amino acid for attachment of the first sequencing adapter. In some embodiments the method comprises functionalising the first reactive amino acid by reacting the oligopeptide with a compound comprising a oligo-reactive functional group and an adapter-reactive functional group under conditions such that the oligo-reactive functional group reacts with the first reactive amino acid.

[0144] For example, in some embodiments the first reactive amino acid is lysine or arginine and the method comprises reacting the oligopeptide with a compound comprising an amine- or guanidine-reactive group capable of reacting with the side chain of the first reactive amino acid, wherein the compound further comprises an adapter-reactive functional group. In some embodiments the first reactive amino acid is lysine and the method comprises reacting the oligopeptide with a compound comprising an aminereactive group capable of reacting with the side chain of the lysine, wherein the compound further comprises an adapter-reactive functional group. In some embodiments the first reactive amino acid is arginine and the method comprises reacting the oligopeptide with a compound comprising an guanidine-reactive group capable of reacting with the side chain of the arginine, wherein the compound further comprises an adapter-reactive functional group.

[0145] Suitable functional groups are described in more detail herein.

[0146] Reacting the N-terminal group with an amine-reactive moiety

[0147] As explained herein, in some embodiments the second end of the oligopeptide generated by action of the proteinase on the polypeptide comprises a N-terminal amine group, and the disclosed methods comprise reacting the N-terminal amine group with an amine-reactive moiety.

[0148] In some embodiments, reacting the N-terminal amine group with an amine-reactive moiety comprises immobilizing the oligopeptide on an substrate. In some embodiments the methods comprise immobilizing the oligopeptide on a substrate at the N-terminal of the oligopeptide. In general, the oligopeptide may be immobilized on the substrate e.g. for purification or enrichment of the oligopeptide, via a purification tag which may be present on the oligopeptide (e.g. on a handle attached to the oligopeptide, as described herein), or on the sequencing adapter.

[0149] Any suitable substrate may be used in such methods. For example, in some embodiments the substrate comprises a chromatography matrix, such as an agarose or sepharose resin. Such resins are commercially available from suppliers such as Sigma Aldrich. In some embodiments the substrate comprises beads (i.e. one or more beads). Magnetic beads are often used as such beads allow for facile purification e.g. using washing with buffer. Functionalised magnetic beads are commercially available with a variety of functionalisations from suppliers such as Sigma Aldrich and Bio-Rad. In some embodiments the substrate comprises a solid surface. Any suitable material can be used. Suitable materials include glass, silica, polymers such as polyester, and ceramics such as hydroxyapatite.

[0150] In some embodiments the substrate is an amine-reactive substrate. Amine-reactive substrates are typically substrates such as those described herein, which comprise an amine-reactive functional group on their surface. Any suitable amine-reactive functional group can be used as disclosed herein. In some embodiments reacting the N-terminal amine group with an amine-reactive moiety comprises immobilizing the oligopeptide on an amine-reactive substrate comprising a reactive carbonyl group. In some embodiments the reactive carbonyl group is an aldehyde group. In some embodiments the amine-reactive substrate is functionalised with 2-PCA (2-pyridinyl carboxaldehyde) or an amine-binding derivative thereof. In some embodiments 2-PCA-maleimide chemistry as described herein can be used.

[0151] In some embodiments the oligopeptide is immobilised on the substrate prior to the reaction of the first reactive amino acid with the first sequencing adapter. In some embodiments the oligopeptide is immobilised on the substrate after the reaction of the first reactive amino acid with the first sequencing adapter.

[0152] In some embodiments the method further comprises releasing the oligopeptide from the substrate. In some embodiments the oligopeptide is comprised in the oligopeptide- adapter adduct and the method comprises releasing the oligopeptide-adapter adduct from the substrate. In some embodiments releasing the oligopeptide or oligopeptide-adapter adduct from the substrate comprises cleaving or eluting the oligopeptide from the substrate. In some embodiments releasing the oligopeptide or oligopeptide-adapter adduct from the substrate comprises contacting the oligopeptide-bound substrate with a suitable condition. In some embodiments the condition comprises a competing amine-containing compound which displaces the oligopeptide or oligopeptide-adapter adduct from the substrate. An example of a suitable compound is hydrazine.

[0153] In some embodiments said releasing (e.g. cleaving or eluting) of the oligopeptide generates a second reactive functional group comprised in said oligopeptide capable of reacting with a complementary functional group on a second sequencing adapter.

[0154] In some embodiments, reacting the N-terminal amine group with an amine-reactive moiety comprises reacting the oligopeptide with a nucleotide or peptide handle comprising a purification tag, under conditions such that the handle is attached to the N-terminus of the oligopeptide; thereby generating a fusion adduct.

[0155] In some embodiments the methods comprise contacting the oligopeptide with a ligase and a peptide handle comprising a purification tag, under conditions such that the ligase ligates the peptide handle to the N-terminus of the oligopeptide. In some embodiments the methods comprise reacting the oligopeptide with a peptide handle comprising a purification tag and an amine-reactive group under conditions such that the amine-reactive group of the handle reacts with the N-terminus of the oligopeptide. In some embodiments the methods comprise reacting the oligopeptide with a peptide handle comprising a purification tag and an adapter-binding group. The adapter binding group may generally be any of the reactive groups described herein that is suitable for reacting with a sequencing adapter as described herein; for example the adapter-binding group may be a click chemistry group for reacting with a complementary click chemistry group comprised in a sequencing adapter as described herein.

[0156] For example, in some embodiments the peptide handle comprises a purification tag as described herein, and an amine-reactive group such as 2-PCA or an amine-binding derivative thereof. In some embodiments the peptide handle comprises a purification tag as described herein, and is attached to the N-terminal amine group of the oligopeptide via 2-PCA or 2-PCA-maleimide chemistry as described herein.

[0157] In embodiments which comprise contacting the oligopeptide with a ligase , any suitable peptide ligase can be used. Some examples are disclosed herein. In some embodiments the ligase is a subtiligase; examples of subtiligases are described in Nature Chemical Biology volume 14, pages 50-57 (2018). In some embodiments the ligase is a thymoligase; examples of thymoligases are described in Org. Biomol. Chem., 2018,16, 609-618. In some embodiments the ligase is an omniligase; an example of an omniligase is described in Comput Struct Biotechnol J. 2021 Feb 9: 19: 1277-1287. In some embodiments the ligase is a ligase of SEQ ID NO: 8, 9, or 10, or a functional analog, fragment or variant thereof.

[0158] In embodiments which comprise contacting the oligopeptide with a ligase and a peptide handle comprising a purification tag, any suitable purification tag can be used. Similarly in embodiments which comprise reacting the oligopeptide with a peptide handle comprising a purification tag and an amine-reactive group under conditions such that the amine-reactive group of the handle reacts with the N-terminus of the oligopeptide, any suitable purification tag can be used. Purification tags can be chosen for subsequent immobilization of the fusion adduct on a complementary substrate, also known as a purification support.

[0159] Any suitable purification tag can be used. For example, the purification tag may comprise or consist of biotin. Biotin is particularly suitable for use in the disclosed methods as it forms a strong non-covalent attachment with streptavidin and related proteins (neutravidin, avidin, etc).

[0160] Another suitable purification tag is desthiobiotin. Desthiobiotin binds to streptavidin (e.g. to purification substrates such as a chromatography matrix or magnetic bead) comprising streptavidin and can be displaced by e.g. biotin (e.g. by reaction of biotin with the purification matrix).

[0161] More often, the purification tag is a peptide purification tags suitable for IMAC (immobilised metal affinity chromatography) chemistry. For example, the purification tag may comprise a poly-His tag (e.g. HHHH (SEQ ID NO: 11), HHHHHH (SEQ ID NO: 12) or HHHHHHHH (SEQ ID NO: 13)). Such tags are suitable for binding to a purification support comprising a metal such as nickel or cobalt. Still other purification tags include peptide tags such as Strep (WSHPQFEK (SEQ ID NO: 14)), FLAG (DYKDDDDK (SEQ ID NO: 15)), Human influenza hemagglutinin (HA) (YPYDVPDYA (SEQ ID NO: 16)), Myc (EQKLISEED (SEQ ID NO: 17)), and V5 (GKPIPNPLLGLDST (SEQ ID NO: 18)), etc.

[0162] Other suitable purification tags include: Biotin-carboxy carrier protein (BCCP); Calmodulin binding peptide (CBP); Chitin binding domain (CBD); Histidine affinity tag (HAT); Polyarginine (Arg-tag); Polyaspartate (Asp-tag); Polylysine (Lys-tag); Polyphenylalanine (Phe-tag); Streptavadin-binding peptide (SBP); Tetrazine tag; TCO tag; Azide tag; and DBCO / Alkyne tag.

[0163] In some embodiments the methods comprise contacting the purification tag with a complementary substrate under conditions such that the fusion adduct is immobilised on the substrate.

[0164] Any suitable substrate can be used in such methods. Substrate materials are described in more detail herein. Substrate materials will typically be functionalised with complementary groups for binding the purification tag. Those skilled in the art will appreciate that the support can be functionalised depending on the purification tag that is used. Alternatively, the purification tag can be chosen depending on the support material to be used. Thus, the choice of purification tag and support material is an operational parameter which can be determined by the user of the disclosed methods.

[0165] In some embodiments the support comprises streptavidin, neutravidin or avidin, or a derivative of streptavidin, neutravidin or avidin such as traptavidin. Such supports are particularly useful the purification tag comprises biotin.

[0166] In some embodiments the support comprises a metal such as nickel or cobalt. The metal ion may be provided with a suitable chelator such as nitriloacetic acid (NTA) or iminodiacetic acid (IDA) For example, the support may comprise Ni-NTA. Such supports are particularly useful when the purification tag comprises a His tag.

[0167] In some embodiments the support comprises streptactin. Such supports are particularly useful when the purification tag comprises a Strep tag.

[0168] In some embodiments the support comprises an antibody for a sequence such as FLAG, HA, Myc or V5 as discussed above.

[0169] In some embodiments the oligopeptide is immobilised on the substrate prior to the reaction of the first reactive amino acid with the first sequencing adapter. In some embodiments the oligopeptide is immobilised on the substrate after the reaction of the first reactive amino acid with the first sequencing adapter.

[0170] In some embodiments the method further comprises releasing a portion of the fusion adduct comprising the oligopeptide from the substrate. In some embodiments the method comprises cleaving or eluting a portion of the fusion adduct comprising the oligopeptide from the substrate.

[0171] The portion of the fusion adduct that is released from the substrate comprises the oligopeptide. In some embodiments the portion consists of the oligopeptide. In some embodiments the portion comprises the oligopeptide and part of the handle. For example, the handle may comprise a cleavage site. A cleavage site comprised in the handle may be the same or different to the cleavage site in the polypeptide. Typically a cleavage site in a handle will be different to a cleavage site in the polypeptide.

[0172] In some embodiments the portion comprises the oligopeptide and the handle.

[0173] In some embodiments the methods comprise attaching a second sequencing adapter to the portion. In some embodiments the methods comprise attaching a second sequencing adapter to the handle or to part of the handle.

[0174] In some embodiments the second sequencing adapter is comprised in the handle. In some embodiments the handle comprises the second sequencing adapter and a purification tag.

[0175] In some embodiments the handle comprises the second sequencing adapter and a purification tag and the second sequencing adapter is not attached to the purification tag via a cleavage site. In some embodiments the second sequencing adapter is attached to the purification tag via a cleavage site and the methods do not comprise cleaving the cleavage site. In such embodiments the handle may be attached to the oligopeptide and the fusion adduct may be immobilized on a suitable support (e.g. for purification and / or enrichment). The immobilized adduct may then be treated with an appropriate condition for eluting the purification tag from the support.

[0176] For example, in some embodiments the handle comprises the second sequencing adapter and a purification tag, such as an affinity purification tag such as a poly-His tag. Exemplary non-limiting chemistry that may be used to attach the handle to the oligopeptide includes 2-PCA (e.g. 2-PCA / maleimide) chemistry. In some embodiments the handle may be attached to the oligopeptide and the fusion adduct may be immobilized on a suitable support (e.g. a chromatography matrix, or magnetic beads) comprising a metal such as cobalt or nickel for binding of the purification tag thereby immobilizing the adduct at the support. In some embodiments the immobilized adduct is treated with an appropriate condition for eluting the purification tag from the support. For example the immobilized adduct may be eluted from the support after enrichment and / or purification. In some embodiments the immobilized adduct is eluted from the support after attachment of the first sequencing adapter. In some embodiments the affinity purification tag is not attached to the second sequencing adapter via a cleavage site. In some embodiments the affinity purification tag is not attached to the second sequencing adapter via a cleavage site and the methods do not comprise cleaving the cleavage site. Accordingly in some embodiments the methods do not comprise cleaving the purification tag from the second sequencing adapter.

[0177] In some embodiments the second sequencing adapter is attached to the purification tag via a cleavage site. In such embodiments the handle may be attached to the oligopeptide and the fusion adduct may be immobilized on a suitable support. The immobilized adduct may then be treated with an appropriate condition for cleaving the cleavage site thereby liberating the second sequencing adapter attached to the oligopeptide, which may be comprised in an oligopeptide-adapter adduct with the first sequencing adapter.

[0178] In some embodiments the methods comprise attaching a second sequencing adapter to the portion after the portion has been released (e.g. by being cleaved or eluted) from the substrate. In some embodiments the methods comprise attaching a second sequencing adapter to the portion after the portion has been released (e.g. by being cleaved or eluted) from the fusion adduct. In some embodiments the methods comprise attaching the second sequencing adapter to the portion whilst the portion is immobilized on the substrate. In some embodiments the methods comprise attaching the second sequencing adapter to the portion whilst the fusion adduct is immobilized on the substrate.

[0179] In some embodiments the handle comprises a second reactive functional group capable of reacting with a complementary functional group on the second sequencing adapter. In some embodiments the said releasing (e.g. cleaving or eluting) of the portion generates a second reactive functional group comprised in said oligopeptide capable of reacting with a complementary functional group on a second sequencing adapter.

[0180] For example, in some embodiments the handle comprises a purification tag (e.g. an affinity purification tag such as a His tag, desthiobiotin tag, or other purification tag described herein) and a second reactive functional group capable of reacting with a complementary functional group on the second sequencing adapter. In some embodiments the second reactive functional group is a click chemistry group (e.g. an azide, tetrazine (e.g. methyl tetrazine, MeTet), alkyne, TCO, etc group) and is capable of reacting with a complementary click chemistry group in the second sequencing adapter. These and other suitable click chemistry reactions are described in more detail herein.

[0181] In some embodiments the handle comprises a purification tag (such as a His tag, desthiobiotin tag, or other purification tag described herein) and a second reactive functional group which is a click chemistry group capable of reacting with a complementary click chemistry group in the second sequencing adapter, wherein one of the click chemistry groups is an alkyne (e.g. BCN) group and the other is an azide group (e.g. the handle may comprise an azide group and the second sequencing adapter may comprise a complementary alkyne (e.g. BCN) group). In some embodiments the handle comprises a purification tag (such as a His tag, desthiobiotin tag, or other purification tag described herein) and a second reactive functional group which is a click chemistry group capable of reacting with a complementary click chemistry group in the second sequencing adapter via an inverse electron demand Diels-Alder reaction (iEDDA). In one embodiment the handle comprises a methyltetrazine group and the sequencing adapter comprises a trans-cyclooctene (TCO) group. In one embodiment the handle comprises a TCO group and the sequencing adapter comprises a methyltetrazine group.

[0182] Those skilled in the art will appreciate that the methods may comprise purification of the oligopeptide via the purification tag before, simultaneously with, or subsequent to, the attachment of the second sequencing adapter to the second reactive functional group. For example, in some embodiments the method comprise purifying the oligopeptide using the purification tag, and then attaching the sequencing adapter to the reactive group on the handle. Following the reaction, the functionalised oligopeptide-sequencing adapter may be released (e.g. eluted) from the purification matrix. The first sequencing adapter may be already present on the oligopeptide prior to attachment of the second sequencing adapter, or may be attached to the oligopeptide whilst the oligopeptide is immobilised on the purification matrix, or may be attached to the oligopeptide after release of the construct formed by attachment of the oligopeptide to the second sequencing adapter from the purification matrix.

[0183] In some embodiments the second sequencing adapter is attached to the portion prior to the attachment of the first sequencing adapter to the oligopeptide. In some embodiments the second sequencing adapter is attached to the portion after to the attachment of the first sequencing adapter to the oligopeptide.

[0184] Second reactive functional groups capable of reacting with a complementary functional group on a second sequencing adapter are described in more detail herein.

[0185] In some embodiments, reacting the N-terminal amine group with an amine-reactive moiety comprises attaching a first sequencing adapter to the first reactive amino acid and to the N-terminal amine group. Accordingly, in some embodiments the methods comprise attaching one molecule of the first sequencing adapter to the first reactive amino acid (e.g. at the C-terminus of the oligopeptide) and attaching a second molecule of the first sequencing adapter to the N-terminal amine group. In some embodiments both the C- terminus and the N-terminus of the oligopeptide therefore comprise the first sequencing adapter.

[0186] In some embodiments the method comprises cleaving the amino acid comprising the reacted N-terminal amine group from the oligopeptide-adapter adduct. In such a manner, the length of the oligopeptide in the oligopeptide-adapter adduct is reduced by one amino acid.

[0187] In some embodiments the reacted N-terminal amine group is cleaved from the oligopeptide-adapter adduct using a chemical reagent. In some embodiments the reacted N-terminal amine group is cleaved from the oligopeptide-adapter adduct using an acid, e.g. trifluoroacetic acid, or citric acid, in some embodiments the embodiments the reacted N- terminal amine group is cleaved from the oligopeptide-adapter adduct using trifluoroacetic acid.

[0188] In some embodiments the N-terminal amino acid is cleaved from the oligopeptide- adapter adduct using an Edmanase enzyme. An example of an Edmanase enzyme is Aminopeptidase I from Pyrococcus furiosiis. which is available from Takara Bio. Edmanases are also disclosed in Borgo and Havranek, Protein Sci. 24, 571-579, 2015, which is incorporated by reference herein.

[0189] In some embodiments cleaving the reacted N-terminal amine group from the oligopeptide-adapter adduct generates a second reactive functional group capable of reacting with a complementary functional group on a second sequencing adapter.

[0190] Reactions between an oligopeptide and a sequencing adapter

[0191] The disclosed methods comprise selectively attaching to each first reactive amino acid a first sequencing adapter, thereby forming an oligopeptide-adapter adduct. Many embodiments of the disclosed methods further comprise attaching a second sequencing adapter to the oligopeptide.

[0192] In embodiments which comprise a second sequencing adapter, the second sequencing adapter may be attached to the oligopeptide prior to attachment of the first sequencing adapter, simultaneously with the attachment of the first sequencing adapter, or subsequently to the attachment of the first sequencing adapter. In some embodiments when the method comprises reacting the oligopeptide with a compound comprising an amine- or guanidine-reactive group capable of reacting with the side chain of the first reactive amino acid, wherein the compound further comprises an adapter-reactive functional group, the method further comprises reacting the first sequencing adapter with the adapter-reactive functional group.

[0193] In some embodiments the attachment between the oligopeptide and a sequencing adapter is covalent. In some embodiments the attachment is non-covalent. In some embodiments the attachment comprises ligating the sequencing adapter to the oligopeptide.

[0194] As will be apparent from the discussion herein, the attachment between a sequencing adapter and the oligopeptide is not especially limited and any suitable attachment means can be used. In some embodiments the attachment means is chosen according to whether or not the oligopeptide is to be cleaved from the construct in the disclosed methods. In some embodiments the attachment means is chosen or designed to be amenable to cleavage under the conditions of the method. In some embodiments the attachment means is designed or chosen to resist cleavage.

[0195] In embodiments which involve the attachment of a second sequencing adapter, the second sequencing adapter is typically attached to the oligopeptide via a second reactive functional group. The second reactive functional group is typically complementary to a reactive functional group on the second sequencing adapter.

[0196] Accordingly, some embodiments of the disclosed methods comprise attaching the second sequencing adapter to the second reactive functional group.

[0197] In some embodiments the second reactive functional group is the N-terminal amine group. In some embodiments the second reactive functional group is the N-terminal amine group of the oligopeptide. In some embodiments the second reactive functional group is the N-terminal amine group of an amino acid comprised in a handle attached to the oligopeptide. In such embodiments, the second sequencing adapter is typically attached to the second reactive functional group via an amine-reactive functional group on the second sequencing adapter. Suitable amine-reactive functional groups are described in more detail herein.

[0198] In some embodiments the second reactive functional group is a group that is introduced e.g. in a handle that may be ligated onto the oligopeptide. In some embodiments the second reactive functional group may be pendant to the handle. In such embodiments the choice of second reactive functional group is not particularly limited and any suitable group can be used to react with a complementary group on a second sequencing adapter.

[0199] In some embodiments the chemistry chosen to attach the second sequencing adapter to the oligopeptide or a fusion adduct comprising the oligopeptide (e.g. to a handle) depends on the first reactive amino acid.

[0200] For example, in some embodiments the first reactive amino acid is lysine. Because the reactive functional group in lysine is an amine group, and the N-terminus of the oligopeptide typically comprises an amine, the N-terminal amine group is typically protected by reacting with an amine-reactive moiety prior to reaction of the first reactive amino acid with the first sequencing adapter. After the reaction of the first reactive amino acid with the first sequencing adapter, a second sequencing adapter can be reacted with an N-terminal reactive group at the N-terminus of the oligopeptide or of a fusion adduct comprising the oligopeptide.

[0201] In some embodiments the first reactive amino acid is arginine. Because the reactive functional group in arginine is an guanidine group, and the N-terminus of the oligopeptide typically comprises an amine, different chemistry can be used to selectively attach the first sequencing adapter to the first reactive amino acid and a second sequencing adapter to the N-terminal amine group.

[0202] In some embodiments the second reactive functional group is orthogonal to the optionally-functionalised first reactive amino acid. Many examples of orthogonal chemical reactive groups are known and are disclosed herein. Typically, orthogonal functional groups do not cross-react. Thus, when the second reactive functional group is orthogonal to the optionally-functionalised first reactive amino acid the first sequencing adapter will not react with the second reactive functional group and the second sequencing adapter will not react with the optionally functionalised first reactive amino acid.

[0203] Choice of appropriate reactive groups is within the capability of the skilled person. Any suitable reactive groups can be used.

[0204] The attachment chemistry between the oligopeptide and a sequencing adapter is not particularly limited and any suitable chemistry can be used. Practitioners are directed to texts such as March's Advanced Organic Chemistry: Reactions, Mechanisms, and Structure (2019), ed. Smith, Wiley, which is hereby incorporated by reference in its entirety. Practitioners are directed particularly to discussion in that text of reactions of amines and guanidines. For example, the adapter or oligopeptide (or a handle attached thereto) may be modified using chemical methods such as the use of 2-PCA derivatives, oxazolone chemistry, photoreductive decarboxylation chemistry and side-chain specific chemistry such as NHS-esters, maleimides, iodoacetamides, fluoroacetamides, chloroacetamides, and others. Further exemplary reactive groups and their corresponding targets include aryl azides which may react with amine, carbodiimides which may react with amines and carboxyl groups, hydrazides which may react with carbohydrates, hydroxmethyl phosphines which may react with amines, imidoesters which may react with amines, isocyanates which may react with hydroxyl groups, carbonyls which may react with hydrazines, maleimides which may react with sulfhydryl groups, NHS-esters which may react with amines, PFP-esters which may react with amines, psoralens which may react with thymine, pyridyl disulfides which may react with sulfhydryl groups, vinyl sulfones which may react with sulfhydryl amines and hydroxyl groups, vinylsulfonamides, and the like.

[0205] Other suitable chemistry for conjugating the adapter and a sequencing adapter includes click chemistry. Many suitable click chemistry reagents are known in the art. Suitable examples of click chemistry include, but are not limited to, the following:

[0206] (a) copper(I)-catalyzed azide-alkyne cycloadditions (azide alkyne Huisgen cycloadditions);

[0207] (b) strain-promoted azide-alkyne cycloadditions; including alkene and azide [3+2] cycloadditions; alkene and tetrazine inverse-demand Diels- Alder reactions; and alkene and tetrazole photoclick reactions;

[0208] (c) copper-free variant of the 1,3 dipolar cycloaddition reaction, where an azide reacts with an alkyne under strain, for example in a cyclooctane ring such as in bicycle[6.1.0]nonyne (BCN);

[0209] (d) the reaction of an oxygen nucleophile on one linker with an epoxide or aziridine reactive moiety on the other; and

[0210] (e) the Staudinger ligation, where the alkyne moiety can be replaced by an aryl phosphine, resulting in a specific reaction with the azide to give an amide bond. Any reactive group may be used in the reaction step. Some suitable reactive groups include [1, 4-Bis[3-(2-pyridyldithio)propionamido]butane; 1,1 1-bis- maleimidotriethyleneglycol; 3,3’-dithiodipropionic acid di (N-hydroxy succinimide ester); ethylene glycol -bis(succinic acid N-hydroxysuccinimide ester); 4,4’- diisothiocyanatostilbene-2,2’-disulfonic acid disodium salt; Bis[2-(4- azidosalicylamido)ethyl] disulphide; 3-(2-pyridyldithio)propionic acid N- hydroxysuccinimide ester; 4-maleimidobutyric acid N-hydroxy succinimide ester; lodoacetic acid N-hydroxysuccinimide ester; S-acetylthioglycolic acid N- hydroxysuccinimide ester; azide-PEG-maleimide; and alkyne-PEG-maleimide. The reactive group may be any of those disclosed in WO 2010 / 086602, particularly in Table 3 of that application.

[0211] In some embodiments a sequencing adapter is attached to an oligopeptide or a handle attached thereto via an amine-reactive group (e.g. on the sequencing adapter, or on a compound used to activate the oligopeptide or handle for attachment of the sequencing adapter).

[0212] In one embodiment an amine-reactive group is a carboxylic acid or an activated derivative thereof (e.g. an NHS-ester), which can form amide bonds with amine groups. Accordingly, in some embodiments the method comprises reacting the oligopeptide with the sequencing adapter, wherein the sequencing adapter comprises a carboxylic acid or activated derivative thereof.

[0213] In one embodiment an amine-reactive group is a thiol or an activated derivative thereof, which can form thioether bonds by reaction with amine groups. Accordingly, in some embodiments the method comprises reacting the oligopeptide with the sequencing adapter, wherein the sequencing adapter comprises a thiol or activated derivative thereof.

[0214] In one embodiment an amine-reactive group is a thiol, which reacts with amines in the presence of a furan and an oxidising agent to form a pyrrole. Accordingly, in some embodiments the method comprises reacting the oligopeptide with the sequencing adapter, wherein the sequencing adapter comprises a thiol or activated derivative thereof.

[0215] In one embodiment an amine-reactive group is a squaramate, which reacts with amines to form a squaramide. Accordingly, in some embodiments the method comprises reacting the oligopeptide with the sequencing adapter, wherein the sequencing adapter comprises a squaramate or activated derivative thereof.

[0216] In one embodiment an amine-reactive group is an amine coupling agent (e.g. N,N'- Disuccinimidyl carbonate, l,l'-Carbonyl diimidazole) which reacts with amines to form a urea linkage. Accordingly, in some embodiments the method comprises reacting the oligopeptide with the sequencing adapter, wherein the sequencing adapter comprises an amine coupling agent or activated derivative thereof.

[0217] In one embodiment an amine-reactive group is an aldehyde or ketone, which reacts with amines via reductive amination (e.g. via reduction using a reducing agent such as NaBH4 or NaBFhCN). Accordingly, in some embodiments the method comprises reacting the oligopeptide with the sequencing adapter, wherein the sequencing adapter comprises an aldehyde or ketone or activated derivative thereof.

[0218] An exemplary amine-reactive aldehyde is 2-PCA (2 -pyridine carboxaldehyde) and its derivatives. 2-PCA has the following structure wherein the wavy line represents a bond to hydrogen.

[0219] Useful derivatives of 2-PCA have the above structure wherein the wavy line represents the bond to a sequencing adapter, handle, purification support or other moiety as disclosed herein. For example, in some embodiments 2-PCA is used as a reactive functional group for reaction with an amine such as the N-terminal amine group of a oligopeptide as disclosed here. In some embodiments an exemplary compound for reaction with an amine group such as the N-terminal amine group thus comprises a handle and / or sequencing adapter as described here, and an optional purification tag which may optionally be attached to the sequencing adapter via a linker such as a cleavable linker or a non-cleavable linker; and a 2PCA group. In some embodiments the compound has the structure: wherein purification tag is a purification tag as described herein; optional linker is an optionally present linker which may or may not be cleavable (e.g. may or may nor comprise a cleavage site as described herein); and sequencing adapter is a sequencing adapter as described herein. Accordingly, in some embodiments the method comprises functionalising an amine group of the oligopeptide by reacting the oligopeptide with a compound comprising an oligo-reactive functional group which is a 2-PCA group or derivative thereof. In some embodiments the method comprises reacting the functionalised amine with the sequencing adapter, wherein the sequencing adapter comprises a 2-PCA group or derivative thereof. In some embodiments amine-reactive chemistry is provided by 2-PCA-maleimide chemistry such as that described in Hanaya et al, Angew. Chem. Int. Ed. 64(5) e202417134. In some embodiments the reaction comprises cycloaddition of a mal eimide group coupled to amine-reaction with a 2-PCA derivative such as a derivative described herein. In some embodiments the reaction is conducted under non-denaturing conditions. In some embodiments the reaction is catalysed by copper (e.g. Cu(II)). In some embodiments the reaction proceeds as shown in the following mechanism wherein R represents the side chain of the first amino acid of the oligopeptide (e.g. the side chain of the N-terminal amino acid of the oligopeptide); the wiggly line represents the derivative of the 2-PCA group (e.g. a handle or sequencing adapter as described herein; or a reactive group for further reaction, such as a sequencing-adapter-reactive group; for example, an azide group for a click reaction with a sequencing adapter or handle as described herein); and R1represents a chemical group. In some embodiments R1may be a chemical moiety such as H, alkyl, etc. In some embodiments R1may comprise or consist of a purification tag such as a purification tag described herein. In some embodiments R1may comprise or consist of a further reactive group for reaction with e.g. a purification tag. In some embodiments the reaction of the oligopeptide with the 2-PCA and maleimide moiety is conducted at ambient temperature (e.g. from about 15 to about 40 °C, such as about 20 to about 35 °C). In some embodiments the reaction of the oligopeptide with the 2-PCA and maleimide moiety is conducted in the presence of Cu(II). In some embodiments the reaction is conducted at a pH of from about 5 to about 7, such as about pH 6. Accordingly, in some embodiments the method comprises functionalising an amine group of the oligopeptide by reacting the oligopeptide with a compound comprising an oligo-reactive functional group which is a 2-PCA group or derivative thereof; and a maleimide compound. In some embodiments the method comprises functionalising an amine group of the oligopeptide by reacting the oligopeptide with a compound comprising an oligo-reactive functional group which is a 2-PCA group or derivative thereof; and subsequently stabilizing the moiety formed by reaction of the oligopeptide with the 2-PCA group (or derivative thereof) by reaction with a maleimide compound. In some embodiments the method comprises reacting the functionalised amine with the sequencing adapter, wherein the sequencing adapter comprises a 2-PCA group or derivative thereof; and stabilizing the moiety formed by reaction of the sequencing adapter with the oligopeptide by reaction with a maleimide compound.

[0220] Accordingly, in some embodiments the disclosed methods comprise reacting one or more amine groups with one or more amine-reactive moieties which comprise a 2-PCA group (or derivative thereof) as described herein. In some embodiments said reaction comprises one or more maleimides such that the reaction of the one or more amine groups with one or more amine-reactive moieties comprises using 2-PCA-maleimide chemistry as described herein.

[0221] In one embodiment an amine may be activated by reaction with a maleimide- containing compound. For example, a maleimide-NHS-ester (e.g. 3-maleimido-propionic NHS ester) may be used, with the amine group reacting with the NHS -ester to form an amide, followed by reaction of the free maleimide group with the sequencing adapter; for example with a thiol group of the sequencing adapter, or with a diene such as a furan group. Accordingly, in some embodiments the method comprises functionalising an amine group of the oligopeptide by reacting the oligopeptide with a compound comprising an oligo-reactive functional group which is a carboxylic acid or activated derivative thereof (such as an NHS-ester) and an adapter-reactive functional group which is a maleimide or activated derivative thereof. In some embodiments the method comprises reacting the functionalised amine with the sequencing adapter, wherein the sequencing adapter comprises a thiol or activated derivative thereof.

[0222] In one embodiment the method comprises reacting the oligopeptide with a sequencing adapter in a click chemistry reaction.

[0223] In one embodiment the click chemistry reaction is a copper-catalysed azide / alkyne cycloaddition (CuAAC) reaction. In one embodiment an amine may be activated by reaction with an azide-containing compound, such as an azidoacetic acid NHS-ester, followed by reaction of the free azide group with the sequencing adapter, for example with an alkyne group of the sequencing adapter. In one embodiment an amine may be activated by reaction with an alkyne-containing compound, followed by reaction of the alkyne group with the sequencing adapter, for example with an azide group of the sequencing adapter, such as an azidoacetic acid NHS-ester. In one embodiment the click chemistry reaction is a strain-promoted azide-alkyne reaction. In one embodiment an amine may be activated by reaction with an azide- containing compound, such as an azidoacetic acid NHS-ester, followed by reaction of the free azide group with the sequencing adapter, for example with a cyclooctynyl group of the sequencing adapter (e.g. a bicyclononyne, BCN, group). In one embodiment an amine may be activated by reaction with an cyclooctynyl -containing compound, such as a bicyclononyne (BCN) group, followed by reaction of the alkyne group with the sequencing adapter, for example with an azide group of the sequencing adapter, such as an azidoacetic acid NHS-ester.

[0224] In one embodiment the click chemistry reaction is a inverse electron demand Diels- Alder reaction (iEDDA). In one embodiment an amine may be activated by reaction with an tetrazine-containing compound, such as a methyltetrazine-containing compound, followed by reaction of the free tetrazine group with the sequencing adapter, for example with a trans-cyclooctene group of the sequencing adapter. In one embodiment an amine may be activated by reaction with a trans-cycloocten-containing compound, followed by reaction of the TCO group with the sequencing adapter, for example with a tetrazine group (e.g. a methyltetrazine group) of the sequencing adapter.

[0225] These methods are illustrated by reference to Figure 6.

[0226] In some embodiments an oligopeptide is functionalised by reaction of a first reactive group such as an amine group comprised in a side chain of a lysine at the C- terminus of the oligopeptide with a azide-containing compound, such as an azidoacetic acid NHS-ester; reaction of a second reactive group such as an N-terminal amine group of the oligopeptide with a 2-PCA derivative comprising a sequencing adapter or handle as described herein; optionally stabilizing the reaction between the N-terminal amine group of the oligopeptide and the 2-PCA derivative with a maleimide compound as described herein; and attaching a first sequencing adapter to the azide-functionalised first reactive group e.g. by an azide-alkyne click reaction, e.g. wherein the first sequencing adapter comprises an alkyne group for reaction with the azide group.

[0227] In some embodiments a sequencing adapter is attached to an oligopeptide or a handle attached thereto via a guanidine-reactive group (e.g. on the sequencing adapter, or on a compound used to activate the oligopeptide or handle for attachment of the sequencing adapter).

[0228] In one embodiment a guanidine-reactive group may be activated by reaction with a glyoxal -containing compound. For example, a compound comprising a glyoxal group coupled to a click chemistry group such as an azide or alkyne group may be used, with the guanidine group reacting with the glyoxal group, followed by reaction of the free click chemistry group with the sequencing adapter; for example with an azide or alkyne group of the sequencing adapter. Accordingly, in some embodiments the method comprises functionalising a guanidine group of the oligopeptide by reacting the oligopeptide with a compound comprising an oligo-reactive functional group which is a glyoxal group or activated derivative thereof and an adapter-reactive functional group such as a click chemistry group. In some embodiments the method comprises reacting the functionalised guanidine with the sequencing adapter, wherein the sequencing adapter comprises a complementary click chemistry group.

[0229] In one embodiment an guanidine-reactive group is a diketone such as a 1,2- diketone or 1,3-diketone. Accordingly, in some embodiments the method comprises reacting the oligopeptide with the sequencing adapter, wherein the sequencing adapter comprises a diketone such as a 1,2-diketone or 1,3-diketone.

[0230] In one embodiment an guanidine-reactive group is an NHS-ester, which can form amide bonds with guanidine groups. Accordingly, in some embodiments the method comprises reacting the oligopeptide with the sequencing adapter, wherein the sequencing adapter comprises a NHS-ester.

[0231] In one embodiment a guanidine-reactive group may be activated by citrullination via arginine deaminase to form a carbamide group, followed by reaction of the carbamide group with the sequencing adapter, for example with a boron-containing functional group such as 2-GBA of the sequencing adapter. Accordingly, in some embodiments the method comprises functionalising a guanidine group of the oligopeptide by reacting the oligopeptide with an arginine deaminase enzyme to form a carbamide group. In some embodiments the method comprises reacting the carbamide group with the sequencing adapter, wherein the sequencing adapter comprises a boron-containing functional group such as 2-GBA.

[0232] In one embodiment a guanidine-reactive group may be activated using an aldehyde such as formaldehyde, followed by reaction of the carbamide group with the sequencing adapter, for example with an amine group of the sequencing adapter. Accordingly, in some embodiments the method comprises functionalising a guanidine group of the oligopeptide by reacting the oligopeptide with an aldehyde such as formaldehyde. In some embodiments the method comprises reacting the product with the sequencing adapter, wherein the sequencing adapter comprises an amine group. In one embodiment the method comprises reacting the arginine-containing oligopeptide with a sequencing adapter using a trypsiligase.

[0233] Sequencing adapters

[0234] The disclosed methods comprise attaching the oligopeptide to a first sequencing adapter and typically further comprise attaching a second sequencing adapter also.

[0235] In some embodiments wherein the disclosed methods comprise attaching a first sequencing adapter and a second sequencing adapter to the oligopeptide, attaching said first sequencing adapter and said second sequencing adapter to the oligopeptide generates an adduct of form:

[0236] S -N~O~C. F wherein F is the first sequencing adapter,W’O’Cis the oligopeptide wherein N- represents the N-terminal end of the oligopeptide and -C represents the C-terminal end of the oligopeptide, and S is the second sequencing adapter.

[0237] Any suitable sequencing adapter may be used as the first and second sequencing adapter, independently. For example, if the oligopeptide-adapter adduct is for characterisation using a nanopore, the choice of the first and second sequencing adapters may be dependent on the nanopore characterisation methods that are envisaged.

[0238] Typically the first sequencing adapter is different to the second sequencing adapter.

[0239] In some embodiments the first sequencing adapter is a polynucleotide sequencing adapter. In some embodiments the second sequencing adapter comprises a polynucleotide and / or a polypeptide. In some embodiments the second sequencing adapter is a polynucleotide sequencing adapter. In some embodiments the second sequencing adapter is a polypeptide sequencing adapter.

[0240] In some embodiments an adapter may function as a barcode or may comprise a barcoding sequence (e.g. an amino acid or polypeptide sequence that functions as a barcode). In some embodiments an adapter may comprise or consist of a barcode. As used herein a barcode may be understood as a moiety (e.g. a polypeptide or polynucleotide sequence) that is associated with or reports on the properties of the oligopeptide. For example, a barcode may be designed or chosen to have properties which are characteristic of one or more characteristics of the oligopeptide to which the barcode is associated. Nanopore analysis as described herein may allow information about the abundance of barcodes in a sample, the length of the barcode, and the identity of the barcode (e.g. its sequence, when the barcode is a polypeptide barcode) to be selectively determined; and this can in some embodiments provide information about the oligopeptide. For example, analysis of the barcode may provide information about one or more characteristics of the oligopeptide such as (i) the length of the oligopeptide, (ii) the identity of the oligopeptide, (iii) the sequence of the oligopeptide, (iv) the secondary structure of the oligopeptide; (v) whether or not and / or to the extent to which the oligopeptide is modified; etc. The utility of barcodes in this manner is described in United Kingdom patent application No.

[0241] 2413109.6 and PCT application number PCT / EP2025 / 075372, the entire contents of which are hereby incorporated by reference.

[0242] In one embodiment, the or each adapter is synthetic or artificial. Typically, the or each adapter comprises a polymer as described herein. In some embodiments, the or each adapter comprises a spacer as described herein. In some embodiments, the or each adapter comprises a polynucleotide. The or each polynucleotide adapter may comprise DNA, RNA, modified DNA (such as abasic DNA), RNA, PNA, LNA, BNA and / or PEG. Usually, the or each adapter comprises single stranded and / or double stranded DNA or RNA.

[0243] In some embodiments, an adapter is a linear adapter. A linear adapter may be bound to either or both ends of a oligopeptide.

[0244] A linear adapter may comprise a leader sequence as described herein. A linear adapter may comprise a portion for hybridisation with a tag (such as a pore tag) as described herein. A linear adapter may be 10 to 150 nucleotides in length, such as from 20 to 120, e.g. 30 to 100, for example 40 to 80 such as 50 to 70 nucleotides in length. A linear adapter may be single stranded. A linear adapter may be double stranded.

[0245] In some embodiments, an adapter may be a Y adapter. A Y adapter is typically a polynucleotide adapter. A Y adapter is typically double stranded and comprises (a) at one end, a region where the two strands are hybridised together and (b), at the other end, a region where the two strands are not complementary. The non-complementary parts of the strands typically form overhangs. The presence of a non-complementary region in the Y adapter gives the adapter its Y shape since the two strands typically do not hybridise to each other unlike the double stranded portion. The two single-stranded portions of the Y adapter may be the same length, or may be different lengths. For example, one singlestranded portion of the Y adapter may be 10 to 150 nucleotides in length, such as from 20 to 120, e.g. 30 to 100, for example 40 to 80 such as 50 to 70 nucleotides in length and the other single stranded portion of the Y adapter may independently by 10 to 150 nucleotides in length, such as from 20 to 120, e.g. 30 to 100, for example 40 to 80 such as 50 to 70 nucleotides in length. The double-stranded “stem” portion of the Y adapter may be e.g. from 10 to 150 nucleotides in length, such as from 20 to 120, e.g. 30 to 100, for example 40 to 80 such as 50 to 70 nucleotides in length. A Y adapter may be attached to either or both ends of a barcode as described herein.

[0246] An adapter may be linked to the oligopeptide by any suitable means known in the art, according to the present disclosure. The adapter may be synthesized separately and chemically attached or enzymatically ligated to the strand as described herein.

[0247] An adapter suitable for use in the described methods may in some embodiments comprise a leader. A leader may be useful to assist the capture of the adapter and thus of the barcode by a nanopore as described herein.

[0248] In some embodiments the leader may be from about 10 to 150 nucleotides (e.g. DNA and / or RNA nucleotides) in length, such as from 20 to 120, e.g. 30 to 100, for example 40 to 80 such as 50 to 70 nucleotides in length, or from about 10 to about 60 nucleotides in length, e.g. from about 20 to about 50, such as from about 20 to about 40, e.g. about 30 nucleotides in length.

[0249] In some embodiments the leader is a charged polymer, e.g. a negatively charged polymer. In some embodiments the leader comprises a polymer such as PEG or a polysaccharide. In such embodiments the leader may be from 10 to 150 monomer units (e.g. ethylene glycol or saccharide units) in length, such as from 20 to 120, e.g. 30 to 100, for example 40 to 80 such as 50 to 70 monomer units (e.g. ethylene glycol or saccharide units) in length.

[0250] An adapter may comprise a polypeptide. An adapter may comprise a leader comprising a polypeptide. The polypeptide may in some embodiments have a net negative charge or it may have a specific recognition sequence, e.g. a sequence specific for a motor protein such as an unfoldase enzyme. In such embodiments the leader may be from 10 to 150 amino acids in length, such as from 20 to 120, e.g. 30 to 100, for example 40 to 80 such as 50 to 70 amino acids in length.

[0251] Motor proteins

[0252] In some embodiments the disclosed methods involve allowing the oligopeptide to move with respect to a nanopore, and taking measurements during said movement thereby characterising the oligopeptide. This is described in more detail herein. The movement of the oligopeptide with respect to the nanopore may be driven by any suitable means. In some embodiments, the movement of the oligopeptide is driven by a physical or chemical force (potential). In some embodiments the physical force is provided by an electrical (e.g. voltage) potential or a temperature gradient, etc.

[0253] In some embodiments, the movement of the oligopeptide comprises mechanically manipulating the oligopeptide thereby moving said oligopeptide with respect to the nanopore. In some embodiments, movement of the oligopeptide by mechanical manipulation does not comprise using a polynucleotide-handling protein.

[0254] In some embodiments the oligopeptide is moved by mechanical manipulation in a direction opposite to a potential applied across said nanopore. In some embodiments, the potential is a voltage potential applied across said nanopore. In some embodiments, the oligopeptide is moved with respect to the nanopore as described in WO 2020 / 128517, the entire contents of which are hereby incorporated by reference, particularly in regards to discussion in that document of movements of polynucleotides with respect to nanoreactors.

[0255] In some embodiments, the oligopeptide moves with respect to the nanopore as an electrical potential is applied across the nanopore. Often the oligopeptide is charged (e.g. negatively charged), and so applying a voltage potential across a nanopore will cause the oligopeptide to move with respect to the nanopore under the influence of the applied voltage potential. For example, if a positive voltage potential is applied to the trans side of the nanopore relative to the cis side of the nanopore, then this will induce a negatively charged oligopeptide to move from the cis side of the nanopore to the trans side of the nanopore. Similarly, if a positive voltage potential is applied to the trans side of the nanopore relative to the cis side of the nanopore then this will impede the movement of a negatively charged oligopeptide from the trans side of the nanopore to the cis side of the nanopore. The opposite will occur if a negative voltage potential is applied to the trans side of the nanopore relative to the cis side of the nanopore. Apparatuses and methods of applying appropriate voltages are described in more detail herein.

[0256] In some embodiments the chemical force is provided by a concentration (e.g. pH) gradient.

[0257] In some embodiments the movement of the oligopeptide with respect to the nanopore is controlled using a method as described in WO 2020 / 016573, the entire contents of which are incorporated herein by reference. In some embodiments the movement of the oligopeptide is controlled using a method as disclosed in any of WO 2021 / 111125, WO 2021 / 133168, or PCT / GB2023 / 052838, the entire contents of which are incorporated herein by reference.

[0258] In some embodiments the movement of the oligopeptide with respect to the nanopore is controlled using a motor protein.

[0259] In some embodiments a motor protein is present (e.g. prior to the contact of the oligopeptide with the nanopore) on the first and / or second sequencing adapter, in some embodiments the disclosed methods comprise loading a motor protein onto the first sequencing adapter and / or the second sequencing adapter if present.

[0260] In some embodiments a motor protein controls the movement of the oligopeptide in the same direction as the physical or chemical force (potential). For example, in some embodiments a positive voltage potential is applied to the trans side of the nanopore relative to the cis side of the nanopore, and a motor protein controls the movement of the oligopeptide from the cis side of the nanopore to the trans side of the nanopore. In some embodiments a positive voltage potential is applied to the cis side of the nanopore relative to the trans side of the nanopore, and a motor protein controls the movement of the oligopeptide from the trans side of the nanopore to the cis side of the nanopore.

[0261] In some embodiments a motor protein controls the movement of the oligopeptide in the opposite direction to the physical or chemical force (potential). For example, in some embodiments a positive voltage potential is applied to the trans side of the nanopore relative to the cis side of the nanopore, and the motor protein controls the movement of the oligopeptide from the trans side of the nanopore to the cis side of the nanopore. In some embodiments a positive voltage potential is applied to the cis side of the nanopore relative to the trans side of the nanopore, and the motor protein controls the movement of the oligopeptide from the cis side of the nanopore to the trans side of the nanopore.

[0262] In some embodiments the movement of the oligopeptide is driven by the motor protein in the absence of an applied potential.

[0263] In embodiments of the disclosed methods which comprise the use of a motor protein, the motor protein is typically capable of controlling the movement of the oligopeptide with respect to a nanopore. In other words, the motor protein is capable of controlling the movement of the oligopeptide.

[0264] Suitable motor proteins are in some embodiments also known as polynucleotide- handling proteins or polynucleotide-handling enzymes, or polypeptide-handling proteins or polypeptide-handling enzymes. Suitable proteins are known in the art and some exemplary motor proteins are described in more detail below.

[0265] In one embodiment, a motor protein is or is derived from a polynucleotide handling enzyme. A polynucleotide handling enzyme is a polypeptide that is capable of interacting with and modifying at least one property of a polynucleotide. The enzyme may modify the polynucleotide by cleaving it to form individual nucleotides or shorter chains of nucleotides, such as di- or trinucleotides. The enzyme may modify the polynucleotide by orienting it or moving it to a specific position.

[0266] In some embodiments, a motor protein can be present on a oligopeptide or an adapter attached thereto prior to its contact with a nanopore. For example, a motor protein can be present on a polynucleotide portion of an adapter.

[0267] In some embodiments the motor protein is designed, configured or selected to remain bound to the oligopeptide. In other words, in some embodiments the motor protein does not dissociate from the oligopeptide.

[0268] In some embodiments the polynucleotide-handling protein is modified to prevent it from disengaging from the oligopeptide (other than by passing off the end of the oligopeptide or a construct comprising the oligopeptide). Such modified polynucleotide- handling proteins are particularly suitable for use in the disclosed methods.

[0269] The motor protein can be adapted in any suitable way. For example, the motor protein can be loaded onto the oligopeptide or an adapter attached thereto and then modified in order to prevent it from disengaging. Alternatively, the motor protein can be modified to prevent it from disengaging before it is loaded onto the oligopeptide or an adapter attached thereto. Modification of a motor protein in order to prevent it from disengaging from a polypeptide can be achieved using methods known in the art, such as those discussed in WO 2014 / 013260, which is hereby incorporated by reference in its entirety, and with particular reference to passages describing the modification of polynucleotide-handling proteins (motor proteins) such as helicases in order to prevent them from disengaging with polynucleotide strands.

[0270] For example, the motor protein may have a polynucleotide-unbinding opening; e.g. a cavity, cleft or void through which a polynucleotide or polypeptide strand may pass when the motor protein disengages from the strand. In some embodiments, the polynucleotide- unbinding opening for a given motor protein (polynucleotide-handling protein) can be determined by reference to its structure, e.g. by reference to its X-ray crystal structure. The X-ray crystal structure may be obtained in the presence and / or the absence of a polynucleotide substrate. In some embodiments, the location of a polynucleotide- unbinding opening in a given motor protein may be deduced or confirmed by molecular modelling using standard packages known in the art. In some embodiments, the polynucleotide-unbinding opening may be transiently produced by movement of one or more parts e.g. one or more domains of the motor protein.

[0271] The motor protein (polynucleotide-handling protein) may be modified by closing the polynucleotide-unbinding opening. Closing the polynucleotide-unbinding opening may therefore prevent the motor protein from disengaging from the oligopeptide as well as preventing it from disengaging from an adapter attached thereto. For example, the motor protein may be modified by covalently closing the polynucleotide-unbinding opening. In some embodiments, a motor protein for addressing in this way is a helicase, as described herein. Accordingly, in some embodiments of the disclosed methods, the motor protein is modified to wholly or partially close an opening existing in at least one conformation state of the unmodified protein through which a polynucleotide or polypeptide strand can unbind.

[0272] In one embodiment, the motor protein is derived from a member of any of the Enzyme Classification (EC) groups 3.1.11, 3.1.13, 3.1.14, 3.1.15, 3.1.16, 3.1.21, 3.1.22, 3.1.25, 3.1.26, 3.1.27, 3.1.30, 3.1.31 and 3.4.21.

[0273] In some embodiments of the claimed methods, the motor protein is a helicase, a polymerase, an exonuclease, a topoisomerase, or a variant thereof.

[0274] In one embodiment, the motor protein is an exonuclease. Suitable enzymes include, but are not limited to, exonuclease I from E. coli. exonuclease III enzyme from E. coli, Red from T. thermophilus and bacteriophage lambda exonuclease, TatD exonuclease and variants thereof.

[0275] In one embodiment, the motor protein is a polymerase. The polymerase may be PyroPhage® 3173 DNA Polymerase (which is commercially available from Lucigen® Corporation), SD Polymerase (commercially available from Bioron®), Klenow from NEB or variants thereof. In one embodiment, the enzyme is Phi29 DNA polymerase or a variant thereof. Modified versions of Phi29 polymerase that may be used in the disclosed methods are disclosed in US Patent No. 5,576,204.

[0276] In some embodiments the motor protein is a polymerase, e.g. a polymerase as described herein.

[0277] In one embodiment the motor protein is a topoisomerase. In one embodiment, the topoisomerase is a member of any of the Moiety Classification (EC) groups 5.99.1.2 and 5.99.1.3. The topoisomerase may be a reverse transcriptase, which are enzymes capable of catalysing the formation of cDNA from a RNA template. They are commercially available from, for instance, New England Biolabs® and Invitrogen®.

[0278] In one embodiment the motor protein is a translocase. Examples include translocases in the FtsK and SpoIII families.

[0279] In one embodiment, the motor protein is a helicase. Any suitable helicase can be used in accordance with the methods provided herein. For example, the or each motor protein used in accordance with the present disclosure may be independently selected from a Hel308 helicase, a RecD helicase, a Tral helicase, a TrwC helicase, an XPD helicase, and a Dda helicase, or a variant thereof. Monomeric helicases may comprise several domains attached together. For instance, Tral helicases and Tral subgroup helicases may contain two RecD helicase domains, a relaxase domain and a C-terminal domain. The domains typically form a monomeric helicase that is capable of functioning without forming oligomers. Particular examples of suitable helicases include Hel308, NS3, Dda, UvrD, Rep, PcrA, Pifl and Tral. These helicases typically work on single stranded DNA. Examples of helicases that can move along both strands of a double stranded DNA include FtsK and hexameric enzyme complexes, or multisubunit complexes such as RecBCD, and are particularly suited to some embodiments disclosed herein. NS3 helicases are particularly suitable for use in the disclosed methods as they are capable of processing both DNA and RNA and so can be used in embodiments of the disclosed methods in which the target double stranded nucleic acid is a DNA-RNA hybrid.

[0280] Hel308 helicases are described in publications such as WO 2013 / 057495, the entire contents of which are incorporated by reference. RecD helicases are described in publications such as WO 2013 / 098562, the entire contents of which are incorporated by reference. XPD helicases are described in publications such as WO 2013 / 098561, the entire contents of which are incorporated by reference. Dda helicases are described in publications such as WO 2015 / 055981 and WO 2016 / 055777, the entire contents of each of which are incorporated by reference.

[0281] In some embodiments a motor protein (e.g. a helicase) can control the movement of a strand in at least two active modes of operation (when the motor protein is provided with all the necessary components to facilitate movement, e.g. fuel and cofactors such as ATP and Mg2+discussed herein) and one inactive mode of operation (when the motor protein is not provided with the necessary components to facilitate movement). When provided with all the necessary components to facilitate movement (i.e. in the active modes), the motor protein (e.g. helicase) moves along a construct comprising the oligopeptide and an adapter in a 5’ to 3’ or a 3’ to 5’ direction (depending on the motor protein). The motor protein can be used to either move the construct away from (e.g. out of) the pore (e.g. against an applied force) or the strand towards (e.g. into) the pore (e.g. with an applied force). For example, when the end of the construct towards which the motor protein moves is captured by a pore, the motor protein works against the direction of the force and pulls the threaded construct out of the pore (e.g. into the cis chamber). However, when the end away from which the motor protein moves is captured in the pore, the motor protein works with the direction of the force and pushes the threaded construct into the pore (e.g. into the trans chamber).

[0282] When the motor protein (e.g. helicase) is not provided with the necessary components to facilitate movement (i.e. in the inactive mode) it can bind to the construct and act as a brake slowing the movement of the construct when it is moved with respect to a nanopore, e.g. by being pulled into the pore by a force. In the inactive mode, it does not matter which end of the construct is captured, it is the applied force which determines the movement with respect to the pore, and the motor protein acts as a brake. When in the inactive mode, the movement control by the motor protein can be described in a number of ways including ratcheting, sliding and braking.

[0283] In another embodiment the motor protein is a protein translocase. Protein translocases are protein-binding polypeptides which are able to control movement of a protein substrate, for example an enzyme, enzyme complex, or a part of an enzyme complex that operates on a protein substrate and moves it relative to the enzyme in a processive manner, i.e. as a function of enzymatic activity.

[0284] In some embodiments the motor protein is a NTP driven unfoldase. NTP driven unfoldases are NTP-dependent enzymes that catalyze protein unfolding. NTP driven unfoldases include ATP-dependent proteases, such as proteasomal ATPases, AAA proteases, AAA+ enzymes; membrane fusion proteins, such as NSF (N-Ethylmal eimidesensitive fusion protein) / Sacl8p (N-Ethylmaleimide-sensitive fusion protein homologue in yeast) or p97 / VCP / Cdc48p (97-kDa valosin-containing protein); Pexlp and Pex6p (peroxisomal ATPase); Katanin and SKD1 (Vps4p homolog in mouse) / Vps4p (Vacuolar protein sorting 4 homolog in yeast); Dynein (motor protein); DNA replication proteins, such as ORC (origin recognition complex), Cdc6 (cell division control protein 6), MCM (minichromosome maintenance protein), DnaA, or RFC (replication factor C) / clamp- loader; RuvB (holliday junction ATP-dependent DNA helicase RuvB, EC=3.6.4.12); TIP49a / TIP49 and TIP49b / TIP48 (eukaryotic RuvB-like protein).

[0285] In some embodiments the motor protein is an AAA+ enzyme, AAA+ enzymes are members of the AAA+ superfamily of enzymes. AAA+ is an abbreviation for ATPases Associated with diverse cellular Activities. They share a common conserved module of approximately 230 amino acid residues. This is a large, functionally diverse protein family belonging to the AAA+ superfamily of ring-shaped P-loop NTPases, which exert their activity through the energy-dependent remodeling or translocation of macromolecules. Examples include ClpAP, ClpXP, ClpCP, HslYU and Lon in bacteria and their homologues in mitochondria and chloroplasts. With the exception of Lon, AAA+ enzymes (sometimes referred to as unfoldases or proteases) consist of regulatory (ATPase) and proteolytic subunits, while Lon is a single polypeptide containing both regulatory and proteolytic domains. ClpX and ClpA dock with ClpP to form ClpXP and ClpAP proteases, whereas HslU docks with HslY to form another protease, HslVU. ClpA and ClpX form hexamers, in contrast to ClpP which forms heptamers. HslU and HslY each form hexamers, although HslU heptamers have also been reported. The regulatory subunits ClpA, ClpX and HslU function as chaperones.

[0286] AAA+ enzymes may also be referred to as AAA+ molecular motors.

[0287] HslU is a member of the HsplOO and Clp family of ATPase. It can also form complex with HslY to act as an unfoldase.

[0288] Lon proteases are ATP-dependent serine peptidases belonging to the MEROPS peptidase family S16 (Ion protease family, clan SF).

[0289] In some embodiments the motor protein is ClpX or is a derivative thereof. ClpX is a member of the HSP (heat-shock protein) 100 family having the Uniprot designation clpX and having the 424 amino acid sequence given there, processed into mature form, as a subunit. ClpX subunits associate to form a six-membered (homohexameric) ring that is stabilized by binding of ATP or nonhydrolysable analogs of ATP. The N-terminal domain of ClpX is a C4-type zinc binding domain (ZBD) involved in substrate recognition. ZBD forms a very stable dimer that is essential for promoting the degradation of some typical ClpXP substrates such as and Mu A.

[0290] In some embodiments the motor protein is E. coli ClpX. E. coli ClpX generates sufficient mechanical force (>20 pN) to denature stable protein folds, and translocates along proteins at a suitable rate for primary sequence analysis by nanopore sensors (up to 80 amino acids per second). ClpX is part of the ClpXP proteasome-like complex. ClpP is composed of a diheptameric cylinder-like protease that binds at one or both ends a regulatory hexameric ATP-dependent unfoldase / translocase complex (e.g. ClpX). ClpX acts as a gate that allows for tagged proteins to enter into the inner lumen of the ClpP protease complex for subsequent degradation. The ATP-dependent unfoldase / translocase activity of the hexameric protein complex, ClpX, is employed to unfold and thread proteins through a nanopore.

[0291] In some embodiments the motor protein is a ClpX-deltaN subunits, lacking N- terminal amino acids 1-60, linked with a 20 amino acid long linker and prepared as a single polypeptide chain.

[0292] In some embodiments the motor protein is a Clp / HsplOO ATPase. Clp / HsplOO ATPases are responsible for selecting protein targets. For example, the two different bacterial ATPases ClpX and ClpA impart distinct substrate preferences to the ClpP peptidase.

[0293] In some embodiments the motor protein is a mitochondrial protein translocase. Examples include TOM or TIM from human or eukaryotic cells, such as TOMM20 (translocase of outer mitochondrial membrane homolog), TOMM22 (mitochondrial import receptor subunit 22 homolog), TOMM40 (translocase of outer mitochondrial membrane 40 homolog), T0M7 (translocase of mitochondrial outer membrane 7), T0MM7 (translocase of outer mitochondrial membrane 7 homolog), TIMM8A (translocase of inner mitochondrial membrane 8 homolog A), TIMM50 (translocase of inner mitochondrial membrane 50 homolog).

[0294] Another alternative protein translocase may be prepared from the Sec family of translocases. These include SecB (chaperone protein), SecA (ATPase), SecY (internal membrane complex in prokaryotes), SecE (interal membrane complex in prokaryotes), SecG (internal membrane complex in prokaryotes) or Sec61 (internal membrane complex in eukaryotes), SecD (membrane protein), and SecF (membrane protein).

[0295] Another alternative protein translocase is Type III Secretion System (TTS) Translocase, such as HrcN and any of the subunits of the TTS translocases, or Secindependent periplasmic protein translocase TatC.

[0296] Examples of suitable protein translocases, such as NTP driven unfoldases as described above, are described in WO 2013 / 123379, hereby incorporated by reference.

[0297] A motor protein typically requires fuel in order to handle the processing of polynucleotides and / or polypeptides. Fuel is typically free nucleotides or free nucleotide analogues. The free nucleotides may be one or more of, but are not limited to, adenosine monophosphate (AMP), adenosine diphosphate (ADP), adenosine triphosphate (ATP), guanosine monophosphate (GMP), guanosine diphosphate (GDP), guanosine triphosphate (GTP), thymidine monophosphate (TMP), thymidine diphosphate (TDP), thymidine triphosphate (TTP), uridine monophosphate (UMP), uridine diphosphate (UDP), uridine triphosphate (UTP), cytidine monophosphate (CMP), cytidine diphosphate (CDP), cytidine triphosphate (CTP), cyclic adenosine monophosphate (cAMP), cyclic guanosine monophosphate (cGMP), deoxyadenosine monophosphate (dAMP), deoxyadenosine diphosphate (dADP), deoxyadenosine triphosphate (dATP), deoxyguanosine monophosphate (dGMP), deoxyguanosine diphosphate (dGDP), deoxyguanosine triphosphate (dGTP), deoxythymidine monophosphate (dTMP), deoxythymidine diphosphate (dTDP), deoxythymidine triphosphate (dTTP), deoxyuridine monophosphate (dUMP), deoxyuridine diphosphate (dUDP), deoxyuridine triphosphate (dUTP), deoxycytidine monophosphate (dCMP), deoxycytidine diphosphate (dCDP) and deoxycytidine triphosphate (dCTP). The free nucleotides are usually selected from AMP, TMP, GMP, CMP, UMP, dAMP, dTMP, dGMP or dCMP. The free nucleotides are typically adenosine triphosphate (ATP).

[0298] A cofactor for the motor protein is a factor that allows the motor protein to function. The cofactor is often a divalent metal cation. The divalent metal cation is often Mg2+, Mn2+, Ca2+or Co2+. The cofactor is most typically Mg2+.

[0299] Detector

[0300] Embodiments described herein refer to movement of a oligopeptide with respect to a nanopore. However, whilst the disclosure provides nanopores as exemplary detectors, the methods provided herein are also amenable to other detectors including (i) a zero-mode waveguide, (ii) a field-effect transistor, optionally a nanowire field-effect transistor; (iii) an AFM tip; (iv) a nanotube, optionally a carbon nanotube and (v) a nanopore. The disclosed methods are particularly amenable to methods in which a polypeptide is moved through a detector or through a structure containing a detector, e.g. a well in a detector chip.

[0301] Nanopore

[0302] As explained above, in some embodiments the disclosed methods comprise taking one or more measurements as a oligopeptide moves with respect to a nanopore. The oligopeptide is typically comprised in an oligopeptide-adapter adduct as described herein.

[0303] Accordingly, in some embodiments the disclosed methods comprise carrying out a method of preparing a oligopeptide-adapter adduct as described herein, and contacting the oligopeptide with a nanopore under conditions such that the oligopeptide moves with respect to the nanopore; and

[0304] - taking one or more measurements characteristic of the oligopeptide as the oligopeptide moves with respect to then nanopore; thereby characterising the oligopeptide.

[0305] In the disclosed methods, any suitable nanopore can be used. In one embodiment a nanopore is a transmembrane pore.

[0306] A transmembrane pore is a structure that crosses the membrane to some degree. It permits hydrated ions driven by an applied potential to flow across or within the membrane. The transmembrane pore typically crosses the entire membrane so that hydrated ions may flow from one side of the membrane to the other side of the membrane. However, the transmembrane pore does not have to cross the membrane. It may be closed at one end. For instance, the pore may be a well, gap, channel, trench or slit in the membrane along which or into which hydrated ions may flow.

[0307] Any suitable transmembrane pore may be used in the methods provided herein. The pore may be biological or artificial. Suitable pores include, but are not limited to, protein pores, polynucleotide pores, and solid state pores.

[0308] A solid state pore may, in one embodiment, comprise a nanochannel. In some embodiments the solid state pore is a pore disclosed in WO 2003 / 003446, WO 2009 / 020682 or WO 2016 / 187519, each of which is incorporated by reference in their entirety.

[0309] In one embodiment, the pore may be a DNA origami pore (Langecker et al.. Science, 2012; 338: 932-936). Suitable DNA origami pores are disclosed in WO2013 / 083983, WO 2018 / 011603 and WO 2020 / 025974, each of which is incorporated by reference in their entirety.

[0310] In one embodiment, the nanopore is a scaffolded polypeptide nanopore. In some embodiments the pore is a scaffolded polypeptide nanopore as disclosed in WO 2020 / 025909 or WO 2020 / 074399, each of which is incorporated by reference in their entirety.

[0311] In one embodiment, the nanopore is a transmembrane protein pore. A transmembrane protein pore is a polypeptide or a collection of polypeptides that permits hydrated ions, such as polynucleotides, to flow from one side of a membrane to the other side of the membrane. In the methods provided herein, the transmembrane protein pore is capable of forming a pore that permits hydrated ions driven by an applied potential to flow from one side of the membrane to the other. The transmembrane protein pore typically permits polynucleotides and polypeptides to flow from one side of the membrane, such as a polymer membrane, to the other. The transmembrane protein pore allows a polynucleotide or polypeptide to be moved through the pore.

[0312] Examples of transmembrane protein pores include Wza, Iota toxin, Anthrax protective antigen, Vibrio cholerae cytolysin, Cytotoxin K (CytK), CELIII, CsgG, CsgF, CsgG-CsgF, Aerolysin, alpha hemolysin, MspA, MspB, MspC, PorARr, PorBRr, PorARc, PilQ, necrotic enteritis B-like toxin (NetB), FraC, portal proteins including G20c, P23 45, T4, SPP1, P22 and Phi29, gamma hemolysin, Monalysin, Lysenin, ClyA, an actinoporin, Clostridium perfringens beta toxin, parasporin-2, epsilon toxin, lectin from the parasitic mushroom Laetiporus sulphureus (LSL), volvatoxin, Cry toxins, Cytl Aa, Cyt2Aa, Complement component 9 (C9), Perfringolysin O, Pleurotolysin, Listeriolysin, Perforin-2, Gasdermin-A3, L-, P- and M-ring protein, Type II secretion system protein D, GspD, InvG, VirB7, SpoIIIAG, Cag8, Cag3, Cag or other proteins in the Type IV secretion system apparatus protein CagY, WzzB, Pentraxin, Afp2, Major vault protein, Thioredoxindependent peroxidase reductase, Arf-GAP, Respiratory syncytial virus ribonucleoprotein, Chikungunya virus nonstructural protein 1, PRC, YaxA, XaxA, HfaB, NfpAB, leukocidin and PrgH.

[0313] In one embodiment, the nanopore is a transmembrane protein pore which is a monomer or an oligomer. The pore is typically made up of several repeating subunits, such as at least 6, at least 7, at least 8, at least 9, at least 10, at least 11, at least 12, at least 13, at least 14, at least 15, or at least 16 subunits. The pore is typically a hexameric, heptameric, octameric or nonameric pore. The pore may be a homo-oligomer or a heterooligomer.

[0314] In one embodiment, the transmembrane protein pore comprises a barrel or channel through which the ions may flow. The subunits of the pore typically surround a central axis and contribute strands to a transmembrane P-barrel or channel or a transmembrane a- helix bundle or channel.

[0315] Typically, the barrel or channel of the transmembrane protein pore comprises amino acids that facilitate interaction with an analyte, such as a target polypeptide (as described herein). These amino acids are typically located near a constriction of the barrel or channel. The transmembrane protein pore typically comprises one or more positively charged amino acids, such as arginine, lysine or histidine, or aromatic amino acids, such as tyrosine or tryptophan. These amino acids typically facilitate the interaction between the pore and nucleotides, polynucleotides, nucleic acids and polypeptides.

[0316] In one embodiment, the nanopore is a transmembrane protein pore derived from 0- barrel pores or a-helix bundle pores. 0-barrel pores comprise a barrel or channel that is formed from 0-strands. Suitable 0-barrel pores include, but are not limited to, 0-toxins, such as a-hemolysin, anthrax toxin, CytK, aerolysin and leukocidins, and outer membrane proteins / porins of bacteria, such as Mycobacterium smegmatis porin (Msp), for example MspA, MspB, MspC or MspD, CsgG, outer membrane porin F (OmpF), outer membrane porin G (OmpG), outer membrane phospholipase A and Neisseria autotransporter lipoprotein (NalP) and other pores, such as lysenin. a-helix bundle pores comprise a barrel or channel that is formed from a-helices. Suitable a-helix bundle pores include, but are not limited to, inner membrane proteins and a outer membrane proteins, such as WZA, FraC and ClyA toxin.

[0317] In one embodiment the nanopore is a transmembrane pore derived from or based on Msp, a-hemolysin (a-HL), lysenin, CsgG, ClyA, Spl or haemolytic protein fragaceatoxin C (FraC).

[0318] In one embodiment, the nanopore is a transmembrane protein pore derived from CsgG, e.g. from CsgG from E. coli Str. K-12 substr. MC4100. Such a pore is oligomeric and typically comprises 7, 8, 9 or 10 monomers derived from CsgG. The pore may be a homo-oligomeric pore derived from CsgG comprising identical monomers. Alternatively, the pore may be a hetero-oligomeric pore derived from CsgG comprising at least one monomer that differs from the others. Examples of suitable pores derived from CsgG are disclosed in WO 2016 / 034591, WO 2017 / 149316, WO 2017 / 149317, WO 2017 / 149318 and WO 2019 / 002893, each of which is hereby incorporated by reference in its entirety.

[0319] In one embodiment, the nanopore is a transmembrane pore derived from lysenin. Examples of suitable pores derived from lysenin are disclosed in WO 2013 / 153359, which is hereby incorporated by reference in its entirety.

[0320] In one embodiment, the nanopore is a transmembrane pore derived from or based on a-hemolysin (a-HL). The wild type a-hemolysin pore is formed of 7 identical monomers or sub-units (i.e., it is heptameric). An a-hemolysin pore may be a-hemolysin- NN or a variant thereof. The variant typically comprises N residues at positions El 11 and K147.

[0321] In one embodiment, the nanopore is a transmembrane protein pore derived from Msp, e.g. from MspA. Examples of suitable pores derived from MspA are disclosed in WO 2012 / 107778.

[0322] In one embodiment, the nanopore is a transmembrane pore derived from or based on ClyA. Examples of suitable pores derived from ClyA are disclosed in Soskine et al., Nano Letters 2012 12 (9), 4895-4900; WO 2014 / 153625; and WO 2017 / 098322, each of which is hereby incorporated by reference.

[0323] In one embodiment, the nanopore is a transmembrane pore derived from Phi29. Examples of suitable pores derived from Phi29 are disclosed in Wendell et al., Nature Nanotech 4, 765-772 (2009), WO 2010 / 062697, WO 2019 / 157365 and WO 2019 / 157424, each of which is hereby incorporated by reference.

[0324] In some embodiments the nanopore is selected from M-ring protein, perforin-2, PlyAB (pleurotolysin), SpoIIIAG, VirB7, Type II secretion system protein D, GspD, InvG, PilQ, pentraxin, and portal proteins including T4, T7, P23 45, G20c and Phi29 nanopores.

[0325] In one embodiment, the nanopore is a transmembrane pore derived from or based on a Rhodococcus species of bacteria, for example Rhodococcus corynebacteroides or Rhodococcus ruber, for example PorARr, PorBRr or PorARc. Examples of such pores are described in Piselli et al., Eur Biophys J 51, 309-323 (2022), and in WO 2024 / 089270 hereby incorporated by reference.

[0326] As explained above, in some embodiments the nanopore comprises a constriction. The constriction is typically a narrowing in the channel which runs through the nanopore which may determine or control the signal obtained when the conjugate moves with respect to the nanopore. As used herein, both protein and solid state nanopores may comprise a “constriction”.

[0327] In some embodiments the nanopore is designed, modified or chosen to have a constriction that is sized according to the diameter of the construct. In some embodiments the pore has a constriction having a diameter of at least 1 nm, e.g. at least 1.5 nm, such as at least 2 nm, e.g. at least 2.5 nm e.g. at least 3 nm. In some embodiments the pore has a constriction having a diameter of from about 1.5 to about 2.5 nm.

[0328] Tags In some embodiments of the methods provided herein, a tag on the nanopore can be used, e.g. to promote the capture of the oligopeptide-adapter adduct.

[0329] The interaction between a tag on a nanopore and a binding site on a construct or barcode (which may for example be a binding site present in the polynucleotide portion of an adaptor attached to the oligopeptide) may be reversible. For example, a polynucleotide can bind to a tag on a nanopore, e.g., via its adaptor, and release at some point, e.g., during characterization of the oligopeptide-adapter adduct by the nanopore and / or during processing by a motor protein. A strong non-covalent bond (e.g., biotin / avidin) is still reversible and can be useful in some embodiments of the methods described herein. For example, a pair of pore tag and polynucleotide adaptor can be designed to provide a sufficient interaction between the complement of a double stranded polynucleotide (or a portion of an adaptor that is attached to the complement) and the nanopore such that the complement is held close to the nanopore (without detaching from the nanopore and diffusing away) but is able to release from the nanopore as it is processed.

[0330] A pore tag and polynucleotide adaptor can be configured such that the binding strength or affinity of a binding site on the polynucleotide (e.g., a binding site provided by an anchor or a leader sequence of an adaptor or by a capture sequence within the duplex stem of an adaptor) to a tag on a nanopore is sufficient to maintain the coupling between the nanopore and polynucleotide until an applied force is placed on it to release the bound polynucleotide from the nanopore.

[0331] In some embodiments, the tags or tethers are uncharged. This can ensure that the tags or tethers are not drawn into the nanopore under the influence of a potential difference if present.

[0332] One or more molecules that attract or bind the barcode or adapter attached thereto may be linked to the nanopore. Any molecule that hybridizes to the conjugate, adaptor and / or polynucleotide may be used. The molecule attached to the pore may be selected from a PNA tag, a PEG linker, a short oligonucleotide, a positively charged amino acid and an aptamer. Pores having such molecules linked to them are known in the art. For example, pores having short oligonucleotides attached thereto are disclosed in Howarka et al (2001) Nature Biotech. 19: 636-639 and WO 2010 / 086620, and pores comprising PEG attached within the lumen of the pore are disclosed in Howarka et al (2000) J. Am. Chem. Soc. 122(11): 2411-2416.

[0333] A short oligonucleotide attached to the nanopore, which comprises a sequence complementary to a sequence in the conjugate (e.g. in a leader sequence or another single stranded sequence in an adaptor) may be used to enhance capture of the barcode or adapter attached thereto in the methods described herein.

[0334] Membrane

[0335] Typically, in the disclosed methods, the nanopore is typically present in a membrane. Any suitable membrane may be used in the system.

[0336] The membrane is typically an amphiphilic layer. An amphiphilic layer is a layer formed from amphiphilic molecules, such as phospholipids, which have both hydrophilic and lipophilic properties. The amphiphilic molecules may be synthetic or naturally occurring. Non-naturally occurring amphiphiles and amphiphiles which form a monolayer are known in the art and include, for example, block copolymers (Gonzalez-Perez et al., Langmuir, 2009, 25, 10447-10450). Block copolymers are polymeric materials in which two or more monomer sub-units that are polymerized together to create a single polymer chain. Block copolymers typically have properties that are contributed by each monomer sub-unit. However, a block copolymer may have unique properties that polymers formed from the individual sub-units do not possess. Block copolymers can be engineered such that one of the monomer sub-units is hydrophobic (i.e. lipophilic), whilst the other subunits) are hydrophilic whilst in aqueous media. In this case, the block copolymer may possess amphiphilic properties and may form a structure that mimics a biological membrane. The block copolymer may be a diblock (consisting of two monomer subunits), but may also be constructed from more than two monomer sub-units to form more complex arrangements that behave as amphipiles. The copolymer may be a triblock, tetrablock or pentablock copolymer.

[0337] In some embodiments, the membrane is one of the membranes disclosed in International Application No. WO2014 / 064443 or WO2014 / 064444.

[0338] The amphiphilic molecules may be chemically-modified or functionalised to facilitate coupling of the polynucleotide. The amphiphilic layer may be a monolayer or a bilayer. The amphiphilic layer is typically planar. The amphiphilic layer may be curved. The amphiphilic layer may be supported.

[0339] Amphiphilic membranes are typically naturally mobile, essentially acting as two dimensional fluids with lipid diffusion rates of approximately 10'8cm s'1. This means that the pore and coupled polynucleotide can typically move within an amphiphilic membrane.

[0340] The membrane may be a lipid bilayer. Lipid bilayers are models of cell membranes and serve as excellent platforms for a range of experimental studies. For example, lipid bilayers can be used for in vitro investigation of membrane proteins by single-channel recording. Alternatively, lipid bilayers can be used as biosensors to detect the presence of a range of substances. The lipid bilayer may be any lipid bilayer. Suitable lipid bilayers include, but are not limited to, a planar lipid bilayer, a supported bilayer or a liposome. The lipid bilayer is typically a planar lipid bilayer. Suitable lipid bilayers are disclosed in WO 2008 / 102121, WO 2009 / 077734 and WO 2006 / 100484.

[0341] In another embodiment, the membrane comprises a solid state layer. Solid state layers can be formed from both organic and inorganic materials including, but not limited to, microelectronic materials, insulating materials such as SislS , AI2O3, and SiO, organic and inorganic polymers such as polyamide, plastics such as Teflon® or elastomers such as two-component addition-cure silicone rubber, and glasses. The solid state layer may be formed from graphene. Suitable graphene layers are disclosed in WO 2009 / 035647. If the membrane comprises a solid state layer, the pore is typically present in an amphiphilic membrane or layer contained within the solid state layer, for instance within a hole, well, gap, channel, trench or slit within the solid state layer. The skilled person can prepare suitable solid state / amphiphilic hybrid systems. Suitable systems are disclosed in WO 2009 / 020682 and WO 2012 / 005857. Any of the amphiphilic membranes or layers discussed above may be used.

[0342] The methods disclosed herein are typically carried out using (i) an artificial amphiphilic layer comprising a pore, (ii) an isolated, naturally-occurring lipid bilayer comprising a pore, or (iii) a cell having a pore inserted therein. The methods are typically carried out using an artificial amphiphilic layer, such as a di- or tri-block copolymer layer. The layer may comprise other transmembrane and / or intramembrane proteins as well as other molecules in addition to the pore. Suitable apparatus and conditions are discussed below. The disclosed methods are typically carried out in vitro.

[0343] Characterisation

[0344] As discussed in more detail herein, in some embodiments the disclosed methods comprise taking one or more measurements as a oligopeptide moves with respect to a nanopore. The oligopeptide is typically comprised in an oligopeptide-adapter adduct as described herein. In some embodiments the disclosed methods comprise carrying out a method of preparing a oligopeptide-adapter adduct as described herein, and contacting the oligopeptide with a nanopore under conditions such that the oligopeptide moves with respect to the nanopore; and

[0345] - taking one or more measurements characteristic of the oligopeptide as the oligopeptide moves with respect to then nanopore; thereby characterising the oligopeptide.

[0346] Many different characteristics of the oligopeptide can be determined. Any suitable measurements can be taken.

[0347] For example, in some embodiments characterising the oligopeptide comprises determining (i) the length of the oligopeptide, (ii) the identity of the oligopeptide, (iii) the sequence of the oligopeptide, (iv) the secondary structure of the oligopeptide; (v) whether or not and / or to the extent to which the oligopeptide is modified, e.g. by one or more post- translational modifications.; (vi) the presence, absence, concentration or relative abundance of the oligopeptide in a sample comprising multiple oligopeptides (e.g. derived from a sample comprising multiple polypeptides). In typical embodiments the measurements are characteristic of the sequence of the oligopeptide or whether or not the oligopeptide is modified, In some embodiments the measurements are characteristics of the sequence of the oligopeptide.

[0348] Because the oligopeptide is derived from the polypeptide, the characteristics of the oligopeptide can inform on the characteristics of the parent polypeptide. Thus, the disclosed methods can be used to inform regarding

[0349] (i) the length of the polypeptide, (ii) the identity of the polypeptide, (iii) the sequence of the polypeptide, (iv) the secondary structure of the polypeptide; (v) whether or not and / or to the extent to which the polypeptide is modified, e.g. by one or more post- translational modifications.; (vi) the presence, absence, concentration or relative abundance of the polypeptide in a sample comprising multiple polypeptides.

[0350] Conditions

[0351] The disclosed methods may be carried out using any apparatus that is suitable for investigating a membrane / pore system in which a pore is inserted into a membrane. The characterisation method may be carried out using any apparatus that is suitable for transmembrane pore sensing. For example, the apparatus may comprise a chamber comprising an aqueous solution and a barrier that separates the chamber into two sections. The barrier may have an aperture in which a membrane containing a transmembrane pore is formed. Transmembrane pores are described herein. The characterisation methods may be carried out using the apparatus described in WO 2008 / 102120, WO 2010 / 122293 or WO 00 / 28312.

[0352] The characterisation methods may comprise optical measurements, for example such as described in WO 2016 / 009180 and WO 2021 / 198695.

[0353] The characterisation methods may involve measuring the ion current flow through the pore, typically by measurement of a current. Alternatively, the ion flow through the pore may be measured optically, such as disclosed by Heron et al: J. Am. Chem. Soc. 9 Vol. 131, No. 5, 2009. Therefore the apparatus may also comprise an electrical circuit capable of applying a potential and measuring an electrical signal across the membrane and pore. The characterisation methods may be carried out using a patch clamp or a voltage clamp. The characterisation methods typically involve the use of a voltage clamp.

[0354] The characterisation methods may be carried out on a silicon-based array of wells where each array comprises 128, 256, 512, 1024, 2000, 3000, 4000, 6000, 10000, 12000, 15000 or more wells.

[0355] The characterisation methods may involve the measuring of a current flowing through the pore. The method is typically carried out with a voltage applied across the membrane and pore. The voltage used is typically from +2 V to -2 V, typically -400 mV to +400mV. The voltage used is typically in a range having a lower limit selected from -400 mV, -300 mV, -200 mV, -150 mV, -100 mV, -50 mV, -20mV and 0 mV and an upper limit independently selected from +10 mV, + 20 mV, +50 mV, +100 mV, +150 mV, +200 mV, +300 mV and +400 mV. The voltage used is more typically in the range 100 mV to 240mV and most typically in the range of 120 mV to 220 mV. It is possible to increase discrimination between different nucleotides by a pore by using an increased applied potential.

[0356] The characterisation methods are typically carried out in the presence of any charge carriers, such as metal salts, for example alkali metal salts, halide salts, for example chloride salts, such as alkali metal chloride salt. Charge carriers may include ionic liquids or organic salts, for example tetramethyl ammonium chloride, trimethylphenyl ammonium chloride, phenyltrimethyl ammonium chloride, or l-ethyl-3 -methyl imidazolium chloride. In the exemplary apparatus discussed above, the salt is present in the aqueous solution in the chamber. Potassium chloride (KC1), sodium chloride (NaCl) or caesium chloride (CsCl) is typically used. KC1 is typical. The salt may be an alkaline earth metal salt such as calcium chloride (CaCh). The salt concentration may be at saturation. The salt concentration may be 3M or lower and is typically from 0.1 to 2.5 M, from 0.3 to 1.9 M, from 0.5 to 1.8 M, from 0.7 to 1.7 M, from 0.9 to 1.6 M or from 1 M to 1.4 M. The salt concentration is typically from 150 mM to 1 M. The characterisation method may be carried out using a salt concentration of at least 0.3 M, such as at least 0.4 M, at least 0.5 M, at least 0.6 M, at least 0.8 M, at least 1.0 M, at least 1.5 M, at least 2.0 M, at least 2.5 M or at least 3.0 M. High salt concentrations provide a high signal to noise ratio and allow for currents indicative of binding / no binding to be identified against the background of normal current fluctuations.

[0357] The characterisation methods are typically carried out in the presence of a buffer. In the exemplary apparatus discussed above, the buffer is present in the aqueous solution in the chamber. Any suitable buffer may be used. Typically, the buffer is HEPES. Another suitable buffer is Tris-HCl buffer. The methods are typically carried out at a pH of from 4.0 to 12.0, from 4.5 to 10.0, from 5.0 to 9.0, from 5.5 to 8.8, from 6.0 to 8.7 or from 7.0 to 8.8 or 7.5 to 8.5. The pH used may be about 7.5.

[0358] The characterisation methods may be carried out at from 0 °C to 100 °C, from 15 °C to 95 °C, from 16 °C to 90 °C, from 17 °C to 85 °C, from 18 °C to 80 °C, 19 °C to 70 °C, or from 20 °C to 60 °C. The characterisation methods are typically carried out at room temperature. The characterisation methods are optionally carried out at a temperature that supports enzyme function, such as about 37 °C.

[0359] Further aspects

[0360] In one embodiment, also provided is a kit for modifying a polypeptide, comprising a first sequencing adapter capable of selectively reacting with an optionally- functionalised first reactive amino acid at the C-terminus of an oligopeptide generated by contacting a polypeptide with a proteinase; and a second sequencing adapter capable of reacting with a second reactive functional group at the N-terminus of said oligopeptide.

[0361] Typically the first and second sequencing adapter are each as described herein.

[0362] In some embodiments the kit further comprises a proteinase capable of selectively cleaving a polypeptide at a cleavage site thereby forming one or more oligopeptides each comprising a first end terminating in a first amino acid motif comprising a first reactive amino acid, and a second end terminating in a second amino acid. In some embodiments the proteinase is as described herein.

[0363] In some embodiments the kit further comprises a compound comprising a oligo- reactive functional group and an adapter-reactive functional group, wherein the compound is capable of reacting with the first reactive amino acid thereby functionalising the first reactive amino acid. In some embodiments the compound (and the reactive functional groups comprised therein) is as described herein.

[0364] In some embodiments the kit further comprises a handle (e.g. a peptide handle) capable of being ligated onto the N-terminus of said oligopeptide. In some embodiments the handle is a handle as described in more detail herein.

[0365] In some embodiments the kit further comprises a ligase capable of ligating a peptide handle onto the N-terminus of said oligopeptide. In some embodiments the ligase is as described herein.

[0366] In some embodiments the kit further comprises a motor protein capable of controlling the movement of the oligopeptide with respect to a nanopore. In some embodiments the motor protein is as described herein.

[0367] In some embodiments the kit further comprises a nanopore. In some embodiments the nanopore is present in an array comprising a plurality of nanopores in a membrane. In some embodiments the nanopore is as described herein.

[0368] The kit may comprise instructions for preparing oligopeptide-adapter adducts as described herein.

[0369] Also provided is a oligopeptide-adapter adduct as described herein. The oligopeptide-adapter adduct may be produced in accordance with the disclosed methods.

[0370] Also provided is a library of oligopeptide-adapter adducts, wherein each oligopeptide-adapter adduct is generated from a polypeptide. The library may be generated in accordance with the methods described herein.

[0371] Also provided is a system, comprising a library as described herein, and a nanopore. In some embodiments the nanopore is as described herein. In some embodiments the system further comprises a motor protein capable of controlling the movement of the oligopeptides in the library with respect to the nanopore. In some embodiments the system comprises computing means configured to detect information characteristic of the oligopeptides in the library and to selectively process the signal obtained as said oligopeptides move with respect to the nanopore. In some embodiments the system comprises receiving means for receiving data from detection of the oligopeptides, processing means for processing the signal obtained as the oligopeptides moves with respect to the nanopore, and output means for outputting the characterisation information thus obtained. Exemplary workflows

[0372] The following workflows are provided to illustrate the methods described herein.

[0373] One embodiment of the disclosed methods is illustrated in Figure 1. This workflow provides an example of the disclosed methods in which reacting the N-terminal amine group with an amine-reactive moiety comprises attaching a first sequencing adapter to the first reactive amino acid and to the N-terminal amine group; the method comprises cleaving the amino acid comprising the reacted N-terminal amine group from the oligopeptide-adapter adduct; and cleaving the reacted N-terminal amine group from the oligopeptide-adapter adduct generates a second reactive functional group capable of reacting with a complementary functional group on a second sequencing adapter.

[0374] With reference to Figure 1, a larger protein is first digested into oligopeptides using a proteinase such as LysC. When LysC is used as the proteinase, the resultant fragments possess reactive lysine residues at their C-termini, and free amino groups at their N- termini.

[0375] The method may comprise the simultaneous modification of the N-termini and the reactive C-terminal lysine. For example, an isothiocyanate (e.g. ethynyl phenyl isothiocyanate, EPITC, or sulfophenyl isothiocyanate, SPITC) can be used to functionalise the amine groups at the N-terminal and in the lysine residue.

[0376] The method may then comprise the selective deprotection of the N-terminus and concomitant loss of a single N-terminal residue, for example by treatment with acid (e.g. trifluoroacetic acid), or with an edmanase enzyme. Additionally, an acidic buffer such as citric acid, may also be able to catalyse these reactions.

[0377] The method may then comprise the attachment of a first sequencing adapter (e.g. a DNA sequencing adapter) to the C-terminal lysine residue. Following this, the newly liberated N-terminus may have a second sequencing adapter attached thereto. Any of the disclosed amine-targeting chemistries disclosed herein can be used at this step.

[0378] One embodiment of the disclosed methods is illustrated in Figure 2. This workflow provides an example of the disclosed methods in which reacting the N-terminal amine group with an amine-reactive moiety comprises immobilizing the oligopeptide on an amine-reactive substrate; the method comprises cleaving or eluting the oligopeptide from the substrate; and said cleaving or eluting generates a second reactive functional group comprised in said oligopeptide capable of reacting with a complementary functional group on a second sequencing adapter. With reference to Figure 2, a larger protein may first be digested into oligopeptides using a proteinase such as LysC. When LysC is used as the proteinase, the resultant fragments possess reactive lysine residues at their C-termini, and free amino groups at their N-termini.

[0379] The method may comprise immobilizing the oligopeptide onto an amine-reactive substrate (e.g. a resin or bead) under conditions such that the N-terminal amine group binds to the substrate. A suitable example is a 2-PCA-modified resin

[0380] The method may then comprise modification of the C-terminal lysine residue and attachment of a first sequencing adapter. For example, the C-terminal lysine may be activated using an amine-reactive compound comprising a click chemistry group, such as . ethynyl phenyl isothiocyanate, EPITC. Alternatively the first sequencing adapter can be directly reacted with the amine group of the C-terminal lysine.

[0381] The method may then comprise releasing the oligopeptide-adapter adduct from the substrate, e.g. by treatment with hydrazine. Following this, the newly liberated N-terminus may have a second sequencing adapter attached thereto. Any of the disclosed amine- targeting chemistries disclosed herein can be used at this step.

[0382] In another iteration of this workflow, an amine-reactive group such as 2-PCA or a derivative thereof may be first coupled to a purification tag such as biotin. The aminereactive portion of the molecule can be reacted with the N-termini of the oligopeptide, and the purification tag subsequently used to immobilize the oligopeptide peptides onto a substrate. Once immobilized, the workflow proceeds as above.

[0383] Another exemplary workflow is disclosed in which reacting the N-terminal amine group with an amine-reactive moiety comprises immobilizing the oligopeptide on an amine-reactive substrate and the method comprises cleaving or eluting the oligopeptide from the substrate.

[0384] A larger protein is first digested into oligopeptides using a proteinase such as LysC or ArgC. When LysC is used as the proteinase, the resultant fragments possess reactive lysine residues at their C-termini, and free amino groups at their N-termini. When ArgC is used, the resultant fragments possess reactive arginine residues at their C-termini, and free amino groups at their N-termini.

[0385] The method may comprise reacting:

[0386] (i) reacting the N-terminal amino group with a second sequencing adapter attached to a purification tag, such as desthiobiotin or a poly-His. The attachment chemistry between the N-terminal amino group and the second sequencing adapter is not especially limited and any of the suitable chemistry described herein could be used. In some embodiments the second sequencing adapter is attached to the N-terminal amine group via 2-PCA (e.g. 2-PCA maleimide) chemistry as described herein;

[0387] (ii) modification of the first reactive group at the C-terminus of the oligopeptides by attachment of a first amino acid. Again, the attachment chemistry between the C-terminal reactive amino acid and the first sequencing adapter is not especially limited and any of the suitable chemistry described herein could be used; and

[0388] (iii) immobilizing the purification tag on a complementary purification substrate (e.g. a resin or bead) under conditions such that the purification tag binds to the substrate.

[0389] These steps may be done simultaneously or sequentially in any order.

[0390] The construct may then be released from the purification substrate. In some embodiments releasing the construct from the purification matrix comprises eluting the purification tag from the substrate.

[0391] After releasing the construct from the purification matrix the construct may be characterised as described herein.

[0392] In one embodiment of this workflow, the N-terminal amine group of the oligopeptide generated by digestion of the larger protein using LysC is reacted with a 2- PCA derivative comprising a 2-PCA moiety attached to a purification tag such as a poly- His or desthiobiotin tag; wherein the 2-PCA derivative has the structure

[0393] In some embodiments maleimide is used to stabilize the reaction of the 2-PCA moiety with the oligopeptide as described herein. As LysC is used to digest the larger protein and the C-terminal amino acid of each oligopeptide is lysine. The lysine at the C-terminal of each oligopeptide may be functionalised by reaction with 2-azidoacetic acid, thereby generating a reactive functional group comprising an azide group at the C-terminus of each oligopeptide. The method may involve reacting the azide group with a first sequencing adapter comprising a complementary BCN functional group for attachment to the azide at the C-terminus of the oligopeptide.

[0394] One embodiment of the disclosed methods is illustrated in Figure 3. This workflow provides an example of the disclosed methods in which reacting the N-terminal amine group with an amine-reactive moiety comprises reacting the oligopeptide with a ligase and a peptide handle comprising a purification tag, under conditions such that the ligase ligates the peptide handle to the N-terminus of the oligopeptide; the method comprises contacting the purification tag with a complementary substrate under conditions such that the fusion adduct is immobilised on the substrate; the method further comprises cleaving or eluting a portion of the fusion adduct comprising the oligopeptide from the substrate adduct thereby generating a second reactive functional group capable of reacting with a complementary functional group on a second sequencing adapter; and attaching a second sequencing adapter to the portion after the portion has been cleaved from the immobilised fusion adduct.

[0395] With reference to Figure 3, a larger protein may first be digested into smaller fragments e.g. using LysC . The resultant fragments would possess lysine residues at their C-termini and free N-termini.

[0396] After the protein is digested, the resulting oligopeptide fragments may be modified by attachment of a peptide handle using a ligase, thereby forming a fusion adduct. The peptide handle may comprise a purification tag for immobilization of the fusion adduct onto a purification substrate such as a resin or bead. The method may involve contacting the adduct with a purification substrate under conditions such that the purification tag binds to the substrate.

[0397] The method may then comprise modification of the C-terminal lysine residue and attachment of a first sequencing adapter. For example, the C-terminal lysine may be activated using an amine-reactive compound comprising a click chemistry group, such as ethynyl phenyl isothiocyanate, EPITC. Alternatively the first sequencing adapter can be directly reacted with the amine group of the C-terminal lysine.

[0398] Following reaction of the C-terminal lysine, the oligopeptide-adapter adduct may be eluted from the purification resin liberating the free N-terminal amine group. Following this, the newly liberated N-terminus may have a second sequencing adapter attached thereto. Any of the disclosed amine-targeting chemistries disclosed herein can be used at this step. One embodiment of the disclosed methods is illustrated in Figure 4. This workflow provides an example of the disclosed methods in which reacting the N-terminal amine group with an amine-reactive moiety comprises reacting the oligopeptide with a peptide handle comprising a purification tag and a second reactive functional group capable of reacting with a complementary functional group on a second sequencing adapter under conditions such that the handle is attached to the N-terminus of the oligopeptide; thereby generating a fusion adduct; the method comprises contacting the purification tag with a complementary substrate under conditions such that the fusion adduct is immobilised on the substrate; the method further comprises cleaving or eluting a portion of the fusion adduct comprising the oligopeptide from the substrate; and attaching a second sequencing adapter to the portion.

[0399] With reference to Figure 4, a larger protein may first be digested into smaller fragments e.g. using LysC . The resultant fragments would possess lysine residues at their C-termini and free N-termini.

[0400] After the protein is digested, the resulting oligopeptide fragments may be modified by attachment of a peptide handle using a ligase, thereby forming a fusion adduct. The peptide handle may comprise a purification tag for immobilization of the fusion adduct onto a purification substrate such as a resin or bead under conditions such that the purification tag binds to the substrate; and may comprise a pendant reactive functional group for attachment of a second sequencing adapter. The pendant reactive functional group may be attached to the purification tag via a cleavage site, such as a protease recognition site.

[0401] Typically, the second functional group is orthogonal to the amine group of the C- terminal lysine. Alternatively, prior to ligation of the peptide handle, the C-terminal lysine may be functionalised to provide an orthogonal reactive group to the reactive group comprised in the peptide handle.

[0402] The method then comprises:

[0403] (i) modification of the C-terminal lysine residue and attachment of a first sequencing adapter. For example, the C-terminal lysine may be activated using an amine-reactive compound comprising a click chemistry group, such as ethynyl phenyl isothiocyanate, EPITC. Alternatively the first sequencing adapter can be directly reacted with the amine group of the C- terminal lysine; (ii) modification of the reactive functional group comprised in the peptide handle with a second sequencing adapter; and

[0404] (iii) cleavage of the cleavage site thereby liberating the oligopeptide-adapter adduct from the resin.

[0405] These steps may be conducted in any order. Thus, in some embodiments the C- terminal sequencing adapter is attached, then the second (N-terminal) sequencing adapter is attached, then the oligopeptide-adapter adduct is released from the resin. In some embodiments the oligopeptide-adapter adduct is released from the resin and then the first and second sequencing adapters are attached. In some embodiments the C-terminal sequencing adapter is attached, the oligopeptide- adapter adduct is released from the resin, and the second (N-terminal) sequencing adapter is attached.

[0406] One embodiment of the disclosed methods is illustrated in Figure 5. This workflow provides an example of the disclosed methods in which reacting the N-terminal amine group with an amine-reactive moiety comprises reacting the oligopeptide with a nucleotide or peptide handle comprising a purification tag, under conditions such that the handle is attached to the N-terminus of the oligopeptide; thereby generating a fusion adduct; contacting the purification tag with a complementary substrate under conditions such that the fusion adduct is immobilised on the substrate; cleaving or eluting a portion of the fusion adduct comprising the oligopeptide from the substrate; and attaching a second sequencing adapter to the portion.

[0407] With reference to Figure 5, a larger protein may first be digested into smaller fragments e.g. using LysC . The resultant fragments would possess lysine residues at their C-termini and free N-termini.

[0408] After the protein is digested, the resulting oligopeptide fragments may be modified by attachment of a peptide handle using a ligase, thereby forming a fusion adduct. The peptide handle may comprise a purification tag for immobilization of the fusion adduct onto a purification substrate such as a resin or bead. The method may involve contacting the adduct with a purification substrate under conditions such that the purification tag binds to the substrate.

[0409] The method may then comprise modification of the C-terminal lysine residue to be orthogonal to an amine group. For example, the C-terminal lysine may be activated using an amine-reactive compound comprising a click chemistry group, such as ethynyl phenyl isothiocyanate, EPITC.

[0410] After activation of the C-terminal lysine, the oligopeptide may be released from the resin, e.g. by cleaving the oligopeptide from the handle, liberating the free N-terminal amine group of the oligopeptide.

[0411] Following this, the method then comprises

[0412] (i) attachment of a first sequencing adapter to the activated C-terminal lysine residue; and

[0413] (ii) the newly liberated N-terminus may have a second sequencing adapter attached thereto. Any of the disclosed amine-targeting chemistries disclosed herein can be used at this step.

[0414] One embodiment of the disclosed methods is illustrated in Figure 7. This workflow provides an example of the disclosed methods in which reacting the N-terminal amine group with an amine-reactive moiety comprises immobilizing the oligopeptide on an amine-reactive substrate by attaching a handle which comprises a sequencing adapter; and attaching a first sequencing adapter to the first reactive amino acid, and cleaving or eluting the oligopeptide from the substrate.

[0415] With reference to Figure 7, a larger protein may first be digested into smaller fragments e.g. using ArgC . The resultant fragments would possess arginine residues at their C-termini and free N-termini.

[0416] The method may then comprise:

[0417] (i) attaching a handle to the N-terminus wherein the handle comprises a sequencing adapter and a purification tag, thereby forming a fusion adduct;

[0418] (ii) immobilizing the fusion adduct onto a substrate (e.g. a resin or bead) under conditions such that the fusion adduct binds to the substrate; and

[0419] (iii) modification of the C-terminal arginine residue and attachment of a first sequencing adapter. For example, the C-terminal arginine may be activated using an guanidine-reactive compound comprising a click chemistry group. Alternatively the first sequencing adapter can be directly reacted with the guanidine group of the C-terminal arginine.

[0420] These steps may be done simultaneously or sequentially in any order.

[0421] The method may then comprise releasing the oligopeptide-adapter adduct from the substrate, e.g. by eluting the fusion adduct or oligopeptide-adapter adduct from the substrate. Alternatively the purification tag may be attached to the sequencing adapter in the handle via a cleavage site and releasing the oligopeptide-adapter adduct from the substrate may comprise cleaving the cleavage site.

[0422] One embodiment of the disclosed methods is illustrated in Figure 11. This workflow provides an example of the disclosed methods in which reacting the N-terminal amine group with an amine-reactive moiety comprises attaching to the oligopeptide a handle which comprises a purification tag and a second reactive functional group for attachment of a second sequencing adapter; attaching the second sequencing adapter to the second reactive functional group; immobilizing the construct thereby formed via the purification tag; attaching a first sequencing adapter to the first reactive amino acid, and cleaving or eluting the oligopeptide from the substrate.

[0423] With reference to Figure 11, a larger protein may first be digested into smaller fragments e.g. using LysC or ArgC. The resultant fragments would possess lysine or arginine residues at their C-termini and free N-termini.

[0424] The method may then comprise:

[0425] (i) attaching a handle to the N-terminus wherein the handle comprises a reactive functional group for attachment to a sequencing adapter, and a purification tag, thereby forming a fusion adduct;

[0426] (ii) attaching a sequencing adapter to the reactive functional group of the handle; and

[0427] (iii) immobilizing the fusion adduct onto a substrate (e.g. a resin or bead) under conditions such that the fusion adduct binds to the substrate

[0428] These steps may be done simultaneously or sequentially in any order.

[0429] The method may then comprise:

[0430] (iv) attaching a first sequencing adapter to the C-terminal lysine or arginine residue (or to a modified derivative of said C-terminal lysine or arginine residue; thus the method may comprise modifying the C-terminal lysine or arginine residue as described herein); and

[0431] (v) releasing the oligopeptide-adapter adduct from the substrate, e.g. by eluting the fusion adduct or oligopeptide-adapter adduct from the substrate. Alternatively the purification tag may be attached to the sequencing adapter in the handle via a cleavage site and releasing the oligopeptide-adapter adduct from the substrate may comprise cleaving the cleavage site. Again these steps may be done simultaneously or sequentially in any order.

[0432] In Figure 11, the second reactive functional group for attachment of the second sequencing adapter at the N-terminus is shown as an N3 (azide) group. Those skilled in the art will appreciate that this is solely for non-limiting illustrative purposes and other reactive groups as described herein could be used also; for example, the N3 group could be replaced by any other suitable click chemistry group such as MeTet, BCN, TCO, etc.

[0433] Similarly, in Figure 11, the purification tag is shown as X (e.g. desthiobiotin). Those skilled in the art will appreciate that this is solely for non-limiting illustrative purposes and other reactive groups as described herein could be used also; for example, the X group could be replaced by any other suitable purification tag such as an affinity purification tag such as a poly-His tag.

[0434] It is to be understood that although particular embodiments, specific configurations as well as materials and / or molecules, have been discussed herein for methods according to the present invention, various changes or modifications in form and detail may be made without departing from the scope and spirit of this invention. The preceding embodiments and subsequent examples are provided for illustration only, and should not be considered limiting the application. The application is limited only by the claims.

[0435] Examples

[0436] EXAMPLE 1

[0437] A section of a model protein, SMT3, (GMSEEKPKEGVKTENDHINLKVAGQDG SVVQFKIKRHTPLSKLMK; SEQ ID NO 19) was digested with LysC (Promega) as per manufacturer’s instructions. An aliquot of this digest was then diluted into HEPES-salt buffer (50 mM, pH 7.5) to give a final concentration of 1 pM. To the digest solution, a 2- PCA derivative bearing an immobilisation tag was added to a final concentration of 5 mM. The mixture was incubated at 60°C for 1 hour, then added to a solid support. Once immobilised, the solid support was resuspended in HEPES-salt buffer (50 mM, pH 8.5) and NHS-azido acetic acid was added to give a final concentration of 5 mM. The mixture was incubated at room temperature for 30 min then the solid support washed with HEPES- salt buffer. The solid support was resuspended in Oligo- l / Oligo-4 (20 pM) and HEPES- salt buffer. The suspension was incubated at 37°C for 1 h then the solid support washed and the DNA-peptide construct eluted. The DNA-peptide construct was immobilised onto SPRI beads. Immobilised DNA-peptide conjugate was modified with azidoacetic acid NHS ester, which was added to a final concentration of 5 mM, and the suspension incubated at 37°C for 30 min. Beads were washed with 75% ethanol. The DNA-peptide conjugate was then further modified with Oligo-2 / Oligo-3 in a HEPES-PEG 8K buffer (28% PEG 8K); the reaction was allowed to proceed for 1 hour at 37°C. Beads were washed with 75% ethanol and the DNA-peptide conjugate was subsequently eluted from the beads in 25 mM HEPES pH 7.5, 50 mM NaCl. The DNA-peptide conjugates were attached to sequencing adapters (Native Barcoding Kit version 14; Oxford Nanopore Technologies) by ligation, following manufacturer’s instructions. Data were collected on a GridlON nanopore sequencing device (Oxford Nanopore Technologies) using custom flowcells. Example current traces produced by the DNA-peptide conjugates are shown in (Figure 8); time in seconds is shown on the x-axis and current in picoamps (pA) is shown on the y-axis.

[0438] EXAMPLE 2

[0439] A model peptide, DLREEYEVVK (SEQ ID NO: 20), representing a possible LysC digestion product was used. The peptide (1.7 mM final concentration), an acyl donor Biotin-Gly-Gly-Lys-Ile-Thr-Thr-azidohomoalanine-carboxamidomethyl ester-Leu-NH2 (Biotin-GGKITThA(N3)-Cam-Leu-NH2; SEQ ID NO: 21; 4 mM final concentration) and potassium phosphate (500 mM, pH 8) were combined. Thymoligase was added to give a final concentration of 0.04 mg / mL. The mixture was incubated at 25°C for 1 h. To the resultant mixture a solution of Oligo-4 / Oligo-5, 20 pM) and HEPES-salt buffer (50 mM, pH 7.5) was added. The mixture was incubated at 37°C for 1 h then subsequently immobilised onto Streptavidin Cl beads (ThermoFisher) as per manufacturer’s instructions. The beads were washed with buffer then resuspended into HEPES-NaCl buffer (50 mM, pH 7.5) and NHS-tetrazine added to give a final concentration of 5 mM. The mixture was incubated at 37°C for 1 h then washed again with buffer. After washing, the beads were resuspended in a solution of Oligo-3 / Oligo-6, (20 pM) and HEPES-PEG buffer (50 mM, pH 7.5, 28% PEG 8K) and suspension was incubated at 37°C for 1 h. The beads were then washed with HEPES-NaCl buffer (50 mM, pH 7.5) and the DNA-peptide conjugate was eluted by digestion with USER enzyme (New England Biolabs). The DNA- peptide conjugate was purified by SPRI, immobilised DNA-peptide conjugates were washed with 75% ethanol and eluted from the beads in 25 mM HEPES pH 7.5, 50 mM NaCl. Sequencing adapters (Native Barcoding Kit version 14; Oxford Nanopore Technologies) were attached to the DNA-peptide conjugates by ligation as per manufacturer’s instructions. Data were collected on a GridlON nanopore sequencing device (Oxford Nanopore Technologies) using custom flowcells. Example current traces produced by the DNA-peptide conjugates are shown in (Figure 9).

[0440] EXAMPLE 3

[0441] A model peptide, N3-EEYESEEEGEAR (SEQ ID NO: 22), representing an expected product from an ArgC digestion and post N-terminal functionalisation was used. The peptide was diluted into HEPES-salt buffer (25 mM, pH 7.5) to a final concentration of 100 pM then Oligo-4 / Oligo-5 (100 pmol) added. The mixture was incubated at 37°C for 1 h, then immobilised onto SPRI beads. The beads were resuspended in borate-PEG buffer (50 mM, pH 9, 28% PEG 8K) then Oligo-3 / Oligo-7 was added. The suspension was incubated at 37°C for 16 h. The DNA-peptide construct was eluted from the beads in 25 mM HEPES pH 7.5, 50 mM NaCl. Sequencing adapters (Native Barcoding Kit version 14; Oxford Nanopore Technologies) were attached to DNA-peptide conjugates by ligation as per manufacturer’s instructions. Data were collected on a GridlON nanopore sequencing device (Oxford Nanopore Technologies) using custom flowcells. Example current traces produced by the DNA-peptide conjugates are shown in (Figure 10).

[0442] EXAMPLE 4

[0443] A model peptide LSEPAELTDAVK (SEQ ID NO: 23; 20 nmol) was labelled at the N-terminus through addition of an oligonucleotide corresponding to SEQ ID NO: 24 (500 pmol), 2-pyridinecarboxylaldehyde (250 nmol) and copper acetate (500 nmol) in citrate- PEG buffer (25 mM, 20% 8K PEG). The mixture was incubated at 50°C for 2 h then purified using Monarch® Spin columns and eluted into 25 mM HEPES pH 8.0. A reactive handle was introduced onto the sidechain of the C-terminal lysine by modification with azido acetic acid NHS-ester (500 nmol). The mixture was incubated at 37°C for 30 min. The reaction mixture was purified using Monarch® Spin columns and eluted into 25 mM HEPES pH 7.5. Oligonucleotides corresponding to SEQ ID NO: 25 and SEQ ID NO: 26 were annealed by heating in a HEPES buffer. The sidechain of the C-terminal lysine of the DNA-modified peptide was then modified with click chemistry to attach the annealed duplex in HEPES-PEG buffer (50% PEG). The mixture was incubated at 50°C for 2 h then immobilised onto Streptavidin Cl beads (ThermoFisher) as per manufacturer’s instructions. The beads were washed with buffer then resuspended into HEPES-NaCl buffer (pH 7.5). The DNA-peptide conjugate was eluted by addition of USER enzyme (New England Biolabs). The DNA-peptide conjugate was purified by SPRI, immobilised DNA-peptide conjugates were washed with 75% ethanol and eluted from the beads in 25 mM HEPES pH 7.5, 50 mM NaCl. Native Adapters (SQK-NBD114; Oxford Nanopore Technologies) were ligated to the DNA-peptide conjugates as described in provided protocols. Data were collected on a GridlON nanopore sequencing device (Oxford Nanopore Technologies) using custom flow cells. Example current traces produced by the DNA-peptide conjugates are shown in Figure 12.

[0444] Sequences The following table provides amino acid sequences of example ligase enzymes (each comprising an optional C-terminal hexa-histidine tag) described herein.

Claims

CLAIMS1. A method of preparing an oligopeptide-adapter adduct for characterisation of said oligopeptide using a nanopore, the method comprising i) contacting a polypeptide comprising said oligopeptide with a proteinase selective for a cleavage site, under conditions such that the proteinase selectively cleaves the polypeptide thereby forming one or more oligopeptides, wherein each oligopeptide comprises a first end terminating in a first amino acid motif comprising a first reactive amino acid, and a second end terminating in a second amino acid; and ii) selectively attaching to each first reactive amino acid a first sequencing adapter, thereby forming an oligopeptide-adapter adduct.

2. A method according to claim 1, wherein the proteinase selectively cleaves the polypeptide at each occurrence of said cleavage site in the polypeptide.

3. A method according to claim 1 or 2, wherein for each oligopeptide formed in step (i), the first end is the C-terminus of the oligopeptide and the second end is the N-terminus of the oligopeptide.

4. A method according to claim 3, wherein the second end comprises a N-terminal amine group.

5. A method according to claim 4, wherein step (ii) comprises reacting the N-terminal amine group with an amine-reactive moiety.

6. A method according to any one of the preceding claims, wherein step (ii) comprises reacting the oligopeptide with the first sequencing adapter under conditions such that the first sequencing adapter reacts with the first reactive amino acid.

7. A method according to any one of claims 1 to 6, wherein step (ii) comprises functionalising the first reactive amino acid by reacting the oligopeptide with a compound comprising a oligo-reactive functional group and an adapter-reactive functional groupunder conditions such that the oligo-reactive functional group reacts with the first reactive amino acid.

8. A method according to claim 7, comprising reacting the first sequencing adapter with the adapter-reactive functional group.

9. A method according to any one of the preceding claims, wherein the cleavage site comprises lysine or arginine.

10. A method according to any one of the preceding claims, wherein the cleavage site, the first amino acid motif and the first reactive amino acid each comprise or consist of lysine.

11. A method according to any one of the preceding claims, wherein the proteinase is LysC or a or functional analog, fragment or variant thereof.

12. A method according to any one of claims 1 to 9, wherein the cleavage site, the first amino acid motif and the first reactive amino acid each comprise or consist of arginine.

13. A method according to claims 1 to 9 or 12, wherein the proteinase is ArgC or a or functional analog, fragment or variant14. A method according to any one of claims 5 to 13, wherein reacting the N-terminal amine group with an amine-reactive moiety comprises immobilizing the oligopeptide on an substrate.

15. A method according to claim 14, wherein the substrate is an amine-reactive substrate comprising a reactive carbonyl group; preferably wherein the reactive carbonyl group is an aldehyde group.

16. A method according to claim 15, wherein the amine-reactive substrate is functionalised with 2-PCA or an amine-binding derivative thereof.

17. A method according to claim any one of claims 14 to 16, comprising cleaving or eluting the oligopeptide from the substrate.

18. A method according to claim 17, wherein said cleaving or eluting generates a second reactive functional group comprised in said oligopeptide capable of reacting with a complementary functional group on a second sequencing adapter.

19. A method according to any one of claims 5 to 13, wherein reacting the N-terminal amine group with an amine-reactive moiety comprises reacting the oligopeptide with a nucleotide or peptide handle comprising a purification tag, under conditions such that the handle is attached to the N-terminus of the oligopeptide; thereby generating a fusion adduct.

20. A method according to claim 19, comprising contacting the oligopeptide with a ligase and a peptide handle comprising a purification tag, under conditions such that the ligase ligates the peptide handle to the N-terminus of the oligopeptide.

21. A method according to claim 19 or 20, comprising contacting the purification tag with a complementary substrate under conditions such that the fusion adduct is immobilised on the substrate.

22. A method according to claim 21, comprising cleaving or eluting a portion of the fusion adduct comprising the oligopeptide from the substrate.

23. A method according to claim 22, comprising attaching a second sequencing adapter to the portion.

24. A method according to claim 22 or 23, comprising attaching a second sequencing adapter to the portion after the portion has been cleaved from the immobilised fusion adduct.

25. A method according to any one of claims 19 to 24, wherein the handle comprises a second reactive functional group capable of reacting with a complementary functional group on the second sequencing adapter.

26. A method according to any one of claims 22 to 25, comprising cleaving the portion from the immobilised fusion adduct thereby generating a second reactive functional group capable of reacting with a complementary functional group on the second sequencing adapter.

27. A method according to any one of claims 19 to 22, wherein a second sequencing adapter is comprised in the handle.

28. A method according to any one of claims 14 to 27, wherein the oligopeptide is attached to the first sequencing adapter.

29. A method according to any one of claims 5 to 13, wherein reacting the N-terminal amine group with an amine-reactive moiety comprises attaching a first sequencing adapter to the first reactive amino acid and to the N-terminal amine group.

30. A method according to claim 29, comprising cleaving the amino acid comprising the reacted N-terminal amine group from the oligopeptide-adapter adduct.

31. A method according to claim 30, wherein cleaving the reacted N-terminal amine group from the oligopeptide-adapter adduct generates a second reactive functional group capable of reacting with a complementary functional group on a second sequencing adapter.

32. A method according to any one of claims 18, 24 to 26, or 31, comprising attaching the second sequencing adapter to the second reactive functional group.

33. A method according to claim 32, wherein the second reactive functional group is an N-terminal amine group.

34. A method according to claim 32, wherein the second reactive functional group is orthogonal to the optionally-functionalised first reactive amino acid.

35. A method according to claim any one of the preceding claims, wherein attaching the first sequencing adapter and a second sequencing adapter to the oligopeptide generates an adduct of form:S - " Oc- F wherein F is the first sequencing adapter,W’O’Cis the oligopeptide wherein N- represents the N-terminal end of the oligopeptide and -C represents the C-terminal end of the oligopeptide, and S is the second sequencing adapter.

36. A method according to any one of the preceding claims, wherein the first sequencing adapter is a polynucleotide sequencing adapter.

37. A method according to any one of claims 18, 22 to 27, or 31 to 35, wherein the second sequencing adapter comprises a polynucleotide and / or a polypeptide; preferably wherein the second sequencing adapter is a polynucleotide sequencing adapter.

38. A method according to claim 37, wherein the first sequencing adapter is different to the second sequencing adapter.

39. A method according to any one of the preceding claims, comprising loading a motor protein onto the first sequencing adapter and / or the second sequencing adapter if present.

40. A method according to any one of the preceding claims, wherein the oligopeptide comprises from about 2 to about 50 amino acids; preferably from about 5 to about 30 amino acids.

41. A method of characterising a polypeptide, comprising carrying out a method according to any one of the preceding claims, and contacting the oligopeptide with a nanopore under conditions such that the oligopeptide moves with respect to the nanopore; and- taking one or more measurements characteristic of the oligopeptide as the oligopeptide moves with respect to then nanopore; thereby characterising the oligopeptide.

42. A method according to claim 41, wherein the oligopeptide is comprised in an adduct of form:S - " Oc- F wherein F is the first sequencing adapter,W’O’Cis the oligopeptide wherein N- represents the N-terminal end of the oligopeptide and -C represents the C-terminal end of the oligopeptide, and S is a second sequencing adapter.

43. A kit for modifying a polypeptide, comprising a first sequencing adapter capable of selectively reacting with an optionally- functionalised first reactive amino acid at the C-terminus of an oligopeptide generated by contacting a polypeptide with a proteinase; and a second sequencing adapter capable of reacting with a second reactive functional group at the N-terminus of said oligopeptide; and optionally comprising one or more of: a proteinase capable of selectively cleaving a polypeptide at a cleavage site thereby forming one or more oligopeptides each comprising a first end terminating in a first amino acid motif comprising a first reactive amino acid, and a second end terminating in a second amino acid; a compound comprising a oligo-reactive functional group and an adapter-reactive functional group, wherein the compound is capable of reacting with the first reactive amino acid thereby functionalising the first reactive amino acid; and a peptide handle capable of being ligated onto the N-terminus of said oligopeptide; a ligase capable of ligating a peptide handle onto the N-terminus of said oligopeptide; a motor protein capable of controlling the movement of the oligopeptide with respect to a nanopore.

Citation Information

Patent Citations

  • phi 29 DNA polymerase

    US5576204A

  • A miniature support for thin films containing single channels or nanopores and methods for using same

    WO2000028312A1

  • Control of solid state dimensional features

    WO2003003446A2

  • Deliver of molecules to a li id bila

    WO2006100484A2

  • Lipid bilayer sensor system

    WO2008102120A1