method
By modifying polypeptides to increase charge and using a protein translocase to control movement through a nanopore, the method addresses inefficiencies in existing characterization techniques, achieving precise single-molecule analysis of polypeptides.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- OXFORD NANOPORE TECH LTD
- Filing Date
- 2026-01-19
- Publication Date
- 2026-07-23
AI Technical Summary
Existing methods for characterizing polypeptides, such as mass spectrometry and Edman degradation, are inefficient and unsuitable for single molecule analysis, lacking the precision needed to distinguish neighboring residues and requiring costly reagents.
A method involving the modification of polypeptides to increase their net charge and conjugation with a leader component, using a protein translocase to control the movement of the polypeptide through a nanopore, allowing for controlled translocation and characterization.
Enables accurate, single-molecule characterization of polypeptides with improved precision and efficiency, overcoming limitations of existing techniques by facilitating regular movement and data accuracy.
Smart Images

Figure IMGF000009_0001 
Figure IMGF000009_0002 
Figure IMGF000015_0001_TABLE
Abstract
Description
[0001] METHOD
[0002] Related applications
[0003] The present application claims priority to United Kingdom Patent Application No. 2500667.7 filed 17 January 2025, United Kingdom Patent Application No.
[0004] 2513808.2 filed 22 August 2025, and United Kingdom Patent Application No.
[0005] 2518554.7 filed 6 November 2025, the entire contents of each of which are hereby incorporated by reference in their entirety.
[0006] Field
[0007] The present disclosure relates to methods of moving a polypeptide with respect to a nanopore using a protein translocase. The present disclosure also relates to related further methods and associated constructs, kits and systems. The methods, constructs, kits and systems are of particular reference to detecting and characterising target polypeptides in a sample using a nanopore.
[0008] Background
[0009] The characterisation of biological molecules is of increasing importance in biomedical and biotechnological applications. For example, sequencing of nucleic acids allows the study of genomes and the proteins they encode and, for example, allows correlation between nucleic acid mutations and observable phenomena such as disease indications. Nucleic acid sequencing can be used in evolutionary biology to study the relationship between organisms. Metagenomics involves identifying organisms present in samples, for example microbes in a microbiome, with nucleic acid sequencing allowing the identification of such organisms.
[0010] Whilst techniques to characterise (e.g. sequence) polynucleotides have been extensively developed, techniques to characterise polypeptides are less advanced, despite being of very significant biotechnological importance. For example, knowledge of a protein sequence can allow structure-activity relationships to be established and has implications in rational drug development strategies for developing ligands for specific receptors. Identification of post-translational modifications is also key to understanding the functional properties of many proteins. For example, typically 30-50% of protein species are phosphorylated in eukaryotes. Some proteins may have multiplephosphorylation sites, serving to activate or inactivate a protein, promote its degradation, or modulate interactions with protein partners.
[0011] Known methods of characterising polypeptides include mass spectrometry and Edman degradation.
[0012] Protein mass spectrometry involves characterising whole proteins or fragments thereof in an ionised form. Known methods of protein mass spectrometry include electrospray ionisation (ESI) and matrix-assisted laser desorption / ionisation (MALDI). Mass spectrometry has some benefits, but results obtained can be affected by the presence of contaminants and it can be difficult to process fragile molecules without their fragmentation. Moreover, mass spectrometry is not a single molecule technique and provides only bulk information about the sample interrogated. Mass spectrometry is unsuitable for characterising differences within a population of polypeptide samples and is unwieldy when seeking to distinguish neighbouring residues.
[0013] Edman degradation is an alternative to mass spectrometry which allows the residue-by-residue sequencing of polypeptides. Edman degradation sequences polypeptides by sequentially cleaving the N-terminal amino acid and then characterising the individually cleaved residues using chromatography or electrophoresis. However, Edman sequencing is slow, involves the use of costly reagents, and like mass spectrometry is not a single molecule technique.
[0014] As such, there remains a pressing need for new techniques to characterise polypeptides, especially at the single molecule level. Single molecule techniques for characterising biomolecules such as polynucleotides have proven to be particularly attractive due to their high fidelity and avoidance of amplification bias.
[0015] Summary
[0016] One attractive method of single molecule characterization of biomolecules such as polypeptides is nanopore sensing. Nanopore sensing is an approach to analyte detection and characterization that relies on the observation of individual binding or interaction events between the analyte molecules and an ion conducting channel.
[0017] Nanopore sensors can be created by placing a single pore of nanometre dimensions in an electrically insulating membrane and measuring voltage-driven ion currents through the pore in the presence of analyte molecules. The presence of an analyte inside or nearthe nanopore will alter the ionic flow through the pore, resulting in altered ionic or electric currents being measured over the channel. The identity of an analyte is revealed through its distinctive current signature, notably the duration and extent of current blocks and the variance of current levels during its interaction time with the pore.
[0018] Nanopore sensing has the potential to allow rapid and cheap polypeptide characterisation.
[0019] Nanopore sensing and characterisation of polypeptides has been proposed in the art. For example, WO 2013 / 123379 discloses the use of an NTP-driven protein processing unfoldase enzyme to process a protein to be translocated through a nanopore. WO 2021 / 111125 discloses methods in which a target polypeptide may be characterised as it moves through a nanopore using a polynucleotide-handling protein. WO 2021 / 133168 discloses protein and polypeptide fingerprinting and sequencing by nanopore translocation. WO 2024 / 094986 discloses methods of characterising a target polypeptide as it moves with relation to a nanopore. Each of these documents is incorporated by reference in their entireties.
[0020] These methods have provided useful techniques for characterising polypeptides using nanopores. However, there remains a need for alternative and / or improved methods of characterising polypeptides.
[0021] The present disclosure relates to methods of moving a polypeptide (e.g. a target polypeptide) with respect to a nanopore using a protein translocase.
[0022] The provided methods comprise modifying the polypeptide to increase the net charge of the polypeptide and contacting the polypeptide with a nanopore having a first opening and a second opening. The polypeptide is contacted with the first opening of the nanopore under conditions such that the polypeptide threads through the nanopore in the direction from the first opening to the second opening and moves with respect to the nanopore in the direction from the first opening to the second opening. The methods further comprise controlling the movement of the polypeptide in the direction from the first opening to the second opening of the nanopore using the protein translocase. The movement of the construct can thus be conceptually considered as movement “into” the pore (from the perspective of the protein translocase), and thus the protein translocasecan thus be conceptually considered as controlling the movement of the construct “into” the pore (from the perspective of the protein translocase).
[0023] In an aspect, the provided methods comprise modifying the polypeptide to increase the net charge of the polypeptide and conjugating the polypeptide with a leader comprising a non-polypeptide component such as a polynucleotide (such as a leader described in more detail herein) to form a construct. As described in more detail here, the modifying of the peptide can be conducted prior to, simultaneously with, or subsequent to the conjugation of the polypeptide to the leader. The methods then comprise contacting the construct with a nanopore having a first opening and a second opening. In more detail, the methods comprise contacting the construct with the first opening of the nanopore under conditions such that the leader threads through the nanopore in the direction from the first opening to the second opening and the construct moves with respect to the nanopore in the direction from the first opening to the second opening. The methods further comprise controlling the movement of the construct in the direction from the first opening to the second opening of the nanopore using the protein translocase.
[0024] Accordingly, provided herein is a method of moving a target polypeptide with respect to a nanopore using a protein translocase;
[0025] the nanopore having a first opening and a second opening;
[0026] the method comprising:
[0027] (i) modifying the target polypeptide to increase the net charge of the polypeptide and (ii) conjugating the target polypeptide with a leader comprising a non-polypeptide component, thereby forming a construct; wherein step (i) can be conducted prior to, simultaneously with or subsequent to step (ii);
[0028] contacting the construct with the first opening of the nanopore under conditions such that the leader threads through the nanopore in the direction from the first opening to the second opening and the construct moves with respect to the nanopore in the direction from the first opening to the second opening; and
[0029] controlling the movement of the construct in the direction from the first opening to the second opening of the nanopore using the protein translocase.In some embodiments, modifying the target polypeptide comprises modifying the side chain(s) of one or more amino acids in the target polypeptide. In some embodiments modifying the target polypeptide comprises covalently attaching said one or more charge-modifying moieties to said side chain(s). In some embodiments modifying the target polypeptide comprises contacting the target polypeptide with one or more amino-acid modifying enzymes and / or with one or more chemical reagents. In some embodiments step (i) comprises modifying the target polypeptide to increase the net negative charge of the polypeptide.
[0030] In some embodiments the protein translocase controls the movement of a portion of the target polypeptide with respect to the nanopore. In some embodiments the portion of the target polypeptide has a length of at least 5, at least 10, at least 20, at least 30, at least 40, at least 50, or at least 100 amino acids.
[0031] In some embodiments the target polypeptide has a length of at least 5, at least 10, at least 20, at least 30, at least 40, at least 50, at least 100, at least 200, at least 300, at least 400 or at least 500 amino acids.
[0032] In some embodiments the target polypeptide has a first end and a second end; the first end of the target polypeptide is attached to the leader; and the protein translocase is oriented on the construct in an orientation to process the target polypeptide in a direction from the first end towards the second end.
[0033] In some embodiments the leader is comprised in an adapter comprising a recognition sequence for the protein translocase. In some embodiments the method comprises loading the protein translocase onto the recognition sequence prior to contacting the construct with the nanopore.
[0034] In some embodiments the method comprises, prior to conjugating the target polypeptide with the adapter, the step of loading the protein translocase onto the adapter. In some embodiments the method comprises the step of loading the protein translocase onto the adapter after conjugating the target polypeptide with the adapter.
[0035] In some embodiments the target polypeptide has a first end and a second end; the first end of the target polypeptide is conjugated to the leader; and the second end of the target polypeptide comprises or is attached to a blocking moiety to prevent the protein translocase from disengaging from the construct and / or to prevent the second end of the target polypeptide from translocating through the nanopore. In some embodiments theblocking moiety limits the movement of the target polypeptide with respect to the polypeptide binding site of the protein translocase and thereby limits the movement of the target polypeptide with respect to the nanopore in the direction from the first opening to the second opening. In some embodiments the method comprises attaching the blocking moiety to the second end of the target polypeptide prior to conjugating the leader to the first end of the target polypeptide. In some embodiments the method comprises attaching the blocking moiety to the second end of the target polypeptide prior after conjugating the adapter to the first end of the target polypeptide.
[0036] In some embodiments the construct comprises the target polypeptide attached to the leader via a recognition sequence for the protein translocase. In some embodiments the leader comprises a polynucleotide. In some embodiments the leader is functionalised for localisation at the nanopore and / or at a membrane comprising the nanopore.
[0037] In some embodiments the construct moves in the direction from the first opening of the nanopore to the second opening of the nanopore under an applied force. In some embodiments the construct moves in the direction from the first opening of the nanopore to the second opening of the nanopore with the applied force and wherein the protein translocase controls the movement of the construct in the direction from the first opening of the nanopore to the second opening of the nanopore with the applied force. In some embodiments the applied force is an electrical or chemical potential applied across the nanopore. In some embodiments the applied force is a voltage potential.
[0038] In some embodiments, the provided methods comprise applying an electroosmotic force across the nanopore. In some embodiments the nanopore is configured to generate an electroosmotic force across the nanopore. In some embodiments the construct moves with the electroosmotic force with respect to the nanopore. In some embodiments the electroosmotic force arises from an electroosmotic flow through the nanopore in the same direction as an electrophoretic force applied across the nanopore.
[0039] In some embodiments the nanopore is a transmembrane nanopore spanning a membrane having a cis side and a trans side, and: (i) the first opening of the nanopore is at the cis side of the membrane and the second opening of the nanopore is at the trans side; and the protein translocase controls the movement of the construct through the nanopore from the cis side to the trans side of the membrane; or (ii) the first opening ofthe nanopore is at the trans side of the membrane and the second opening of the nanopore is at the cis side; and the protein translocase controls the movement of the construct through the nanopore from the trans side to the cis side of the membrane. In some embodiments the nanopore is comprised in a membrane separating a first volume from a second volume, wherein the first volume contacts the first opening of the nanopore and the second volume contacts the second opening of the nanopore; and wherein the protein translocase is retained in the first volume during said method.
[0040] In some embodiments the protein translocase is an NTP-driven unfoldase.
[0041] Also provided is a method of characterising a target polypeptide,
[0042] the target polypeptide being comprised in a construct comprising the target polypeptide conjugated to a leader comprising a non-polypeptide component, wherein the construct further comprises said protein translocase;
[0043] the method comprising
[0044] moving the target polypeptide with respect to a nanopore as defined herein; and
[0045] taking one or more measurements characteristic of the target polypeptide as the protein translocase controls the movement of the construct with respect to the nanopore,
[0046] thereby characterising the target polypeptide.
[0047] In some embodiments the one or more measurements are characteristic of one or more characteristics of the target polypeptide selected from (i) the length of the polypeptide, (ii) the identity of the polypeptide, (iii) the sequence of the polypeptide, (iv) the secondary structure of the polypeptide and (v) whether or not the polypeptide is modified.
[0048] Also provided is a method of controlling the movement of a target polypeptide; the target polypeptide being comprised in a construct comprising the target polypeptide conjugated to a leader comprising a non-polypeptide component, wherein the construct further comprises said protein translocase;
[0049] the method comprising moving the target polypeptide with respect to a nanopore as defined herein; and controlling the speed of the movement of the construct with respect to the nanopore using the protein translocase;thereby controlling the movement of the target polypeptide with respect to the nanopore.
[0050] Also provided is a construct comprising a polypeptide conjugated to an adapter comprising (a) a leader comprising a non-polypeptide component and optionally (b) a recognition sequence for a protein translocase, wherein the construct further comprises said protein translocase; wherein
[0051] the polypeptide comprises a first end and a second end;
[0052] the first end of the polypeptide is attached to the leader;
[0053] the second end of the polypeptide is attached to a blocking moiety capable of limiting the movement of the polypeptide with respect to the polypeptide binding site of the protein translocase; and
[0054] the protein translocase is oriented on the construct in an orientation for processing the polypeptide in a direction from the first end to the second end.
[0055] Also provided is a kit, comprising
[0056] a first adapter having a first end comprising a leader comprising a non- polypeptide component; an optional recognition sequence for a protein translocase; and a second end comprising an attachment point for attaching to a first end of a polypeptide analyte; and
[0057] a second adapter having a first end comprising an attachment point for attaching to a second end of the polypeptide analyte; and a second end comprising a blocking moiety;
[0058] wherein the first adapter comprises a protein translocase in an orientation for processing said adapter in a direction from the first end to the second end;
[0059] and wherein the blocking moiety of the second adapter is suitable for preventing the protein translocase from disengaging from the second end of the polypeptide analyte when the first and second adapters are attached to the polypeptide analyte.
[0060] Also provided is a system for characterising a target polypeptide comprising: first and second adapters as defined herein;
[0061] a nanopore for characterising the target polypeptide as the target polypeptide moves with respect to the nanopore; and
[0062] a protein translocase for controlling the movement of the target polypeptide with respect to the nanopore.In some embodiments of the construct, the kit and the system provided here, the polypeptide, the leader, the protein translocase, the blocking moiety, and the nanopore are each optionally and independently as defined herein.
[0063]
[0064] Figure 1 shows a non-limiting exemplary embodiment of an analyte suitable for use in the methods disclosed herein. In Figure 1, (1) is an unfoldase recognition sequence, (2) is an oligonucleotide leader, (3) is a linking moiety introduced by ybbR tagging, and (4) is the protein analyte.
[0065] Figure 2 shows a non-limiting exemplary embodiment of the disclosed methods. In Figure 2, a protein unfoldase (1) interacts with an analyte with an oligonucleotide leader and recognition sequence (2); the leader promotes capture by a nanopore and initiates threading into the nanopore (3); and the unfoldase controls the translocation of the analyte through the nanopore (4).
[0066] Figure 3 shows a non-limiting exemplary embodiment of a workflow suitable for use in accordance with the methods disclosed herein.
[0067] Figure 4 shows by way of non-limiting illustration some exemplary reactions of amino acids that can be made in accordance with the provided methods
[0068] Figure 5 shows exemplary electrophysiology data demonstrating successful controlled translocation of a polypeptide analyte through a nanopore, as described in the Example.
[0069] Figure 6 shows exemplary electrophysiology data demonstrating that modification of the polypeptide analyte to increase its net negative charge led to successful controlled translocation through the nanopore, as described in the Example.
[0070] Figure 7 shows plots of charge distribution for unmodified (A) and modified (B) polypeptide analytes as described in the Example.
[0071] Detailed
[0072]
[0073] The present invention will be described with respect to particular embodiments and with reference to certain drawings but the invention is not limited thereto but only by the claims. Any reference signs in the claims shall not be construed as limiting the scope. Of course, it is to be understood that not necessarily all aspects or advantagesmay be achieved in accordance with any particular embodiment of the invention. Thus, for example those skilled in the art will recognize that the invention may be embodied or carried out in a manner that achieves or optimizes one advantage or group of advantages as taught herein without necessarily achieving other aspects or advantages as may be taught or suggested herein.
[0074] The invention, both as to organization and method of operation, together with features and advantages thereof, may best be understood by reference to the following detailed description when read in conjunction with the accompanying drawings. The aspects and advantages of the invention will be apparent from and elucidated with reference to the embodiment(s) described hereinafter. Reference throughout this specification to "one embodiment" or "an embodiment" means that a particular feature, structure or characteristic described in connection with the embodiment is included in at least one embodiment of the present invention. Thus, appearances of the phrases "in one embodiment" or "in an embodiment" in various places throughout this specification are not necessarily all referring to the same embodiment, but may. Similarly, it should be appreciated that in the description of exemplary embodiments of the invention, various features of the invention are sometimes grouped together in a single embodiment, figure, or description thereof for the purpose of streamlining the disclosure and aiding in the understanding of one or more of the various inventive aspects. This method of disclosure, however, is not to be interpreted as reflecting an intention that the claimed invention requires more features than are expressly recited in each claim. Rather, as the following claims reflect, inventive aspects lie in less than all features of a single foregoing disclosed embodiment.
[0075] It should be appreciated that “embodiments” of the disclosure can be specifically combined together unless the context indicates otherwise. The specific combinations of all disclosed embodiments (unless implied otherwise by the context) are further disclosed embodiments of the claimed invention.
[0076] In addition as used in this specification and the appended claims, the singular forms “a”, “an”, and “the” include plural referents unless the content clearly dictates otherwise. Thus, for example, reference to “a polynucleotide” includes two or more polynucleotides, reference to “a motor protein” includes two or more such proteins, reference to “a helicase” includes two or more helicases, reference to “a monomer”refers to two or more monomers, reference to “a pore” includes two or more pores and the like.
[0077] All publications, patents and patent applications cited herein, whether supra or infra, are hereby incorporated by reference in their entirety.
[0078] Definitions
[0079] Where an indefinite or definite article is used when referring to a singular noun e.g. "a" or "an", "the", this includes a plural of that noun unless something else is specifically stated. Where the term "comprising" is used in the present description and claims, it does not exclude other elements or steps. Furthermore, the terms first, second, third and the like in the description and in the claims, are used for distinguishing between similar elements and not necessarily for describing a sequential or chronological order. It is to be understood that the terms so used are interchangeable under appropriate circumstances and that the embodiments of the invention described herein are capable of operation in other sequences than described or illustrated herein. The following terms or definitions are provided solely to aid in the understanding of the invention. Unless specifically defined herein, all terms used herein have the same meaning as they would to one skilled in the art of the present invention. Practitioners are particularly directed to Sambrook et al., Molecular Cloning: A Laboratory Manual, 4thed., Cold Spring Harbor Press, Plainsview, New York (2012); and Ausubel et al., Current Protocols in Molecular Biology (Supplement 114), John Wiley & Sons, New York (2016), for definitions and terms of the art. The definitions provided herein should not be construed to have a scope less than understood by a person of ordinary skill in the art.
[0080] "About" as used herein when referring to a measurable value such as an amount, a temporal duration, and the like, is meant to encompass variations of ± 20 % or ± 10 %, more preferably ± 5 %, even more preferably ± 1 %, and still more preferably ± 0.1 % from the specified value, as such variations are appropriate to perform the disclosed methods.
[0081] “Nucleotide sequence”, “DNA sequence” or “nucleic acid molecule(s)” as used herein refers to a polymeric form of nucleotides of any length, either ribonucleotides or deoxyribonucleotides. This term refers only to the primary structure of the molecule.Thus, this term includes double- and single-stranded DNA, and RNA. The term “nucleic acid” as used herein, is a single or double stranded covalently-linked sequence of nucleotides in which the 3' and 5' ends on each nucleotide are joined by phosphodiester bonds. The polynucleotide may be made up of deoxyribonucleotide bases or ribonucleotide bases. Nucleic acids may be manufactured synthetically in vitro or isolated from natural sources. Nucleic acids may further include modified DNA or RNA, for example DNA or RNA that has been methylated, or RNA that has been subject to post-transcriptional modification, for example 5 ’-capping with 7-m ethylguanosine, 3 ’-processing such as cleavage and polyadenylation, and splicing. Nucleic acids may also include synthetic nucleic acids (XNA), such as hexitol nucleic acid (HNA), cyclohexene nucleic acid (CeNA), threose nucleic acid (TNA), glycerol nucleic acid (GNA), locked nucleic acid (LNA) and peptide nucleic acid (PNA). Sizes of nucleic acids, also referred to herein as “polynucleotides” are typically expressed as the number of base pairs (bp) for double stranded polynucleotides, or in the case of single stranded polynucleotides as the number of nucleotides (nt). One thousand bp or nt equal a kilobase (kb). Polynucleotides of less than around 40 nucleotides in length are typically called “oligonucleotides” and may comprise primers for use in manipulation of DNA such as via polymerase chain reaction (PCR).
[0082] The term “amino acid” in the context of the present disclosure is used in its broadest sense and is meant to include organic compounds containing amine (NH2) and carboxyl (COOH) functional groups, along with a side chain (e.g., a R group) specific to each amino acid. In some embodiments, the amino acids refer to naturally occurring L a-amino acids or residues. The commonly used one and three letter abbreviations for naturally occurring amino acids are used herein: A=Ala; C=Cys; D=Asp; E=Glu;
[0083] F=Phe; G=Gly; H=His; I=Ile; K=Lys; L=Leu; M=Met; N=Asn; P=Pro; Q=Gln; R=Arg; S=Ser; T=Thr; V=Val; W=Trp; and Y=Tyr (Lehninger, A. L., (1975) Biochemistry, 2d ed., pp. 71-92, Worth Publishers, New York). The general term “amino acid” further includes D-amino acids, retro-inverso amino acids as well as chemically modified amino acids such as amino acid analogues, naturally occurring amino acids that are not usually incorporated into proteins such as norleucine, and chemically synthesised compounds having properties known in the art to be characteristic of an amino acid, such as P-amino acids. For example, analogues or mimetics of phenylalanine or proline,which allow the same conformational restriction of the peptide compounds as do natural Phe or Pro, are included within the definition of amino acid. Such analogues and mimetics are referred to herein as "functional equivalents" of the respective amino acid. Other examples of amino acids are listed by Roberts and Vellaccio, The Peptides:
[0084] Analysis, Synthesis, Biology, Gross and Meiehofer, eds., Vol. 5 p. 341, Academic Press, Inc., N.Y. 1983, which is incorporated herein by reference.
[0085] The terms “polypeptide”, and “peptide” are interchangeably used herein to refer to a polymer of amino acid residues and to variants and synthetic analogues of the same. Thus, these terms apply to amino acid polymers in which one or more amino acid residues is a synthetic non-naturally occurring amino acid, such as a chemical analogue of a corresponding naturally occurring amino acid, as well as to naturally-occurring amino acid polymers. Polypeptides can also undergo maturation or post-translational modification processes that may include, but are not limited to: glycosylation, proteolytic cleavage, lipidization, signal peptide cleavage, propeptide cleavage, phosphorylation, and such like. A peptide can be made using recombinant techniques, e.g., through the expression of a recombinant or synthetic polynucleotide. A recombinantly produced peptide is typically substantially free of culture medium, e.g., culture medium represents less than about 20 %, more typically less than about 10 %, and most typically less than about 5 % of the volume of the protein preparation.
[0086] The term “protein” is used to describe a folded polypeptide having a secondary, tertiary, or quaternary structure. The protein may be composed of a single polypeptide, or may comprise multiple polypeptides that are assembled to form a multimer. The multimer may be a homooligomer, or a heterooligmer. The protein may be a naturally occurring, or wild type protein, or a modified, or non-naturally, occurring protein. The protein may, for example, differ from a wild type protein by the addition, substitution or deletion of one or more amino acids.
[0087] A “variant” of a protein encompass peptides, oligopeptides, polypeptides, proteins and enzymes having amino acid substitutions, deletions and / or insertions relative to the unmodified or wild-type protein in question and having similar biological and functional activity as the unmodified protein from which they are derived. The term "amino acid identity" as used herein refers to the extent that sequences are identical on an amino acid-by-amino acid basis over a window of comparison. Thus, a "percentageof sequence identity" is calculated by comparing two optimally aligned sequences over the window of comparison, determining the number of positions at which the identical amino acid residue (e.g., Ala, Pro, Ser, Thr, Gly, Vai, Leu, He, Phe, Tyr, Trp, Lys, Arg, His, Asp, Glu, Asn, Gin, Cys and Met) occurs in both sequences to yield the number of matched positions, dividing the number of matched positions by the total number of positions in the window of comparison (i.e., the window size), and multiplying the result by 100 to yield the percentage of sequence identity.
[0088] For all aspects and embodiments of the present invention, a “variant” has at least 40%, 50%, 60%, 70%, 80%, 90%, 95% or 99% complete sequence identity to the amino acid sequence of the corresponding wild-type protein. Sequence identity can also be to a fragment or portion of the full length polynucleotide or polypeptide. Hence, a sequence may have only 50 % overall sequence identity with a full length reference sequence, but a sequence of a particular region, domain or subunit could share 80 %, 90 %, or as much as 99 % sequence identity with the reference sequence.
[0089] The term “wild-type” refers to a gene or gene product isolated from a naturally occurring source. A wild-type gene is that which is most frequently observed in a population and is thus arbitrarily designed the “normal” or “wild-type” form of the gene. In contrast, the term “modified”, “mutant” or “variant” refers to a gene or gene product that displays modifications in sequence (e.g., substitutions, truncations, or insertions), post-translational modifications and / or functional properties (e.g., altered characteristics) when compared to the wild-type gene or gene product. It is noted that naturally occurring mutants can be isolated; these are identified by the fact that they have altered characteristics when compared to the wild-type gene or gene product. Methods for introducing or substituting naturally-occurring amino acids are well known in the art. For instance, methionine (M) may be substituted with arginine (R) by replacing the codon for methionine (ATG) with a codon for arginine (CGT) at the relevant position in a polynucleotide encoding the mutant monomer. Methods for introducing or substituting non-naturally-occurring amino acids are also well known in the art. For instance, non-naturally-occurring amino acids may be introduced by including synthetic aminoacyl-tRNAs in the IVTT system used to express the mutant monomer. Alternatively, they may be introduced by expressing the mutant monomer in E. coli that are auxotrophic for specific amino acids in the presence of synthetic (i.e.non-naturally-occurring) analogues of those specific amino acids. They may also be produced by native chemical ligation if the mutant monomer is produced using partial peptide synthesis. Conservative substitutions replace amino acids with other amino acids of similar chemical structure, similar chemical properties or similar side-chain volume. The amino acids introduced may have similar polarity, hydrophilicity, hydrophobicity, basicity, acidity, neutrality or charge to the amino acids they replace. Alternatively, the conservative substitution may introduce another amino acid that is aromatic or aliphatic in the place of a pre-existing aromatic or aliphatic amino acid. Conservative amino acid changes are well-known in the art and may be selected in accordance with the properties of the 20 main amino acids as defined in Table 1 below. Where amino acids have similar polarity, this can also be determined by reference to the hydropathy scale for amino acid side chains in Table 2.
[0090] Table 1 - Chemical properties of amino acids
[0091]
[0092] Table 2 - Hydropathy scale
[0093] Side Chain Hydropathy
[0094] He 4.5
[0095] Vai 4.2
[0096] Leu 3.8
[0097] Phe 2.8
[0098] Cys 2.5
[0099] Met 1.9
[0100] Ala 1.8
[0101] Gly -0.4
[0102] Thr -0.7
[0103] Ser -0.8
[0104] Trp -0.9
[0105] Tyr -1.3
[0106] Pro -1.6
[0107] His -3.2
[0108] Glu -3.5
[0109] Gin -3.5
[0110] Asp -3.5
[0111] Asn -3.5
[0112] Lys -3.9
[0113] Arg -4.5
[0114] A mutant or modified protein, monomer or peptide can also be chemically modified in any way and at any site. A mutant or modified monomer or peptide may be chemically modified by attachment of a molecule to one or more cysteines (cysteine linkage), attachment of a molecule to one or more lysines, attachment of a molecule to one or more non-natural amino acids, enzyme modification of an epitope or modification of a terminus. Suitable methods for carrying out such modifications are well-known in the art. The mutant of modified protein, monomer or peptide may be chemically modified by the attachment of any molecule. For instance, the mutant of modified protein, monomer or peptide may be chemically modified by attachment of a dye or a fluorophore.
[0115] As used herein, an alkylene group is a bidentate moiety derived by abstraction of two hydrogen atoms from a linear or branched alkyl group. Typically an alkylene group comprises from 1 to 10 carbon atoms and is referred to as a Ci-io alkylene group. A Ci-io alkylene group is often a Ci-4 alkylene group, or a C1-3 alkylene group. Examples of C1-4 alkylene groups include methylene, ethylene, n-propylene, iso-propylene, n-butylene, sec-butylene, and tert-butylene.As used herein, an alkenylene group is a bidentate moiety derived by abstraction of two hydrogen atoms from a linear or branched linear alkenyl group having one or more, e.g. one or two, typically one double bonds. Typically an alkenylene group comprises from 2 to 10 carbon atoms and is referred to as a C2-10 alkenylene group. A C2-10 alkenylene group is often a C2 to C4 alkenylene group or a C2 to C3 alkenylene group. Examples of C2 to C4 alkenylene groups include ethenylene, propenylene and butenylene.
[0116] As used herein, a alkynylene group is a bidentate moiety derived by abstraction of two hydrogen atoms from a linear or branched linear alkynyl group having one or more, e.g. one or two, typically one triple bonds. Typically an alkynylene group comprises from 2 to 10 carbon atoms and is referred to as a C2-10 alkynylene group. A C2-10 alkenylene group is often a C2 to C4 alkynylene group or a C2 to C3 alkynylene group. Examples of C2 to C4 alkynylene groups include ethynylene, propynylene and butynylene.
[0117] An arylene group is a bidentate moiety derived from an aryl group. As used herein, an aryl group is often a Ce to C10 aryl group which may be a substituted or unsubstituted, monocyclic or fused polycyclic aromatic group containing from 6 to 10 carbon atoms in the ring portion. Examples include monocyclic groups such as phenyl and fused bicyclic groups such as naphthyl and indenyl.
[0118] A heteroarylene group is a bidentate moiety derived from a heteroaryl group. As used herein, an heteroaryl group is often a 5- to 10- membered heteroaryl group which may be a substituted or unsubstituted monocyclic or fused polycyclic aromatic group containing from 5 to 10 atoms in the ring portion, including at least one heteroatom, for example 1, 2 or 3 heteroatoms, typically selected from O, S and N. A heteroaryl group is typically a 5- or 6-membered heteroaryl group or a 9- or 10- membered heteroaryl group. Examples include imidazole, pyridine, pyrimidine and pyrazine.
[0119] A carbocyclylene group is a bidentate moiety derived from a carbocyclyl group. As used herein, a carbocyclyl group is often a 4-10- or 4-6 membered carbocyclic group containing from 4 to 10 carbon atoms. A carbocyclic group may be saturated or partially unsaturated, but is typically saturated. Examples of carbocyclic groups include cyclobutyl, cyclopentyl and cyclohexyl groups.A heterocyclylene group is a bidentate moiety derived from a heterocyclyl group. As used herein, a heterocyclyl group is often a 4-10- or 4-6 membered heterocyclic group containing from 4 to 10 atoms in the ring portion, including at least one heteroatom, for example 1, 2 or 3 heteroatoms, typically selected from O, S and N. A heterocyclic group may be saturated or partially unsaturated, but is typically saturated. Examples of heterocyclic groups include azetidine, morpholine, 1,4-oxazepane, octahydropyrrolo[3,4-c]pyrrole, piperazine, piperidine, and pyrrolidine.
[0120] Disclosed Methods
[0121] Provided herein is a method of moving a target polypeptide with respect to a nanopore using a protein translocase;
[0122] the nanopore having a first opening and a second opening;
[0123] the method comprising:
[0124] modifying the target polypeptide to increase the net charge of the polypeptide, contacting the target polypeptide with the first opening of the nanopore under conditions such that the polypeptide threads through the nanopore in the direction from the first opening to the second opening and moves with respect to the nanopore in the direction from the first opening to the second opening; and
[0125] controlling the movement of the polypeptide in the direction from the first opening to the second opening of the nanopore using the protein translocase. In an aspect, the provided methods comprise moving a target polypeptide with respect to a nanopore using a protein translocase;
[0126] the nanopore having a first opening and a second opening;
[0127] the method comprising:
[0128] (i) modifying the target polypeptide to increase the net charge of the polypeptide and (ii) conjugating the target polypeptide with a leader comprising a non-polypeptide component, thereby forming a construct; wherein step (i) can be conducted prior to, simultaneously with or subsequent to step (ii);
[0129] contacting the construct with the first opening of the nanopore under conditions such that the leader threads through the nanopore in the directionfrom the first opening to the second opening and the construct moves with respect to the nanopore in the direction from the first opening to the second opening; and
[0130] controlling the movement of the construct in the direction from the first opening to the second opening of the nanopore using the protein translocase.
[0131] The disclosure thus relates to methods of moving polypeptides which are charge modified to increase the net charge, and using a protein translocase to control the movement of the polypeptide (e.g. of a construct containing the polypeptide) with respect to a nanopore.
[0132] Modifying the charge of a polypeptide may be associated with advantages compared to methods for characterising polypeptides known in the art. By way of example, polypeptides may in some embodiments have a low charge density and therefore in some embodiments modifying the polypeptide to increase the net charge of the polypeptide may facilitate translocation of the polypeptide through a nanopore under an applied force such as a voltage force. This can facilitate the regular movement of the polypeptide with respect to the nanopore. Regular controlled movement of a polypeptide with respect to a nanopore may lead to more accurate characterisation data for polypeptides characterised in accordance with the disclosed methods as compared to previously known methods. Any suitable method for modifying the charge of the polypeptide can be used. For example, in some embodiments the polypeptide can be modified by contacting the polypeptide with one or more amino-acid modifying enzymes and / or with one or more chemical reagents. Methods of modifying polypeptides in accordance with the present disclosure are described in more detail herein. Often, the methods comprise modifying the polypeptide to increase the net negative charge of the polypeptide. This can be useful when the method is set up such that the protein translocase controls the movement of the polypeptide through the nanopore with an applied force which is a voltage potential, such as a positive voltage potential applied across the nanopore.
[0133] The use of a protein translocase to control the movement of the construct, and thus the movement of the polypeptide, may also be associated with advantages compared to methods for characterising polypeptides known in the art. By way ofexample, protein translocases are capable of processing the handling of long polypeptides. This means that more extensive characterisation data may be obtained for polypeptides characterised in accordance with the disclosed methods as compared to previously known methods. Protein translocases suitable for use in the disclosed methods are described in more detail herein.
[0134] The methods disclosed herein involve using the protein translocase to control the movement of the polypeptide (e.g. a construct comprising the polypeptide and a leader as described herein) in a first direction with respect to a detector such as a nanopore in the disclosed methods. The direction is in the direction from the first opening to the second opening of the nanopore and the polypeptide (or construct) is contacted with the first opening.
[0135] The construct comprises the polypeptide with a leader comprising a nonpolypeptide component. The non-polypeptide component may for example be a polynucleotide such that the leader may comprise or consist of a polynucleotide. The leader may facilitate the threading of the construct through the nanopore. In the disclosed methods any suitable leader may be used.
[0136] In some embodiments the leader is attached to the polypeptide in the construct using a linker. In some embodiments the leader is attached to a first end of the polypeptide. Suitable leaders are described in more detail herein. In the disclosed methods, the leader can be conjugated to the polypeptide using any suitable means. Some exemplary means are described in more detail herein.
[0137] The construct is contacted with the first opening of the nanopore. The contacting between the construct and the nanopore takes place such that the leader threads through the nanopore in the direction from the first opening to the second opening and the construct moves with respect to the nanopore in the direction from the first opening to the second opening. By moving the construct in the direction from the first opening to the second opening of the nanopore the polypeptide portion of the construct is moved in the direction from the first opening to the second opening of the nanopore. In some embodiments, measurements characteristic of the polypeptide are taken as the construct moves in this direction. In some embodiments, such measurements allow a decision to be taken whether or not the construct is a construct of interest that should be further characterised in accordance with the present disclosure.The protein translocase is then used to control the movement of the construct in the direction from the first opening to the second opening of the nanopore. By moving the construct in the direction from the first opening to the second opening of the nanopore the polypeptide portion of the construct is moved in the direction from the first opening to the second opening of the nanopore. In some embodiments, measurements characteristic of the polypeptide are taken as the construct moves in this direction.
[0138] As discussed in more detail herein, the direction of the movement from the first opening to the second opening of the nanopore is thus “into” the nanopore (from the “viewpoint” of the protein translocase). This is described in more detail herein.
[0139] As described in more detail herein, the protein translocase may be loaded onto the construct prior to or during the disclosed methods. The protein translocase may be loaded onto any portion of the construct, for example at a loading site as described in more detail herein. As described in more detail, in some embodiments the construct comprises a recognition site (e.g. a recognition sequence) for the protein translocase. In some embodiments the protein translocase may be loaded onto the recognition site (e.g. the recognition sequence). In some embodiments the protein translocase may be loaded onto a different portion of the construct and localise on the recognition site prior to the methods being conducted.
[0140] In some embodiments the polypeptide has a first end and a second end; the first end of the polypeptide is attached to the leader; and the protein translocase is oriented on the construct in an orientation to process the polypeptide in a direction from the first end towards the second end. Conceptually, therefore, if the protein translocase is considered essentially immobile (e.g. by being held in contact with the nanopore) then the construct will move with respect to the protein translocase in the same direction as the processing of the construct by the protein translocase. Thus, if the protein translocase is oriented on the construct in an orientation to process the polypeptide in a direction from the first end towards the second end, the construct will move with respect to the protein translocase in the direction from the first end towards the second end.
[0141] In some embodiments, a second end of the polypeptide comprises or is attached to a blocking moiety. In some embodiments the blocking moiety is for preventing the protein translocase from disengaging from the construct. In some embodiments theblocking moiety is for preventing the second end of the polypeptide from translocating through the nanopore. In some embodiments the blocking moiety is for preventing the protein translocase from disengaging from the construct and preventing the second end of the polypeptide from translocating through the nanopore. Accordingly, in some embodiments the disclosed methods comprise attaching a blocking moiety to the construct. In some embodiments the disclosed methods comprise attaching the blocking moiety to the polypeptide. In some embodiments the disclosed methods comprise attaching the blocking moiety to the second end of the polypeptide.
[0142] Any suitable blocking moiety can be used and some suitable blocking moieties are described herein. In some embodiments the blocking moiety limits the movement of the polypeptide with respect to the polypeptide binding site of the protein translocase. In some embodiments the blocking moiety limits the movement of the polypeptide with respect to the nanopore (e.g. with respect to the first opening of the nanopore). In this way, the blocking moiety can limits the movement of the polypeptide with respect to the nanopore in the direction from the first opening to the second opening. This is described in more detail herein.
[0143] In some embodiments the protein translocase is loaded onto the construct and then a blocking moiety is attached to the construct. In some embodiments a protein translocase is loaded onto a construct comprising a blocking moiety.
[0144] Any suitable polypeptide can be characterised using the methods disclosed herein. In some embodiments the target polypeptide is a protein or naturally occurring polypeptide. In some embodiments the polypeptide is a synthetic polypeptide.
[0145] Polypeptides which can be characterised in accordance with the disclosed methods are described in more detail herein.
[0146] As discussed herein, the disclosed methods comprise controlling the movement of the construct (or a portion thereof) using a protein translocase. The protein translocase is capable of controlling the movement of polypeptide with respect to a nanopore. Exemplary protein translocases are described in more detail herein.
[0147] The protein translocase controls the movement of the construct, and thus of the polypeptide, with respect to a nanopore. Nanopores suitable for use in the disclosed methods are described in more detail herein.In some embodiments the methods comprise the use of electroosmotic force to facilitate the movement of the polypeptide (e.g. construct) through the nanopore. In some embodiments the methods comprise applying an electroosmotic force across the nanopore. In some embodiments the nanopore is configured to generate, enhance or modify an electroosmotic force across the nanopore. This is described in more detail herein.
[0148] It will be apparent that the disclosed methods allow the speed of the construct, and thus of the polypeptide, to be controlled with respect to the nanopore. Thus, in some embodiments the disclosure provides a method of controlling the movement of a target polypeptide; wherein the target polypeptide is comprised in a construct comprising the target polypeptide attached to a leader comprising a non-polypeptide component, wherein the construct further comprises a protein translocase; the method comprising moving the target polypeptide with respect to a nanopore in accordance with the provided disclosure; and controlling the speed of the movement of the construct with respect to the pore using the protein translocase; thereby controlling the movement of the target polypeptide with respect to the nanopore.
[0149] In some embodiments the disclosed methods comprise taking one or more measurements characteristic of the polypeptide as the construct moves with respect to the nanopore. The one or more measurements can be any suitable measurements. Typically, the one or more measurements are electrical measurements, e.g. current measurements, and / or are one or more optical measurements. Apparatuses for recording suitable measurements, and the information that such measurements can provide, are described in more detail herein.
[0150] Controlling movement of a target polypeptide
[0151] The method can be understood by reference to Figures 1 and 2, which illustrate non-limiting examples of the disclosed method.
[0152] With reference to Figure 1, a construct comprises a polypeptide and a leader comprising a non-polypeptide component (e.g. a leader comprising or consisting of a polynucleotide). With reference to Figure 2, the construct is contacted with a nanopore having a first opening and a second opening (e.g. having cis and trans openings). The construct may be captured by the nanopore. More specifically, the construct may becaptured by the nanopore such that the leader of the construct threads through the first opening of the nanopore. Because the leader contacts the first opening of the nanopore and threads through the nanopore the leader threads through the nanopore in the direction from the first opening to the second opening of the nanopore. For example, the first opening may be the cis opening of the nanopore, and the leader of the construct may thread through the nanopore in the direction from the cis opening towards the trans opening of the nanopore.
[0153] The construct moves through the nanopore in the direction from the first opening towards the second opening. In some non-limiting embodiments, a blocking moiety is present on the construct at the second end of the polypeptide portion of the construct, wherein the leader is attached to a first end of the polypeptide. However, the use of a blocking moiety is not essential and in some embodiments no blocking moiety is present, e.g. in some embodiments no blocking moiety is present at the second end of the polypeptide.
[0154] The construct may move through the nanopore in the direction from the first opening towards the second opening under an applied force. The applied force may be an electrophoretic force. The applied force may operate on the leader. The applied force may operate on the polypeptide portion of the construct. Modifying the charge of the polypeptide of the construct can lead to the applied force having a greater interaction with the polypeptide portion of the construct. For instance, in some embodiments the leader may be negatively charged, the polypeptide may be modified to increase the net negative charge, and the applied force may be a positive charge applied across the nanopore to draw the leader, and thus the construct, into the nanopore with the applied force. (Of course the opposite motion is also possible and the leader could be a positively charged leader, the polypeptide may be modified to increase the net positive charge, and the applied force may be a negative charge applied across the nanopore to draw the leader, and thus the construct, into the nanopore with the applied force.) A protein translocase is then used to control the movement of the construct in the same direction, i.e. in the direction from the first opening (e.g. the cis opening) of the nanopore towards the second opening (e.g. the trans opening) of the nanopore. In some embodiments the protein translocase is loaded onto the construct before the construct is contacted with the nanopore. In some embodiments the construct iscontacted with the nanopore and the protein translocase is then loaded onto the construct, e.g. after the leader has threaded through the nanopore but before the remainder of the construct has threaded through the nanopore. Typically, the protein translocase is loaded onto a loading site of the construct. Often, the loading site is a peptide sequence. The loading site may be referred to as a recognition site or recognition sequence.
[0155] The protein translocase controls the construct in moving in the direction from the first opening of the nanopore towards the second opening of the nanopore. For example, if the leader threads through the nanopore in the direction from the cis opening of the nanopore to the trans opening of the nanopore, the protein translocase may control the movement of the construct in the direction from the cis opening towards the trans opening of the nanopore. The movement may be movement with the force applied across the nanopore. As the protein translocase processes the construct, the construct is moved with respect to the nanopore and so the polypeptide is moved with respect to the nanopore. As the polypeptide is moved with respect to the nanopore (e.g. by being passed through the openings of the nanopore) it may optionally be characterised by taking one or more measurements characteristic of the polypeptide.
[0156] In some embodiments the protein translocase controls the movement of the construct until the protein translocase reaches the second end of the polypeptide. In some embodiments when the protein translocase reaches the second end of the polypeptide the protein translocase disengages from the construct (e.g. by no longer binding to the construct and being able to freely dissociate into the reaction medium).
[0157] In some embodiments the protein translocase controls the movement of the construct until the protein translocase contacts a blocking moiety if present, e.g. at the second end of the polypeptide portion of the construct. In some embodiments, contacting the protein translocase with the blocking moiety causes the protein translocase to unbind from the construct. It is important to distinguish “unbinding” of the protein translocase from the construct from the separate concept of a protein translocase “disengaging” from the construct. As described in more detail herein, the “unbinding” of the protein translocase from the construct is a transient status in which the protein translocase remains engaged with the construct. Thus, the proteintranslocase does not dissociate from the construct during the unbinding of the protein translocase from the construct.
[0158] In some embodiments the construct may be ejected from the nanopore once the protein translocase has controlled the movement of the construct with respect to the nanopore. This allows a subsequence molecule of the construct to be moved with respect to the nanopore as described herein.
[0159] In some embodiments the construct may be reversed through the nanopore without being fully ejected. For example, in some embodiments only a portion of the polypeptide may be ejected from the nanopore and a portion of the polypeptide may remain translocating the nanopore. In some embodiments the protein translocase can then be used to re-control the portion of the polypeptide that has been ejected from the nanopore back through the nanopore. This allows the same molecule of the construct to be characterised multiple times. This can increase measurement accuracy.
[0160] In the example illustrated in Figures 1 and 2 the protein translocase moves the construct “into” the first opening of the nanopore, from the “viewpoint” of the protein translocase. For example, as shown the nanopore may span a membrane having a cis side and a trans side. The first opening of the nanopore may be at the cis side of the membrane and the second opening of the nanopore may be at the trans side such that the protein translocase controls the movement of the construct through the nanopore in the direction “into” of the nanopore (from the “viewpoint” of the protein translocase). Thus, the protein translocase controls the movement of the construct through the nanopore in the direction from the cis side to the trans side of the membrane. In such embodiments the movement of the construct is thus movement from the cis side to the trans side of the nanopore. The movement of the construct controls the movement of the polypeptide portion of the construct comprised therein with respect to the nanopore.
[0161] Of course the opposite setup can also be used in the methods disclosed herein. For example, as shown the nanopore may span a membrane having a cis side and a trans side. The first opening of the nanopore may be at the trans side of the membrane and the second opening of the nanopore may be at the cis side such that the protein translocase controls the movement of the construct through the nanopore in the direction “into” of the nanopore (from the “viewpoint” of the protein translocase). Thus the protein translocase may control the movement of the construct through the nanoporein the direction from the trans side to the cis side of the membrane. In such embodiments the movement of the construct is thus in the direction from the trans side to the cis side of the nanopore. The movement of the construct controls the movement of the polypeptide portion of the construct comprised therein with respect to the nanopore.
[0162] Typically the protein translocase does not translocate through the nanopore in the disclosed methods. In some embodiments the nanopore is comprised in a membrane separating a first volume from a second volume, and the first volume contacts the first opening of the nanopore and the second volume contacts the second opening of the nanopore. Typically, the protein translocase is retained in the first volume during said method. Thus, in some embodiments the method comprises retaining the protein translocase in the first volume and using the protein translocase to control the movement of the construct, and thus of the polypeptide, through the nanopore from the first volume into the second volume.
[0163] An exemplary non-limiting workflow is now described. A nanopore such as a nanopore disclosed herein can be contacted with a construct under conditions such that the leader of the construct is captured by the nanopore.
[0164] In some embodiments a determination is made as to whether or not the construct is a construct of interest. In some embodiments such determination is made by taking one or more measurements characteristic of the construct as the leader moves in the direction of from the first opening of the nanopore towards the second opening of the nanopore. This movement may for example comprise movement of the leader with an applied force across the nanopore such as a voltage force. The one or more measurements may be for example measurements of the length or composition of the leader. In some embodiments the measurements are measurements sufficient to distinguish the construct from e.g. impurities which may be present in a sample.
[0165] If the construct is not a construct of interest, then in some embodiments the methods may involve ejecting the construct from the nanopore. In some embodiments ejecting the construct from the nanopore comprises applying a force across the nanopore in the opposite direction to a force applied across the nanopore during the capture of the construct by the nanopore. For example, if the leader is a negatively charged leaderthen the construct may be captured by applying a positive voltage potential across the nanopore and if the construct is not a construct of interest it may be ejected from the nanopore by applying a negative voltage potential across the nanopore. (Of course the opposite setup can also be used; for example, if the leader is a positively charged leader then the construct may be captured by applying a negative voltage potential across the nanopore and if the construct is not a construct of interest it may be ejected from the nanopore by applying a positive voltage potential across the nanopore.) Other ejection methods may also be used.
[0166] If the construct is of interest then its movement with respect to the nanopore may be controlled in accordance with the methods described herein. This may in some embodiments be described as “reading” the construct.
[0167] One or more measurements characteristic of the construct may be taken during this reading, i.e. as the protein translocase controls the movement of the construct into the nanopore. Taking one or more measurements in this way may comprise analysing a signal that may be recorded across the nanopore during such movement. For example, the signal may be an signal from an optical or electrical measurement. Suitable measurements are described in more detail herein.
[0168] After data acquisition, as shown in Figure 3, the user of the methods may consider whether or not more data is required. If more data is required then the construct may be further characterised as described herein. If more data is not required then the construct may be ejected from the nanopore as described above.
[0169] In some embodiments of the disclosed methods, the protein translocase controls the movement of a portion of the target polypeptide with respect to the nanopore. In some embodiments the protein translocase controls the movement of from about 10% to about 100% (e.g. from about 20% to about 99%, such as from about 50% to about 97%) of the target polypeptide with respect to the nanopore. In some embodiments the protein translocase controls the movement of at least 10%, at least 20%, at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 97%, at least 98% or at least 99% of the target polypeptide with respect to the nanopore. In some embodiments the protein translocase controls the entire length of the target polypeptide with respect to the nanopore.In some embodiments the portion of the target polypeptide controlled by the protein translocase has a length of at least 5, at least 10, at least 20, at least 30, at least 40, at least 50, or at least 100 amino acids. In some embodiments the portion of the target polypeptide controlled by the protein translocase has a length of at most about 100,000 amino acids, such as at most about 50,000 amino acids, e.g. at most about 10,000 amino acids, such as at most about 5000 amino acids, e.g. at most about 4000 amino acids, e.g. at most about 3000 amino acids, e.g. at most about 2000 amino acids, e.g. at most about 1000 amino acids. Thus, in some embodiments the portion of the polypeptide has a length of from about 5 to about 50,000 amino acids, such as from about 10 to about 10,000 amino acids, e.g. from about 20 to about 5000 amino acids, e.g. from about 50 to about 1000 amino acids.
[0170] Applied forces during movement
[0171] As explained above, in some embodiments of the disclosed methods, a force may be applied across the nanopore. In some embodiments the construct moves in the direction from the first opening of the nanopore to the second opening of the nanopore under an applied force.
[0172] The force can be controlled in order to control the methods. For example, by increasing the force the movement of the construct with respect to the nanopore can be increased or decreased, e.g. the rate at which the construct moves through the nanopore can be controlled.
[0173] In the methods provided herein, any suitable force can be applied. The force may be a potential applied across the first and second openings of the nanopore. In some embodiments, no external force is applied across the first and second openings of the nanopore. For example, in some embodiments no electrical potential is applied. Such embodiments are in some embodiments particularly suited to methods in which optical measurements are taken as the construct moves with respect to the nanopore.
[0174] In some embodiments, the movement of the construct is driven by a physical or chemical force (potential). In some embodiments the chemical force is provided by a concentration (e.g. pH) gradient.
[0175] In some embodiments the physical force is provided by an electrical (e.g. voltage) potential or a temperature gradient, etc. In some embodiments, the force maybe a voltage potential applied across the first and second openings of the nanopore. A voltage potential may be applied using any suitable apparatus, such as an apparatus described herein. Suitable voltage potentials are described in more detail herein.
[0176] In some embodiments the force is applied across a membrane in which the nanopore is embedded. The force is typically applied from the cis side to the trans side of the membrane; i.e. from the cis side to the trans side of the nanopore. The force may be a positive voltage applied across the nanopore or a negative voltage applied across the nanopore.
[0177] As explained herein, the protein translocase typically controls the movement of the construct with the applied force. Thus, in some embodiments, the construct moves in the direction from the first opening of the nanopore to the second opening of the nanopore with the applied force and the protein translocase controls the movement of the construct in the direction from the first opening of the nanopore to the second opening of the nanopore with the applied force.
[0178] Typically the force is a positive voltage applied across the first and second openings of the nanopore such that the trans side of the pore is positive relative to the cis side of the pore. In such embodiments, the force thus attracts negatively charged moi eties such as a negatively charged leader and / or a negatively charged polypeptide (e.g. a polypeptide modified to increase the net negative charge of the polypeptide as described herein) comprised in the construct to move from the cis side to the trans side of the pore. In such embodiments, the methods provided herein typically comprise using the protein translocase at the cis side of the pore to control the movement of the construct in the direction from the cis side of the pore to the trans side of the nanopore with the applied force; i.e. in the same direction as the applied force. Of course, if the leader is positively charged and / or the polypeptide is positively charged (e.g. if the polypeptide is modified to increase the net positive charge of the polypeptide as described herein) then a positive potential applied across the first and second openings of the nanopore may attract the construct to move from the trans side to the cis side of the pore. In such embodiments, the methods provided herein typically comprise using the protein translocase at the trans side of the pore to control the movement of the construct in the direction from the trans side of the pore to the cis side of the nanopore with the applied force; i.e. in the same direction as the applied force.In other embodiments the force is a negative voltage applied across the first and second openings of the nanopore such that the trans side of the pore is negative relative to the cis side of the pore. In such embodiments, the force thus attracts positively charged moi eties such as a positively charged leader and / or a positively charged polypeptide (e.g. if the polypeptide is modified to increase the net positive charge of the polypeptide as described herein) comprised in the construct to move from the cis side to the trans side of the pore. In such embodiments, the methods provided herein typically comprise using the protein translocase at the cis side of the pore to control the movement of the construct in the direction from the cis side of the pore to the trans side of the nanopore with the applied force; i.e. in the same direction as the applied force. Of course, if the leader is negatively charged and / or the polypeptide is negatively charged (e.g. if the polypeptide is modified to increase the net negative charge of the polypeptide as described herein) then a negative potential applied across the first and second openings of the nanopore may attract the construct to move from the trans side to the cis side of the pore. In such embodiments, the methods provided herein typically comprise using the protein translocase at the trans side of the pore to control the movement of the construct in the direction from the trans side of the pore to the cis side of the nanopore with the applied force; i.e. in the same direction as the applied force.
[0179] Polypeptide
[0180] As explained above, the disclosed methods comprise moving a polypeptide within a construct with respect to a nanopore using a protein translocase. In some embodiments the polypeptide may be referred to as a target polypeptide as it is the subject of the disclosed methods. In some embodiments the methods comprise characterising a polypeptide (i.e. characterising a target polypeptide).
[0181] Any suitable polypeptide can be addressed in the disclosed methods.
[0182] As used herein, the term polypeptide is used synonymously with terms such as peptide, oligopeptide, protein or protein fragment, unless implied otherwise by the context.
[0183] In some embodiments the target polypeptide is a polypeptide comprised in a cellular lysate, an extracellular fluid, a reaction mixture, or a biological fluid.Accordingly, in some embodiments the test sample is a cellular lysate, an extracellular fluid, a reaction mixture, or a biological fluid.
[0184] In some embodiments the target polypeptide is the product of a proteolysis reaction. Accordingly, in some embodiments the test sample is subject to proteolysis in the disclosed methods.
[0185] In some embodiments the polypeptide is secreted from cells. Alternatively, the polypeptide can be produced inside cells such that it must be extracted from cells for characterisation by the disclosed methods. The polypeptide may comprise the products of cellular expression of a plasmid, e.g. a plasmid used in cloning of proteins in accordance with the methods described in Sambrook et al., Molecular Cloning: A Laboratory Manual, 4thed., Cold Spring Harbor Press, Plainsview, New York (2012); and Ausubel et al., Current Protocols in Molecular Biology (Supplement 114), John Wiley & Sons, New York (2016).
[0186] In some embodiments, for example, the polypeptide is obtained by expression in a bacterial cell, a yeast cell, an insect cell or a mammalian cell; or by a cell free expression method such as a translation system selected from rabbit reticulocyte lysate, wheat germ extract, and E. coli cell-free systems (available commercially, such as from the PURExpress® systems available from New England Biolabs (Ipswich, MA, USA)). In some embodiments the expression is from the genomic DNA of an organism.
[0187] In some embodiments the polypeptide is a polypeptide product of a chemical or biochemical reaction. For example, the polypeptide may be the product of a ligation reaction or a proteolysis reaction. The polypeptide may be the product of a chemical reaction or an enzymatic reaction. The polypeptide may be the main product of the reaction or a side product of the reaction.
[0188] The polypeptide may be obtained from or extracted from any organism or microorganism. The polypeptide may be obtained from a human or animal, e.g. from urine, lymph, saliva, mucus, seminal fluid or amniotic fluid, or from whole blood, plasma or serum. The polypeptide may be obtained from a plant e.g. a cereal, legume, fruit or vegetable.
[0189] The polypeptide can be provided as an impure mixture of one or more polypeptides and one or more impurities. Impurities may comprise truncated forms of the polypeptide which are distinct from the “target polypeptides” for characterisation inthe disclosed methods. For example, the target polypeptide may be a full length protein and impurities may comprise fractions of the protein. Impurities may also comprise proteins other than the target protein e.g. which may be co-purified from a cell culture or obtained from a sample.
[0190] A polypeptide may comprise any combination of any amino acids, amino acid analogs and modified amino acids (i.e. amino acid derivatives). Amino acids (and derivatives, analogs etc) in the polypeptide can be distinguished by their physical size and charge.
[0191] The amino acids / derivatives / analogs can be naturally occurring or artificial. In some embodiments the polypeptide may comprise any naturally occurring amino acid. Twenty amino acids are encoded by the universal genetic code. These are alanine (A), arginine (R), asparagine (N), aspartic acid (D), cysteine (C), glutamic acid / glutamate (E), glutamine (Q), glycine (G), histidine (H), isoleucine (I), leucine (L), lysine (K), methionine (M), phenylalanine (F), proline (P), serine (S), threonine (T), tryptophan (W), tyrosine (Y) and valine (V). Other naturally occurring amino acids include selenocysteine and pyrrolysine.
[0192] In some embodiments, prior to being modified in accordance with the disclosed methods the polypeptide is an unmodified protein or a portion thereof, or a naturally occurring polypeptide or a portion thereof. In some embodiments the target polypeptide is modified with one or more modifications separate to those made in accordance with the disclosed methods. For example, the target polypeptide may be modified separately to being modified in accordance with the disclosed methods by contacting the target polypeptide with a chaotrope such as high molar concentrations of urea (e.g. about 8M urea), guanidinium chloride (e.g. about 6 M guanidinium chloride), etc.
[0193] In some embodiments the disclosed methods are for characterising modifications in the target polypeptide (apart from those made to alter the charge of the polypeptide in accordance with the disclosed methods). For example, in some embodiments one or more of the amino acids / derivatives / analogs in the polypeptide is post-translationally modified. As such, the methods disclosed herein can be used to detect the presence, absence, number of positions of post-translational modifications in a polypeptide. The disclosed methods can be used to characterise the extent to which a polypeptide has been post-translationally modified.Any one or more post-translational modifications may be present in the polypeptide. Typical post-translational modifications include modification with a hydrophobic group, modification with a cofactor, addition of a chemical group, glycation (the non-enzymatic attachment of a sugar), biotinylation and pegylation. Post-translational modifications can also be non-natural, such that they are chemical modifications done in the laboratory for biotechnological or biomedical purposes. This can allow monitoring the levels of the laboratory made peptide, polypeptide or protein in contrast to the natural counterparts.
[0194] Examples of post-translational modification with a hydrophobic group include myristoylation, attachment of myristate, a Ci4 saturated acid; palmitoylation, attachment of palmitate, a Ci6 saturated acid; isoprenylation or prenylation, the attachment of an isoprenoid group; famesylation, the attachment of a farnesol group; geranylgeranylation, the attachment of a geranylgeraniol group; and glypiation, and glycosylphosphatidylinositol (GPI) anchor formation via an amide bond.
[0195] Examples of post-translational modification with a cofactor include lipoylation, attachment of a lipoate (Cs) functional group; flavination, attachment of a flavin moiety (e.g. flavin mononucleotide (FMN) or flavin adenine dinucleotide (FAD)); attachment of heme C, for instance via a thioether bond with cysteine; phosphopantetheinylation, the attachment of a 4'-phosphopantetheinyl group; and retinylidene Schiff base formation.
[0196] Examples of post-translational modification by addition of a chemical group include acylation, e.g. O-acylation (esters), N-acylation (amides) or S-acylation (thioesters); acetylation, the attachment of an acetyl group for instance to the N-terminus or to lysine; formylation; alkylation, the addition of an alkyl group, such as methyl or ethyl; methylation, the addition of a methyl group for instance to lysine or arginine; amidation; butyrylation; gamma-carboxylation; glycosylation, the enzymatic attachment of a glycosyl group for instance to arginine, asparagine, cysteine, hydroxylysine, serine, threonine, tyrosine or tryptophan; polysialylation, the attachment of polysialic acid; malonylation; hydroxylation; iodination; bromination; citrulination; nucleotide addition, the attachment of any nucleotide such as any of those discussed above, ADP ribosylation; oxidation; phosphorylation, the attachment of a phosphate group for instance to serine, threonine or tyrosine (O-linked) or histidine (N-linked);adenylyl ati on, the attachment of an adenylyl moiety for instance to tyrosine (O-linked) or to histidine or lysine (N-linked); propionylation; pyroglutamate formation; S-glutathionylation; Sumoylation; S-nitrosylation; succinylation, the attachment of a succinyl group for instance to lysine; sei enoyl ati on, the incorporation of selenium; and ubiquitinilation, the addition of ubiquitin subunits (N-linked).
[0197] It is within the scope of the methods provided herein that the polypeptide is labelled with a molecular label. A molecular label may be a modification to the polypeptide which promotes the detection of the polypeptide in the methods provided herein. For example the label may be a modification to the polypeptide which alters the signal obtained as construct is characterised. For example, the label may interfere with a flux of ions through the nanopore. In such a manner, the label may improve the sensitivity of the methods.
[0198] In some embodiments the polypeptide contains one or more cross-linked sections, e.g. C-C bridges. In some embodiments the polypeptides is not cross-linked prior to being characterised using the disclosed methods.
[0199] In some embodiments the polypeptide comprises sulphide-containing amino acids and thus has the potential to form disulphide bonds. Typically, in such embodiments, the polypeptide is reduced using a reagent such as DTT (Dithiothreitol) or TCEP (tris(2-carboxyethyl)phosphine) prior to being characterised using the disclosed methods.
[0200] In some embodiments the polypeptide is a full length protein or naturally occurring polypeptide. In some embodiments a protein or naturally occurring polypeptide is fragmented prior to conjugation to the leader to form the construct for use in the disclosed methods. In some embodiments the protein or polypeptide is chemically or enzymatically fragmented. In some embodiments polypeptides or polypeptide fragments can be conjugated to form a longer target polypeptide.
[0201] The polypeptide can be a polypeptide of any suitable length. In some embodiments the polypeptide has a length of at least 5, at least 10, at least 20, at least 30, at least 40, at least 50, at least 100, at least 200, at least 300, at least 400 or at least 500 amino acids. In some embodiments the length of the polypeptide may be expressed as peptide units, with each peptide unit typically comprising an amino acids.
[0202] Accordingly, in some embodiments the polypeptide has a length of at least 5, at least10, at least 20, at least 30, at least 40, at least 50, at least 100, at least 200, at least 300, at least 400 or at least 500 peptide units. In some embodiments the target polypeptide has a length of from about 2 peptide units to about 100, about 200, about 300, about 400, about 500, about 1000, about 5,000, about 10,000, about 15,000, about 20,000, about 30,000 or about 40,000 peptide units. In some embodiments the target polypeptide has a length of at least 25, at least 30, at least 40, at least 50, at least 100, at least 150, at least 200, at least 300 or at least 400 amino acids. In some embodiments the target polypeptide has a length of from about 50 to about 40,000 peptide units. In some embodiments the target polypeptide has a length of from about 50 to about 35,000 peptide units. In some embodiments the target polypeptide has a length of from about 50 to about 20,000 peptide units. In some embodiments the target polypeptide has a length of from about 50 to about 10,000 peptide units. In some embodiments the target polypeptide has a length of from about 100 to about 5,000 peptide units, for example from about 200 to about 1000 peptide units, e.g. from about 300 to about 500 peptide units, such as from about 200 to about 400 peptide units.
[0203] Any number of target polypeptides can be used in the disclosed methods. For instance, the method may comprise processing 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 50, 100, 200, 500, 1000, 2000 or more target polypeptides or about 10, about 50, about 100, about 200, about 500, about 1000, or about 2000 target polypeptides. If two or more target polypeptides are used, they may be different target polypeptides or two or more instances of the same target polypeptide.
[0204] In some embodiments the target polypeptide to be processed in accordance with the disclosed methods is present in a sample. In some embodiments the sample comprises a plurality of different polypeptides. In some embodiments the plurality may comprise at least 2, at least 3, at least 4, at least 5, at least 10, at least 20, at least 50, at least 100, at least 500, at least 1000, at least 5000, at least 10000, at least 100000, at least 1000000 or more polypeptides.
[0205] Modification of the polypeptide
[0206] As explained here, the disclosed methods comprise modifying the target polypeptide to increase the net charge of the polypeptide.As those skilled in the art will appreciate, increasing the net charge of the polypeptide may comprise making the polypeptide more positively charged or more negatively charged. Thus, in some embodiments modifying the target polypeptide to increase the net charge of the polypeptide comprises increasing the positive charge of the polypeptide. In some embodiments modifying the target polypeptide to increase the net charge of the polypeptide comprises increasing the negative charge of the polypeptide.
[0207] In some embodiments the target polypeptide before modification has a net positive charge and the disclosed methods comprise modifying the target polypeptide to increase the net positive charge of the polypeptide. In some embodiments the target polypeptide before modification has no net positive charge and the disclosed methods comprise modifying the target polypeptide to increase the net positive charge of the polypeptide. In some embodiments the target polypeptide before modification has no net charge (in some embodiments the target polypeptide before modification is neutrally charged) and the disclosed methods comprise modifying the target polypeptide to increase the net positive charge of the polypeptide. In some embodiments the target polypeptide before modification has a net negative charge and the disclosed methods comprise modifying the target polypeptide to increase the net positive charge of the polypeptide. For example, in some embodiments the target polypeptide before modification has a net negative charge and the disclosed methods comprise modifying the target polypeptide to remove the net negative charge and introduce a net positive charge to the polypeptide. In some embodiments the target polypeptide before modification has a net negative charge and the disclosed methods comprise modifying the target polypeptide to increase the net positive charge of the polypeptide, wherein the target polypeptide is modified so that the overall charge density of the polypeptide is increased.
[0208] In some embodiments the target polypeptide before modification has a net negative charge and the disclosed methods comprise modifying the target polypeptide to increase the net negative charge of the polypeptide. In some embodiments the target polypeptide before modification has no net negative charge and the disclosed methods comprise modifying the target polypeptide to increase the net negative charge of the polypeptide. In some embodiments the target polypeptide before modification has nonet charge (in some embodiments the target polypeptide before modification is neutrally charged) and the disclosed methods comprise modifying the target polypeptide to increase the net negative charge of the polypeptide. In some embodiments the target polypeptide before modification has a net positive charge and the disclosed methods comprise modifying the target polypeptide to increase the net negative charge of the polypeptide. For example, in some embodiments the target polypeptide before modification has a net positive charge and the disclosed methods comprise modifying the target polypeptide to remove the net positive charge and introduce a net negative charge to the polypeptide. In some embodiments the target polypeptide before modification has a net positive charge and the disclosed methods comprise modifying the target polypeptide to increase the net negative charge of the polypeptide, wherein the target polypeptide is modified so that the overall charge density of the polypeptide is increased.
[0209] In some embodiments increasing the net charge of the polypeptide comprises increasing the charge density of the polypeptide. In some embodiments, the charge density of the polypeptide is identified as the charge on the polypeptide per unit length of the polypeptide.
[0210] In some embodiments increasing the net charge of the polypeptide comprises increasing the negative charge density of the polypeptide. In some embodiments, therefore, the disclosed methods comprise increasing the negative charge of the polypeptide per unit length of the polypeptide. In some embodiments increasing the net charge of the polypeptide comprises increasing the positive charge density of the polypeptide. In some embodiments, therefore, the disclosed methods comprise increasing the positive charge of the polypeptide per unit length of the polypeptide.
[0211] In some embodiments one or more amino acids in the polypeptide are modified. In some embodiments the one or more amino acids in the polypeptide which are modified are referred to as target amino acids.
[0212] As described herein, in some embodiments a plurality of amino acids in the target polypeptide are modified. In some embodiments the disclosed methods comprise modifying at least one, at least two, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 20, at least 50, at least 100, at least 500, at least1000, at least 5000, at least 10000, or more amino acids in the or each target polypeptide.
[0213] The one or more target amino acids which are modified in the disclosed methods may be located at any position within the target polypeptide.
[0214] In some embodiments the target amino acids which are modified in the disclosed methods may be identified by consideration of the native (unmodified) structure of the target polypeptide. For example, in some embodiments the target amino acids can be chosen to perturb charge-charge interaction in the unmodified peptide: for instance, it can in some embodiments be inferred that a positively charged amino acid may interact with a negatively charged amino acid if the amino acids are located proximate to each other in the sequence, and such amino acids may be selected for modification in some embodiments. However, in some embodiments the amino acids for modifying in the disclosed methods are distributed randomly or pseudo-randomly throughout the target polypeptide. For example, the disclosed methods may comprise selectively modifying one or more types of amino acid in a polypeptide and in this case it may not be required to know the location of the amino acids in the polypeptide in advance of the methods. Indeed, in some embodiments the disclosed methods can be used to characterise (e.g. to locate) the modified amino acids as such modifications may be detectable in the disclosed detection methods.
[0215] In some embodiments the amino acids in the target polypeptide that are modified are separated from each other by between about 1 and about 20 amino acids, such as between about 2 and about 15 amino acids e.g. between about 3, about 4, or about 5 amino acids and about 8, about 9 or about 10 amino acids. In some embodiments the amino acids that are modified are adjacent to each other. In some embodiments the amino acids in the in the target polypeptide that are modified are separated from each other by more than about 20, such as more than about 30, e.g. more than about 40 or more than about 50 amino acids. For example, in some embodiments two or more amino acids which interact in the target polypeptide may be modified, wherein the two or more amino acids are spatially proximate in the target polypeptide before its modification. In some embodiments this can occur despite the two or more amino acids being separate from each other in the amino acid sequence of the target polypeptide.In some embodiments modifying one or more amino acids in the polypeptide to increase the net charge of the polypeptide may comprise disrupting the 3D structure of the unmodified polypeptide. In some embodiments this can be referred to as destabilizing the polypeptide, as typically increasing the net charge of the polypeptide may alter the 3D structure of the polypeptide when comprised in the construct. In some embodiments the one or more amino acids that can be modified in accordance with the disclosed methods are selected for modification as described in United Kingdom patent application GB 2500667.7, the entire contents of which are hereby incorporated by reference.
[0216] In some embodiments one or more amino acids which are modified in the disclosed methods interact via one or more charge-mediated interactions with one or more other amino acids in the target polypeptide. In some embodiments the one or more amino acids are charged amino acids. In some embodiments the one or more amino acids are negatively charged and / or positively charged amino acids and the disclosed methods comprise modifying one or more negatively charged amino acids and / or one or more positively charged amino acids in the target polypeptide.
[0217] In some embodiments one or more amino acids which are modified in the disclosed methods via one or more hydrophobic interactions with one or more other amino acids in the target polypeptide. In some embodiments the one or more amino acids are uncharged and / or aromatic amino acids. In some embodiments the disclosed methods comprise modifying one or more uncharged and / or one or more aromatic amino acids in the target polypeptide.
[0218] In some embodiments one or more amino acids which are modified in the disclosed methods interact via hydrogen-bonding with one or more other amino acids in the target polypeptide. In some embodiments the one or more amino acids are polar amino acids. In some embodiments amino acids suitable for modification in such methods have side chains comprising hydrogen-bond donor groups or hydrogen bond acceptor groups. In some embodiments the disclosed methods comprise modifying one or more amino acids having hydrogen-bond donor groups and / or one or more amino acids having hydrogen bond acceptor groups in the target polypeptide.In some embodiments one or more amino acids which are modified in the disclosed methods interact via one or more covalent interactions (e.g. disulphide bonds) with one or more other amino acids in the target polypeptide. In some embodiments amino acids suitable for modification in such methods are include cysteine and nonnatural thiol -containing amino acids. In some embodiments the disclosed methods comprise modifying one or more cysteine amino acids and / or one or more non-natural thiol-containing amino acids in the target polypeptide.
[0219] In some embodiments each occurrence of a target amino acid in a polypeptide may be modified in the disclosed methods. For example, as described in more detail herein, in some embodiments chemical reagents may be used to modify the side chains of each occurrence of an amino acid in the target polypeptide. By way of illustration, in some embodiments the chemical reagent is an amine-reactive reagent and the method may comprise modifying of each occurrence of lysine in the target amino acid. In some embodiments the chemical reagent is a carboxylic acid- (or carboxylate)-reactive reagent and the method may comprise modifying of each occurrence of glutamic acid (glutamate) and aspartic acid (aspartate) in the target amino acid.
[0220] In some embodiments the one or more amino acids that are modified in the disclosed methods are each comprised in a target motif. In some embodiments the target motif comprises from about 2, 3, 4, 5, 6, 7, 8, 9 or 10 amino acids comprising the amino acid to be modified. For example, in some embodiments the target motif may comprise or consist of a recognition sequence for an amino-acid modifying enzyme which may be used to modify the target polypeptide. Through appropriate selection of reagents, the skilled operator of the methods is capable of modifying specific amino acids in the target polypeptide.
[0221] The choice or selection of amino acids to be modified in the disclosed methods is not particularly limited. In some embodiments, the or each amino acid modified in the provided methods is selected from Ser, Thr, Tyr, Cys, Lys, Arg, His, Asp, Glu, Pro, Asn, Gin, Met, Trp, Phe, Ala, Vai, Leu and He. In some embodiments the or each amino acid modified in the provided methods is selected from Ser, Thr, Tyr, Cys, Lys, Arg, His, Asp, Glu, Asn, Gin, Met, and Trp. In some embodiments the or each amino acid modified in the provided methods is selected from Ser, Thr, Tyr, Cys, Lys, Arg,Asp, Glu, Asn, and Gin. In some embodiments the or each amino acid modified in the provided methods is selected from Cys, Lys, Arg, Asp, Glu, Asn, and Gin.
[0222] In some embodiments modifying the polypeptide comprises modifying one or more target amino acids with one or more charge-modifying moieties.
[0223] In some embodiments the one or more target amino acids which are modified in the disclosed methods comprise one or more charged amino acids. In some embodiments the one or more amino acids that are modified with one or more chargemodifying moieties are selected from histidine (His), lysine (Lys), arginine (Arg), glutamic acid / glutamate (Glu) and aspartic acid / aspartate (Asp). In some embodiments the one or more amino acids that are modified with one or more charge-modifying moieties are selected from Lys, Arg, Glu and Asp.
[0224] In some embodiments the one or more amino acids that are modified with one or more charge-modifying moieties comprise one or more positively charged amino acids and increasing the net charge of the polypeptide comprises reducing the positive charge of said amino acids, thereby increasing the net negative charge of the polypeptide.
[0225] Thus, in some embodiments the one or more amino acids that are modified with one or more charge-modifying moieties are selected from His, Lys and Arg, and the method comprises modifying the side chain of said one or more amino acids to reduce the positive charge of the said amino acids. In some embodiments the one or more amino acids are one or more lysines. In some embodiments modifying said positively charged amino acids may comprise neutralising the positive charge of said amino acids. In some embodiments the modifying said positively charged amino acids may comprise introducing negatively charged groups to said amino acids thereby making said amino acids negatively charged.
[0226] In some embodiments the one or more amino acids that are modified with one or more charge-modifying moieties comprise one or more negatively charged amino acids and increasing the net charge of the polypeptide comprises reducing the negative charge of said amino acids, thereby increasing the net positive charge of the polypeptide. Thus, in some embodiments the one or more amino acids that are modified with one or more charge-modifying moieties are selected from Asp and Glu, and the method comprises modifying the side chain of said one or more amino acids to reduce the negative charge of the said amino acids. In some embodiments modifying said negatively chargedamino acids may comprise neutralising the negative charge of said amino acids. In some embodiments the modifying said negatively charged amino acids may comprise introducing positively charged groups to said amino acids thereby making said amino acids positively charged.
[0227] In some embodiments both one or more positively charged amino acids (e.g. one or more amino acids selected from His, Lys and Arg, particularly Lys) and one or more negatively charged amino acids (e.g. one or more amino acids selected from Asp and Glu) are modified as described herein. This is compatible with the disclosed methods wherein the net charge of the overall polypeptide is increased.
[0228] In some embodiments the one or more amino acids that are modified with one or more charge-modifying moieties comprise one or more amino acids that stabilize the secondary and / or tertiary structure of the target polypeptide by hydrogen bonding. In some embodiments the one or more amino acids that are modified with one or more charge-modifying moieties are selected from Ser, Thr, Tyr, Asn and Gin and the method comprises increasing the charge of said amino acids. In some embodiments modifying said amino acids comprises increasing the positive charge of said amino acids. In some embodiments modifying said amino acids may comprise introducing positively charged groups to said amino acids thereby making said amino acids positively charged. In some embodiments modifying said amino acids comprises increasing the negative charge of said amino acids. In some embodiments modifying said amino acids may comprise introducing negatively charged groups to said amino acids thereby making said amino acids negatively charged.
[0229] In some embodiments the one or more amino acids that are modified with one or more charge-modifying moieties comprise one or more amino acids that stabilize the secondary and / or tertiary structure of the target polypeptide by hydrophobic interactions. In some embodiments the one or more amino acids that are modified with one or more charge-modifying moieties are selected from Ala, Vai, Leu, He, Met, Trp and Phe, particularly Trp and Phe which can engage in pi-pi interactions. In some embodiments modifying said amino acids comprises introducing positive charge to said amino acids. In some embodiments modifying said amino acids may comprise introducing positively charged groups to said amino acids thereby making said amino acids positively charged. In some embodiments modifying said amino acids comprisesintroducing negative charge to said amino acids. In some embodiments modifying said amino acids may comprise introducing negatively charged groups to said amino acids thereby making said amino acids negatively charged.
[0230] In some embodiments, one or more amino acids of the target polypeptide are modified by modifying the side chains of said amino acids with one or more chargemodifying moieties.
[0231] Any suitable charge-modifying moiety can be used. The charge-modifying moiety alters the native charge of the one or more amino acids as described herein, (e.g. to introduce charge to a neutral amino acid; to decrease the positive charge of a positively charged amino acid or to decrease the negative charge of a negatively charged amino acid).
[0232] Suitable charge-modifying moieties are thus typically capable of reacting with functional groups present in the side chain of the one or more amino acids.
[0233] In some embodiments a charge-modifying moiety may comprise a reactive group for reacting with the side chain of the amino acid being modified. The reactive group can be chosen or determined by the user of the disclosed methods to ensure selective reaction with the or each amino acid.
[0234] In some embodiments the or each amino acid natively comprises a reactive side chain for reacting with a charge-modifying moiety. In some embodiments the or each amino acid is modified to comprise a reactive side chain for reacting with a chargemodifying moiety. In some embodiments the methods comprise activating the side chain of the or each first amino acid for reaction with a charge-modifying moiety.
[0235] For example, in some embodiments the side chain of the amino acid comprises a nucleophilic group and the charge-modifying moiety comprises an electrophilic group. In some embodiments the side chain of the amino acid comprises an electrophilic group and the charge-modifying moiety comprises a nucleophilic group.
[0236] In some embodiments modifying the side chains of said one or more amino acids with one or more charge-modifying moieties comprises covalently attaching said one or more charge-modifying moieties to said side chains. For example, in some embodiments modifying the side chains of said one or more amino acids with one or more charge-modifying moieties comprises one or more of oxidation, esterification, O-glycosylation, alkylation, nitration, nitrosylation, succinylation, hydroxylation, amidation, deamidation, acylation, sulfhydration, carb amyl ati on, glycation, metal coordination, Schiff-base formation, and ADP-ribosylation of said side chains.
[0237] Reagents to achieve these modifications are well known to those skilled in the art. Some exemplary reagents are described herein.
[0238] In some embodiments, modifying the side chains of said one or more amino acids with one or more charge-modifying moieties comprises contacting said one or more amino acids with one or more chemical reagents. The attachment chemistry between the side chain of the amino acid and the charge-modifying moiety is not particularly limited and any suitable chemistry can be used. For example, in some embodiments the one or more chemical reagents are selected from oxidising agents, halogenating agents, nitrating agents, alkylating agents, phosphoric acid derivatives, glycosyl donors, acids, bases, NO donors, sulphide donors, glycation agents, nucleotides, acetyl anhydride, glyoxal, formaldehyde, cyanate, metal ions, succinic anhydride, hydroxyl radicals, and fatty acids. Practitioners are directed to texts such as March's Advanced Organic Chemistry: Reactions, Mechanisms, and Structure (2019), ed. Smith, Wiley; and to G. Hermanson, Bioconjugate Techniques, 3rd Edition (2013), each of which are hereby incorporated by reference in their entirety. Practitioners are directed particularly to discussion in those texts of reactions of amines and guanidines, carboxylic acids, alcohols, and thiols, especially amines and guanidines.
[0239] Reactive groups and their corresponding targets include aryl azides which may react with amine, carbodiimides which may react with amines and carboxyl groups, hydrazides which may react with carbohydrates, hydroxmethyl phosphines which may react with amines, imidoesters which may react with amines, isocyanates which may react with hydroxyl groups, carbonyls which may react with hydrazines, maleimides which may react with sulfhydryl groups, NHS-esters which may react with amines, PFP-esters which may react with amines, psoralens which may react with thymine, pyridyl disulfides which may react with sulfhydryl groups, vinyl sulfones which may react with sulfhydryl amines and hydroxyl groups, vinylsulfonamides, and the like. In some embodiments a reactive group (e.g. a first and / or second reactive group) is selected from NHS esters, maleimides, imido esters, aldehydes, hydrazides, epoxides,isocyanates, activated carboxylic acids, azides, thiols, iodoacetamides, bromoacetamides, diazirines, photoreactive benzophenones, alkyl halides, succinimides, pyridyldisulfides, amines, carbodiimides, and oxiranes.
[0240] Other suitable chemistry for attaching a charge-modifying moiety to the side chain of an amino acid includes click chemistry. Many suitable click chemistry reagents are known in the art. Suitable examples of click chemistry include, but are not limited to, the following:
[0241] (a) copper(I)-catalyzed azide-alkyne cycloadditions (azide alkyne Huisgen cycloadditions);
[0242] (b) strain-promoted azide-alkyne cycloadditions; including alkene and azide [3+2] cycloadditions; alkene and tetrazine inverse-demand Diels- Alder reactions; and alkene and tetrazole photoclick reactions;
[0243] (c) copper-free variant of the 1,3 dipolar cycloaddition reaction, where an azide reacts with an alkyne under strain, for example in a cyclooctane ring such as in bicycle[6.1.0]nonyne (BCN);
[0244] (d) the reaction of an oxygen nucleophile on one linker with an epoxide or aziridine reactive moiety on the other; and
[0245] (e) the Staudinger ligation, where the alkyne moiety can be replaced by an aryl phosphine, resulting in a specific reaction with the azide to give an amide bond. Thus, any reactive group may be used in the reaction between the side chain of the amino acid and the charge-modifying moiety. Some suitable reactive groups include [1, 4-Bis[3-(2-pyridyldithio)propionamido]butane; 1,1 1-bis-maleimidotriethyleneglycol; 3,3’-dithiodipropionic acid di(N-hydroxysuccinimide ester); ethylene glycol-bis(succinic acid N-hydroxysuccinimide ester); 4,4’-diisothiocyanatostilbene-2,2’-disulfonic acid disodium salt; Bis[2-(4-azidosalicylamido)ethyl] disulphide; 3-(2-pyridyldithio)propionic acid N-hydroxysuccinimide ester; 4-maleimidobutyric acid N-hydroxysuccinimide ester; lodoacetic acid N-hydroxysuccinimide ester; S-acetylthioglycolic acid N-hydroxysuccinimide ester; azide-PEG-maleimide; and alkyne-PEG-maleimide. The reactive group may be any of those disclosed in WO 2010 / 086602, particularly in Table 3 of that application.In some embodiments the side chain of the amino acid to be modified comprises an amine group. Examples of such amino acids include lysine. In some embodiments such an amine groups may be reacted with a charge-modifying moieties which comprises an amine-reactive group.
[0246] Examples of amine-reactive groups include: carboxylic acids and activated derivatives thereof (e.g. NHS-esters), which can form amide bonds with amine groups; thiols or activated derivatives thereof, which can form thioether bonds by reaction with amine groups; squaramates, which react with amines to form squaramides; amine coupling agents (e.g. N,N’-Disuccinimidyl carbonate, l,l'-Carbonyldiimidazole) which react with amines to form a urea linkage; aldehydes and ketones, which react with amines via reductive amination (e.g. via reduction using a reducing agent such as NaBEE or NaBEECN); and isothiocyanates (e.g. 4-sulfophenyl isothiocyanate (SPITC)), which react with amines to form thiourea linkages. In some embodiments the disclosed methods comprise modifying amine groups in the target polypeptide (e.g. comprised in lysine amino acid residues of the target polypeptide) using an amine-reactive group as disclosed herein, such as an isothiocyanate group. In some embodiments the disclosed methods comprise modifying amine groups in the target polypeptide (e.g. comprised in lysine amino acid residues of the target polypeptide) using 4-sulfophenyl isothiocyanate (SPITC)). Modifying lysine residues with 4-sulfophenyl isothiocyanate removes the positive charge at each reacted lysine residue and introduces a negative charge.
[0247] In one embodiment an amine group comprised in the side chain of an amino acid may be activated for reaction with the charge-modifying moiety. For example, an amine group (e.g. comprised in the side chain of a lysine amino acid) may be modified by reaction with a maleimide-containing compound such as a maleimide-NHS-ester (e.g. 3-maleimido-propionic NHS ester), with the amine group reacting with the NHS-ester to form an amide, optionally followed by reaction of the free maleimide group with a further moiety to alter the charge of the residue, for example, with a thiol group, or with a diene such as a furan group. The charge-modifying moiety may be neutral and thus disrupt the charge of the lysine by reducing the positive charge of the side chain of the lysine, or may be negatively charged (e.g. may comprise a carboxylic acid moiety) and thus disrupt the charge of the lysine by making the side chain of the lysine negatively charged.In another embodiment, an amine group (e.g. comprised in the side chain of a lysine amino acid) may be modified for subsequent reaction (e.g. with the chargemodifying moiety) in a click chemistry reaction. Examples include copper-catalysed azide / alkyne cycloaddition (CuAAC) reactions. For example, in one embodiment an amine may be activated by reaction with an azide-containing compound, such as an azidoacetic acid NHS-ester, followed by reaction of the free azide group with the charge-modifying moiety, e.g. with an alkyne group of the charge-modifying moiety. In one embodiment an amine may be activated by reaction with an alkyne-containing compound, followed by reaction of the alkyne group with the charge-modifying moiety, for example with an azide group of the charge-modifying moiety, such as an azidoacetic acid NHS-ester. In one embodiment the click chemistry reaction is a strain-promoted azide-alkyne reaction. In one embodiment an amine may be activated by reaction with an azide-containing compound, such as an azidoacetic acid NHS-ester, followed by reaction of the free azide group with the charge-modifying moiety, for example with a cyclooctynyl group of the charge-modifying moiety (e.g. a bicyclononyne, BCN, group). In one embodiment an amine may be activated by reaction with an cyclooctynyl-containing compound, such as a bicyclononyne (BCN) group, followed by reaction of the alkyne group with the charge-modifying moiety, for example with an azide group of the charge-modifying moiety, such as an azidoacetic acid NHS-ester. In one embodiment the click chemistry reaction is an inverse electron demand Diels-Alder reaction (iEDDA). In one embodiment an amine may be activated by reaction with a tetrazine-containing compound, such as a methyltetrazine-containing compound, followed by reaction of the free tetrazine group with the charge-modifying moiety, for example with a trans-cyclooctene group of the charge-modifying moiety. In one embodiment an amine may be activated by reaction with a trans-cycloocten-containing compound, followed by reaction of the TCO group with the charge-modifying moiety, for example with a tetrazine group (e.g. a methyltetrazine group) of the chargemodifying moiety.
[0248] In some embodiments the side chain of the amino acid to be modified comprises a guanidine group. Examples of such amino acids include arginine. In some embodiments such a guanidine group may be reacted with a charge-modifying moiety which comprises a guanidine-reactive group. Examples of guanidine-reactive groupsinclude diketones such as a 1,2-diketone or 1,3 -diketone; and NHS-esters, which can form amide bonds with guanidine groups. In one embodiment a guanidine group may be modified by citrullination via arginine deaminase to form a carbamide group; by using an aldehyde such as formaldehyde; or by reaction with a glyoxal -containing compound. The resulting activated groups may optionally be further modified by reaction with further groups to modify the charge of the residue. Analogous strategies can be used to those described above in the context of activating amine groups.
[0249] In some embodiments the side chain of the amino acid to be modified comprises a carboxyl group. Examples of such amino acids include Asp and Glu. In some embodiments such a carboxyl group may be reacted with a charge-modifying moiety which comprises a carboxyl -reactive group. Examples of carboxyl -reactive groups include amines, which can form amide bonds with carboxyl groups; and alcohols, which can form esters with carboxyl groups. In one embodiment a carboxyl group may be modified using a reagent such as a carbodiimide, followed by reaction with a nucleophilic group (e.g. an amine). In one embodiment a compound comprising a nucleophilic group such as an amine group coupled to a click chemistry group such as an azide or alkyne group may be used, with the carboxyl group being activated by the carbodiimide, the amine group reacting with the activated carboxyl group; and followed by reaction of the free click chemistry group with a charge modifying moiety; for example with a moiety comprising an azide or alkyne group. Analogous strategies can be used to those described above in the context of activating amine groups.
[0250] In some embodiments the side chain of the amino acid to be modified comprises a hydroxyl group. Examples of such amino acids include Ser, Thr and Tyr. In some embodiments such a hydroxyl group may be reacted with a charge-modifying moiety which comprises a hydroxyl-reactive group. Examples of hydroxyl -reactive groups include carboxyl groups which can react with hydroxyl groups to form esters; isocyanates which react with hydroxyl groups to form urethanes; and vinyl sulfones. A hydroxyl group may be activated for further reaction, e.g. by using a compound comprising a carboxyl, isocyanate or vinyl sulphone group coupled to a click chemistry group such as an azide or alkyne group, with the hydroxyl group being activated, followed by reaction of the free click chemistry group; for example with an azide oralkyne group. Analogous strategies can be used to those described above in the context of activating amine groups.
[0251] In some embodiments the side chain of the amino acid to be modified comprises a thiol group. Examples of such amino acids include Cys. In some embodiments such a thiol group may be reacted with a charge-modifying moiety which comprises a thiolreactive group. Examples of thiol -reactive groups include maleimides, haloacetamides, pyridyl disulfides and vinyl sulfones. In one embodiment a thiol group may be activated for further reaction, e.g. using a compound comprising a maleimide group, a haloacetamide or a pyridyl disulfides or vinyl sulfone group coupled to a click chemistry group such as an azide or alkyne group, with the thiol group being activated, followed by reaction of the free click chemistry group; for example with an azide or alkyne group of the linker. Analogous strategies can be used to those described above in the context of activating amine groups.
[0252] In some embodiments the target polypeptide is modified to prevent crossreaction of functional groups in the target polypeptide apart from the target amino acid(s) the charge of which is altered in the disclosed methods. For example, in some embodiments the N-terminal amine group of the target polypeptide is modified, e.g. by being capped. In some embodiments the N-terminal amine group of the target polypeptide is modified by being acetylated. In some embodiments the C-terminal amine group of the target polypeptide is modified, e.g. by being capped. In some embodiments the C-terminal amine group of the target polypeptide is modified by being ami dated. Capping of the N- and / or C-terminals of the target polypeptide can prevent reaction of such groups, which can be useful e.g in embodiments where it may be desirable for the N- and / or C-terminals of the target polypeptide to remain unmodified so that they can be modified with one or more sequencing adapters as described herein.
[0253] In some embodiments, modifying the side chains of said one or more amino acids with one or more charge-modifying moieties comprises contacting said one or more amino acids with one or more amino-acid modifying enzymes.
[0254] Any suitable such enzymes can be used. As those skilled in the art will appreciate, choice of a suitable amino-acid modifying enzyme is an operational parameter of the disclosed methods which can be made by the user of the method according to the polypeptide to be processed.In some embodiments an amino-acid modifying enzyme is selective for a recognition sequence for the amino-acid modifying enzyme in the target polypeptide. Choice of a suitable amino-acid modifying enzyme will include consideration of any recognition sequence for the amino-acid modifying enzyme required. In brief, however, and by way of non-limiting illustration, the statistical distribution of a given recognition sequence for the amino-acid modifying enzyme in a polypeptide comprising a random or pseudo-random arrangement of amino acids (e.g. in a full length protein or other polypeptide as described herein) can be calculated based on the length and sequence of the recognition sequence for the amino-acid modifying enzyme and the composition of the polypeptide. This means that a skilled person can choose appropriate amino-acid modifying enzymes to use according to the frequency of the amino acid that it is desired to modify.
[0255] Suitable amino-acid modifying enzymes include kinases, sulfotransferases, O-GlcNAc transferases (OGTs), methyltransferases, peroxidases, palmitoyltransferases (PATs), nitrosylases, acetyltransferases, hydroxylases, deiminases, succinyltransferases, glutamylases, oligosaccharyltransferases (OSTs), glycosyltransferases, and glutaminase.
[0256] In some embodiments modifying the side chains of said one or more amino acids with one or more charge-modifying moieties comprises conducting a reaction as set out in the table in Figure 4.
[0257] In some embodiments an amino acid comprising an amine group in the side chain (e.g. a lysine amino acid) may be modified by acetylation, alkylation (e.g. methylation), ubiquitination, or hydroxylation. Suitable reagents include acetyl anhydride, alkyl halides (e.g. methyl iodide), glyoxal, aldehydes (e.g. formaldehyde). Suitable enzymes include methyltransferases, acetyltransferases, hydroxylases, and the like.
[0258] In some particular embodiments an amino acid comprising a hydroxyl group in the side chain (e.g. Ser, Thr, Tyr) may be modified by phosphorylation, O-GlcNAcylation, sulfation, alkylation (e.g. methylation), or nitration. Suitable reagents include phosphoric acid derivatives (e.g. ATP), glycosyl donors, oxidising agents (e.g. peroxides such as hydrogen peroxide, hypocholorites such as sodium hypochlorite, permanganates such as potassium permanganate), halogen sources (such as Ch and h), nitrating agents (such as peroxynitrites and nitryl chloride), and Fenton reagents.Suitable enzymes include kinases, sulfotransferases, O-GlcNAc transferases (OGTs), methyltransferases, peroxidases, and the like.
[0259] In some particular embodiments an amino acid comprising a carboxyl group in the side chain (e.g. Asp, Glu) may be modified by succinylation or alkylation (e.g. methylation). Suitable reagents include succinic anhydride, alkyl halides such as methyl iodide, and metal ion salts. Suitable enzymes include methyltransferases and succinyltransferases.
[0260] In some particular embodiments an amino acid comprising an imidazole group in the side chain (e.g. His) may be modified by phosphorylation or alkylation (e.g. methylation). Suitable reagents include nitrating agents and transition metals (e.g. copper and iron salts). Suitable enzymes include methyltransferases and kinases.
[0261] In some particular embodiments an amino acid comprising an guanidine group in the side chain (e.g. Arg) may be modified by alkylation (e.g. methylation) or citrullination. Suitable reagents include cyanates, nitric oxide donors (e.g. NO2, ONO?' etc) and alkyl halides such as methyl iodide. Suitable enzymes include methyltransferases, nitrosylases and deiminases.
[0262] In some particular embodiments an amino acid comprising a thiol group in the side chain (e.g. Cys) may be modified by S-palmitoylation, S-nitrosylation, disulphide bond formation, or sulfydration. Suitable reagents include alkyl halides (e.g. iodoacetamide) NO donors, oxidising agents (e.g. peroxides such as hydrogen peroxide, 02, etc), and sulphide donors such as NaHS. Suitable enzymes include palmitoyltransferases (PATs), nitrosylases, peroxidases, and the like.
[0263] In some particular embodiments an amino acid comprising a secondary amide group in the side chain (e.g. Asn, Gin) may be modified by N-linked glycosylation, ADP-ribosylation, or deamidation. Suitable reagents include aldehydes (e.g. formaldehyde), glycation agents such as glucose and fructose, NAD+, and elevated heat and pH conditions (e.g. treatment with alkali such as sodium hydroxide). Suitable enzymes include glycosyltransferases, glutaminase, ribotransferases, etc.
[0264] In some embodiments the charge of the target polypeptide can be altered or controlled by adjusting the pH. For example, increasing the pH (e.g. to pH above pH 7, e.g. from about pH 8 to about pH 11) can increase the relative negative charge of thepolypeptide. Alternatively, decreasing the pH (e.g. to pH below pH 7, e.g. from about pH 3 to about pH 6) increases the relative positive charge of the polypeptide.
[0265] Formation of the construct
[0266] As explained in more detail herein, the polypeptide is typically comprised in a construct comprising a leader comprising a non-polypeptide component (e.g. the leader may comprise or consist of a polynucleotide) conjugated to the target polypeptide.
[0267] When present, the polypeptide can be conjugated to the leader at any suitable position. For example, the polypeptide can be conjugated to the leader at the N-terminus or the C-terminus of the polypeptide.
[0268] In some embodiments the leader comprises or consists of a polynucleotide and the polypeptide is conjugated to the 3’ end of the polynucleotide of the leader. In some such embodiments the first end of the polypeptide is the C terminus of the polypeptide and the first end of the polypeptide is attached to the 3’ end of the polynucleotide. In some embodiments the first end of the polypeptide is the N terminus of the polypeptide and the first end of the polypeptide is attached to the 3’ end of the polynucleotide.
[0269] In other embodiments the leader comprises or consists of a polynucleotide and the polypeptide is conjugated to the 5’ end of the polynucleotide of the leader. In some embodiments the first end of the polypeptide is the C terminus of the polypeptide and the first end of the polypeptide is attached to the 5’ end of the polynucleotide. In some embodiments the first end of the polypeptide is the N terminus of the polypeptide and the first end of the polypeptide is attached to the 5’ end of the polynucleotide.
[0270] In some embodiments the first end of the polypeptide is attached to the leader and a second end of the polypeptide is attached to a blocking moiety. As explained in more detail herein, the blocking moiety may function to prevent the protein translocase from disengaging from the construct. Additionally or alternatively, the blocking moiety may function to prevent the construct from fully translocating the nanopore. In some embodiments the blocking moiety is comprised in an adapter for attaching at the second end of the polypeptide. This is described in more detail herein. In some embodiments the adapter comprises a first end comprising an attachment point for attaching to a second end of a polypeptide analyte, and second end comprising a blocking moiety.Thus, in some embodiments the construct comprises a polypeptide conjugated to a leader comprising a polynucleotide component wherein a first end of the polypeptide is conjugated to the 3’ end of the polynucleotide of the leader and a second end of the polypeptide is attached to a blocking moiety. In some embodiments the first end of the polypeptide is the C terminus of the polypeptide, the first end of the polypeptide is attached to the 3’ end of the polynucleotide, and the second end of the polypeptide is the N terminus of the polypeptide and is attached to a blocking moiety. In some embodiments the first end of the polypeptide is the N terminus of the polypeptide, the first end of the polypeptide is attached to the 3’ end of the polynucleotide, and the second end of the polypeptide is the C terminus of the polypeptide and is attached to a blocking moiety.
[0271] In other embodiments the construct comprises a polypeptide conjugated to a leader comprising a polynucleotide component wherein a first end of the polypeptide is conjugated to the 5’ end of the polynucleotide of the leader and a second end of the polypeptide is attached to a blocking moiety. In some embodiments the first end of the polypeptide is the C terminus of the polypeptide, the first end of the polypeptide is attached to the 5’ end of the polynucleotide, and the second end of the polypeptide is the N terminus of the polypeptide and is attached to a blocking moiety. In some embodiments the first end of the polypeptide is the N terminus of the polypeptide, the first end of the polypeptide is attached to the 5’ end of the polynucleotide, and the second end of the polypeptide is the C terminus of the polypeptide and is attached to a blocking moiety.
[0272] In some embodiments the leader is directly attached to the first end of the polypeptide. In some embodiments the leader is comprised in an adapter for attaching at the first end of the polypeptide. This is described in more detail herein. In some embodiments the adapter comprises a first end comprising the leader comprising a polynucleotide; and a second end comprising an attachment point for attaching to a first end of a polypeptide analyte. In some embodiments the leader is attached to the first end of the polypeptide by a linker. In some embodiments the adapter comprises a recognition sequence for the protein translocase. In some embodiments the leader is attached to the first end of the polypeptide via a recognition site (e.g. a recognition sequence) for the protein translocase. This is further described herein.The polypeptide may be conjugated to the leader using any suitable chemistry. In some embodiments the leader may comprise a reactive group for reacting with the polypeptide. The reactive group can be chosen or determined by the user of the disclosed methods to ensure selective reaction with the polypeptide.
[0273] In some embodiments the polypeptide comprises a terminal reactive group. In some embodiments the terminal reactive group is comprised in the terminal amino acid of the polypeptide. In some embodiments the terminal reactive group is provided by the N- terminus or the C-terminus of the polypeptide. In some embodiments the terminal reactive group is provided by the side chain of the N-terminal or C-terminal amino acid of the polypeptide. Thus in some embodiments the terminal reactive group is comprised in the terminal amino acid or amino acid analog of the polypeptide.
[0274] In some embodiments the leader is attached to the polypeptide via the terminal reactive group. In some embodiments the leader is attached to the polypeptide via a linker. In some embodiments the leader is attached to the polypeptide via a linker attached to the terminal reactive group.
[0275] In some embodiments a linker may be attached to a side chain of an amino acid in the polypeptide. In some embodiments a linker may be attached to a terminal amino acid of the polypeptide. In some embodiments a linker may be attached to the N-terminus of the polypeptide. In some embodiments a linker may be attached to the N-terminal amine group of the polypeptide. In some embodiments a linker may be attached to the C-terminus of the polypeptide. In some embodiments a linker may be attached to the C-terminal carboxy group of the polypeptide.
[0276] In some embodiments a side chain of an amino acid of the polypeptide (e.g. a terminal amino acid of the polypeptide) comprises a nucleophilic group and the leader comprises an electrophilic group. In some embodiments the side chain of the amino acid (e.g. a terminal amino acid of the polypeptide) comprises an electrophilic group and the leader comprises a nucleophilic group.
[0277] In some embodiments the leader is covalently attached to the polypeptide.
[0278] Practitioners are directed to texts such as March's Advanced Organic Chemistry:
[0279] Reactions, Mechanisms, and Structure (2019), ed. Smith, Wiley; and to G. Hermanson, Bioconjugate Techniques, 3rd Edition (2013), each of which are hereby incorporated byreference in their entirety. Practitioners are directed particularly to discussion in those texts of reactions of amines and guanidines, carboxylic acids, alcohols, and thiols.
[0280] In some embodiments the leader comprises an amine-reactive group and is conjugated to the polypeptide at the N-terminus of the polypeptide. Examples of amine-reactive groups include: carboxylic acids and activated derivatives thereof (e.g. NHS-esters), which can form amide bonds with amine groups; thiols or activated derivatives thereof, which can form thioether bonds by reaction with amine groups; squaramates, which react with amines to form squaramides; amine coupling agents (e.g. N,N’-Disuccinimidyl carbonate, l,l'-Carbonyldiimidazole) which react with amines to form a urea linkage; and aldehydes and ketones, which react with amines via reductive amination (e.g. via reduction using a reducing agent such as NaBEU or NaBHjCN). Amine groups may be activated for reaction with groups such as carboxylic acids, e.g. by reacting the amine group with an NHS ester. Amine groups may be modified for subsequent reaction in a click chemistry reaction. Examples include copper-catalysed azide / alkyne cycloaddition (CuAAC) reactions. For example, in one embodiment an amine may be activated by reaction with an azide-containing compound, such as an azidoacetic acid NHS-ester, followed by reaction of the free azide group with the leader, e.g. with an alkyne group of the leader. Alternatively, an amine may be activated by reaction with an alkyne-containing compound, followed by reaction of the alkyne group with the leader, for example with an azide group of the leader, such as an azidoacetic acid NHS-ester. Other suitable click chemistry is discussed above.
[0281] In some embodiments the leader comprises a carboxyl -reactive group and is conjugated to the polypeptide at the C-terminus of the polypeptide. Examples of carboxyl -reactive groups include amines, which can form amide bonds with carboxyl groups; and alcohols, which can form esters with carboxyl groups. In one embodiment a carboxyl group may be modified using a reagent such as a carbodiimide, followed by reaction with a nucleophilic group (e.g. an amine). In one embodiment a compound comprising a nucleophilic group such as an amine group coupled to a click chemistry group such as an azide or alkyne group may be used, with the carboxyl group being activated by the carbodiimide, the amine group reacting with the activated carboxyl group; and followed by reaction of the free click chemistry group.In some embodiments the polypeptide is conjugated to the leader via a side chain group of a residue (e.g. an amino acid residue) in the polypeptide. In some embodiments the polypeptide has a naturally occurring reactive functional group which can be used to facilitate conjugation to the leader. For example, a cysteine residue can be used to form a disulphide bond to the leader or to a modified group thereon.
[0282] In some embodiments the polypeptide is modified in order to facilitate its conjugation to the leader (in addition to being modified to increase the net charge of the peptide as described herein). In some embodiments a residue in the polypeptide is modified to facilitate attachment of the polypeptide to the leader. In some embodiments a residue (e.g. an amino acid residue) in the polypeptide is chemically modified for attachment to the leader. In some embodiments a residue (e.g. an amino acid residue) in the polypeptide is enzymatically modified for attachment to the leader.
[0283] In some embodiments the polypeptide is modified by attaching a moiety comprising a reactive functional group for attaching to the leader. For example, in some embodiments the polypeptide can be extended at the N-terminus or the C-terminus by one or more residues (e.g. amino acid residues) comprising one or more reactive functional groups for reacting with a corresponding reactive functional group on the leader. For example, in some embodiments the polypeptide can be extended at the N-terminus and / or the C-terminus by one or more cysteine residues. Such residues can be used for attachment to the leader, e.g. by maleimide chemistry (e.g. by reaction of cysteine with an azido-maleimide compound such as azido-[Pol]-maleimide wherein [Pol] is typically a short chain polymer such as PEG, e.g. PEG2, PEG3, or PEG4; followed by coupling to appropriately functionalised group on the leader, e.g. a BCN group for reaction with the azide). For avoidance of doubt, when the polypeptide comprises an appropriate naturally occurring residue at the N- and / or C-terminus (e.g. a naturally occurring cysteine residue at the N- and / or C-terminus) then such residue(s) can be used for attachment to the polynucleotide.
[0284] The conjugation chemistry between the leader and the polypeptide in the construct is not particularly limited. Any suitable combination of reactive functional groups can be used. Many suitable reactive groups and their chemical targets are known in the art. Some exemplary reactive groups and their corresponding targets include aryl azides which may react with amine, carbodiimides which may react withamines and carboxyl groups, hydrazides which may react with carbohydrates, hydroxmethyl phosphines which may react with amines, imidoesters which may react with amines, isocyanates which may react with hydroxyl groups, carbonyls which may react with hydrazines, maleimides which may react with sulfhydryl groups, NHS-esters which may react with amines, PFP-esters which may react with amines, psoralens which may react with thymine, pyridyl disulfides which may react with sulfhydryl groups, vinyl sulfones which may react with sulfhydryl amines and hydroxyl groups, vinylsulfonamides, and the like.
[0285] Other suitable chemistry for conjugating the polypeptide to the leader includes click chemistry. Many suitable click chemistry reagents are known in the art. Suitable examples of click chemistry include, but are not limited to, the following:
[0286] (f) copper(I)-catalyzed azide-alkyne cycloadditions (azide alkyne Huisgen cycloadditions);
[0287] (g) strain-promoted azide-alkyne cycloadditions; including alkene and azide [3+2] cycloadditions; alkene and tetrazine inverse-demand Diels- Alder reactions; and alkene and tetrazole photoclick reactions;
[0288] (h) copper-free variant of the 1,3 dipolar cycloaddition reaction, where an azide reacts with an alkyne under strain, for example in a cyclooctane ring such as in bicycle[6.1.0]nonyne (BCN);
[0289] (i) the reaction of an oxygen nucleophile on one linker with an epoxide or aziridine reactive moiety on the other; and
[0290] (j) the Staudinger ligation, where the alkyne moiety can be replaced by an aryl phosphine, resulting in a specific reaction with the azide to give an amide bond. Any reactive group may be used to form the construct. Some suitable reactive groups include [1, 4-Bis[3-(2-pyridyldithio)propionamido]butane; 1,1 1-bis-maleimidotriethyleneglycol; 3,3’-dithiodipropionic acid di(N-hydroxysuccinimide ester); ethylene glycol-bis(succinic acid N-hydroxysuccinimide ester); 4,4’-diisothiocyanatostilbene-2,2’-disulfonic acid disodium salt; Bis[2-(4-azidosalicylamido)ethyl] disulphide; 3-(2-pyridyldithio)propionic acid N-hydroxysuccinimide ester; 4-maleimidobutyric acid N-hydroxysuccinimide ester; lodoacetic acid N-hydroxysuccinimide ester; S-acetylthioglycolic acid N-hydroxysuccinimide ester; azide-PEG-maleimide; and alkyne-PEG-maleimide. Thereactive group may be any of those disclosed in WO 2010 / 086602, particularly in Table 3 of that application.
[0291] In some embodiments a reactive functional group is comprised in the leader and a complementary functional group is comprised in the polypeptide prior to conjugation of the leader and the polypeptide. In other embodiments the reactive functional group is comprised in the polypeptide and the complementary functional group is comprised in the leader prior to conjugation of the leader and the polypeptide. In some embodiments the reactive functional group is attached directly to the polypeptide. In some embodiments the reactive functional group is attached to the polypeptide via a spacer. Any suitable spacer can be used. Suitable spacers include for example alkyl diamines such as ethyl diamine, etc.
[0292] As mentioned above, in some embodiments the leader is conjugated directly to the polypeptide. In some embodiments the leader is ligated to a linker which is conjugated to the polypeptide.
[0293] Generally speaking, in some embodiments a linker may comprise a multifunctional molecule. As used herein, a multi-functional molecule is a molecule having at least two reactive functional groups and is thus capable of binding to a polypeptide and to a leader, e.g. to a polynucleotide. In some embodiments the multi-functional molecule has a first reactive group capable of attaching to the leader and a second reactive group capable of attaching to an amino acid in a polypeptide. The first reactive group may be capable of attaching to any complementary functional group in the leader. For example, the first reactive group may be capable of attaching to a nucleotide in the leader.
[0294] A linker may comprise a linear or branched, unsubstituted or substituted alkylene, alkenylene, alkynylene, arylene, heteroarylene, carbocyclylene or heterocyclylene moiety. A linker may comprise an unsubstituted or substituted alkylene, alkenylene, or alkynylene moiety. A linker may comprise an unsubstituted or substituted alkylene or alkenylene moiety. A linker may comprise an unsubstituted or substituted alkylene moiety. Typically, an alkylene group is a Ci-io alkylene group. Typically, an alkenylene group is a C2-10 alkenylene group. Typically, an alkynylene group is a C2-10 alkynylene group. Typically, an arylene group is a Ce-12 arylene group. Typically, a heteroarylene group is a 5- to 12- membered heteroarylene group.Typically, a carbocyclylene group is a C5-12 carbocyclylene group. Typically, a heterocyclylene group is a 5- to 12- membered heterocyclylene group.
[0295] An alkylene, alkenylene, or alkynylene moiety may be uninterrupted or interrupted by or terminate in one or more atoms or groups selected from maleimide, O, N(R), S, C(O), C(O)NR, C(O)O, phosphate, thiophosphate, unsubstituted or substituted arylene, unsubstituted or substituted heteroarylene, unsubstituted or substituted carbocyclylene and unsubstituted or substituted heterocyclylene; wherein R is selected from H, unsubstituted or substituted alkyl, and unsubstituted or substituted aryl.
[0296] In some embodiments a multifunctional molecule which can be used as a linker in the disclosed methods may comprise a linear or branched, unsubstituted or substituted alkylene, alkenylene, alkynylene, arylene, heteroarylene, carbocyclylene or heterocyclylene moiety comprising at least two reactive groups. In some embodiments a first reactive group may be a group capable of reacting with a complementary functional group of a leader, e.g. to a group comprised in or attached to a polynucleotide. A second reactive group may be a group capable of reacting with a complementary group of the polypeptide. In some embodiments the first and second reactive groups are each independently selected from maleimide groups, carboxyl groups, amine groups, hydroxy groups, azide groups and alkyne groups. Click chemistry reactions between linkers and leaders and polypeptides are particularly useful as it is possible to readily ensure orthogonal reactivity between the reactions of the linker with the polypeptide and the leader. Click chemistry reagents are described herein.
[0297] In some embodiments a linker may be comprised from a multifunctional molecule such as glutaraldehyde, disuccinimidyl suberate (DSS), bismaleimidoethane (BMOE), sulfo-SMCC (sulfosuccinimidyl 4-(N-maleimidomethyl)cyclohexane-l-carboxylate), sulfo-EMCS (N-s-maleimidocaproyl-oxysulfosuccinimide ester), EDC (1-ethyl-3-carbodiimide), DMP (Dess-Martin periodinane; l,l,l-Tris(acetyloxy)-l,l-dihydro-l,2-benziodoxol-3-(lH)-one), bis(sulfosuccinimidyl) suberate (BS3), DTME (dithiobismaleimidoethane), SPDP (succinimidyl 3-(2-pyridyldithio)propionate), BMH (bismaleimidohexane), sulfo-GMBS (N-y-maleimidobutyryl-oxysulfosuccinimide ester), MPBH (4-(4-N-maleimidophenyl)butyric acid hydrazide), NHS-PC (e.g., PCBiotin-PEG-NHS ester, e.g. CAS number 2353409-93-3), and NHS-PEG-maleimide (e.g. CAS number 1325208-25-0).
[0298] In some embodiments, a linker is or comprises a polymer. In some embodiments, a linker is or comprises a biopolymer. Suitable polymers include polynucleotides, polypeptides, polysaccharides, polyethylene glycols) and the like. In some embodiments the or each linker independently comprises a polynucleotide, a polypeptide and / or a polysaccharide. In some embodiments the or each linker independently comprises a polynucleotide, and / or a polypeptide. In some embodiments a linker comprises a synthetic polymer. In some embodiments a linker comprises or consists of a polymer such as polyethylene glycol (PEG), poly(vinyl alcohol) (PVA), polyethylene oxide) (PEO), poly(acrylic acid) (PAA), polypropylene glycol) (PPG), poly(caprolactone) (PCL), polydimethylsiloxane (PDMS), poly(methacrylate) and derivatives thereof, polyurethane, poly(2-oxazoline), poly(N-isopropylacrylamide) (PNIPAM), and / or polyethyleneimine (PEI). In some embodiments a linker comprises a dendrimer.
[0299] When the linker is a polymer, the length of the linker is not particularly limited and can be chosen or selected as required by the operator of the disclosed methods. The choice of linker is an operational parameter within the control of the skilled user.
[0300] In some embodiments, however, the linker is a polymer (e.g. a polynucleotide and / or a polypeptide) comprising from about 2 to about 1000 monomer units, such as from about 5 to about 500 monomer units, e.g. from about 10 to about 100 monomer units, e.g. from about 10 to about 50 monomer units, e.g. about 10, about 15, about 20, about 30, about 40 or about 50 monomer units (e.g. peptide units or nucleotide units).
[0301] Leader
[0302] In some embodiments the leader is comprised in an adapter conjugated to the polypeptide. In some embodiments the leader is comprised in an adapter conjugated to the first end of the polypeptide. In some embodiments the second end of the polypeptide is attached to a further adapter (e.g. a further adapter comprising a blocking moiety as described herein). In some embodiments the second end of the polypeptide is not attached to an adapter.In some embodiments the adapter comprises a recognition sequence for the protein translocase as described herein. In some embodiments the adapter is configured such that the leader is attached to the polypeptide via the recognition sequence for the protein translocase. This is not essential, however, and in some embodiments the adapter is configured such that the leader is not attached to the polypeptide via the recognition sequence for the protein translocase; for example, the recognition sequence for the protein translocase may be pendant to the leader, or a first end of the leader may be attached to the polypeptide and a second end of the leader may be attached to the recognition sequence for the protein translocase. In some embodiments the recognition sequence for the protein translocase is comprised in the leader. For example, in some embodiments the leader comprises a polypeptide portion and the polypeptide portion comprises a recognition sequence for the protein translocase.
[0303] Any suitable leader may be used, as explained herein. The leader is capable of threading through the first opening of the nanopore. In some embodiments the leader is capable of threading through the second opening of the nanopore. In some embodiments the leader is capable of threading through the first opening and the second opening of the nanopore. In some embodiments the detector is or comprises a nanopore and the leader is capable of translocating the nanopore.
[0304] The leader comprises or consists of a non-polypeptide component. The nonpolypeptide component may for example be a polynucleotide or a modified polynucleotide such as a spacer group as described herein. The leader may comprise or consist of a polynucleotide or a modified polynucleotide as described herein.
[0305] In some embodiments the leader is charged. In some embodiments the leader is uncharged. In some embodiments the leader is negatively charged or positively charged, typically negatively charged. A charged leader may be useful e.g. in threading the polypeptide through the first and / or second openings of the detector.
[0306] In some embodiments the leader may comprise or consist of DNA or RNA. In some embodiments the leader may comprise or consist of ssDNA or dsDNA. In some embodiments the leader may comprise or consist of one or more nucleotides in addition to (or instead of) other monomer units; for example the leader may comprise (or consist of) one or more spacers. In some embodiments a spacer may provide an energy barrier which impedes movement of a protein translocase. For example, in some embodimentsa spacer may stall a protein translocase or may physically block movement of a protein translocase, for instance by introducing a bulky chemical group to physically impede the movement of the protein translocase on the polypeptide.
[0307] In some embodiments the leader may comprise or consist of one or more abasic spacers i.e. one or more spacers in which the bases are removed from one or more nucleotides. In some embodiments the leader may comprise peptide nucleic acid (PNA), glycerol nucleic acid (GNA), threose nucleic acid (TNA), locked nucleic acid (LNA) or a synthetic polymer with nucleotide side chains. In some embodiments the leader may comprise one or more nitroindoles, one or more inosines, one or more acridines, one or more 2-aminopurines, one or more 2-6-diaminopurines, one or more 5-bromo-deoxyuri dines, one or more inverted thymidines (inverted dTs), one or more inverted dideoxy -thymidines (ddTs), one or more dideoxy-cytidines (ddCs), one or more 5-methylcytidines, one or more 5 -hydroxymethylcytidines, one or more 2’-O-Methyl RNA bases, one or more Iso-deoxycytidines (Iso-dCs), one or more Isodeoxyguanosines (Iso-dGs), one or more C3 (OC3H6OPO3) groups, one or more photo-cleavable (PC) [OC3H6-C(O)NHCH2-CeH3NO2-CH(CH3)OPO3] groups, one or more hexandiol groups, one or more spacer 9 (iSp9) [(OCEhCEh^OPCh] groups, or one or more spacer 18 (iSp 18) [(OCEhCEh^OPCh] groups; or one or more thiol connections. A leader may comprise any combination of these groups. Many of these groups are commercially available from IDT® (Integrated DNA Technologies®). For example, C3, iSp9 and iSp 18 spacers are all available from IDT®. Other spacers may comprise one or more pendant chemical groups, which may for example be attached to one or more nucleobases in a polynucleotide. Suitable groups include, but are not limited to, fluorophores, streptavidin and / or biotin, cholesterol, methylene blue, dinitrophenols (DNPs), digoxigenin and / or anti-digoxigenin and dibenzylcyclooctyne groups. A leader may comprise any number of the above groups. For example, a leader may comprise from about 5 to about 100, such as from about 10 to about 50, e.g. from about 20 to about 40 spacers as described herein, e.g. C3, iSp9 and / or iSp 18.
[0308] In some embodiments the leader comprises a polymer such as PEG or a polysaccharide. In such embodiments the leader may be from 10 to 150 monomer units (e.g. ethylene glycol or saccharide units) in length, such as from 20 to 120, e.g. 30 to100, for example 40 to 80 such as 50 to 70 monomer units (e.g. ethylene glycol or saccharide units) in length.
[0309] In some embodiments the leader is directly attached to the polypeptide. In some embodiments the leader is directly attached to the first end of the polypeptide in the construct. In some embodiments the leader is attached to the first end of the polypeptide by a chemical bond. In some embodiments the leader is attached to the first end of the polypeptide by a covalent bond. In some embodiments the leader is attached to the first end of the polypeptide by a linker as described herein.
[0310] As explained in more detail, in some embodiments the leader is attached to the first end of the polypeptide by a recognition site for the protein translocase. In some embodiments the leader comprises a recognition site for the protein translocase and is attached to the first end of the polypeptide. In some embodiments the leader is directly attached to the first end of the polypeptide and the leader comprises a recognition site for the protein translocase.
[0311] In some embodiments, the leader does not comprise a blocking moiety. In some embodiments the leader does not comprise a blocking moiety for preventing the translocation of the construct through the first and second openings of the nanopore. In some embodiments the leader does not comprise a blocking moiety as described herein. For example, a blocking moiety attached to a leader may in some embodiments prevent the leader from threading the first and / or second openings of the detector, e.g. may prevent the leader from translocating through a nanopore. In some embodiments the construct may comprise a blocking moiety e.g. at the second end of the polypeptide portion of the construct but in some embodiments the leader attached at the first end of the polypeptide does not comprise a blocking moiety. In some embodiments the construct comprises a polypeptide conjugated at a first end to a leader, and attached at a second end to a blocking moiety.
[0312] In some embodiments the leader of the construct is functionalised. For example, in some embodiments a polynucleotide can be functionalised by hybridising to the polynucleotide a further polynucleotide strand comprising a functionality, such as an anchor for localising the construct to a membrane in which the nanopore may be present, and / or to a tether for localising the construct to the nanopore. Accordingly, insome embodiments the leader is functionalised for localisation at the nanopore and / or at a membrane comprising the nanopore. In some embodiments prior to step (i) the leader is hybridised to a second polynucleotide-containing strand comprising a membrane anchor and / or a nanopore tether. Anchors and tethers are described in more detail herein.
[0313] Generally speaking, the leader may be functionalised by being comprised in an adapter. An adapter may also be referred to as a sequencing adapter. However, the disclosure also embraces methods in which a separate sequencing adapter is attached to a construct comprising the polypeptide and leader.
[0314] A sequencing adapter (such as an adapter comprising a leader as described herein) may in some embodiments be synthetic or artificial. Typically, an adapter comprises a polymer as described herein. In some embodiments, an adapter comprises a spacer as described herein. In some embodiments, an adapter comprises a polynucleotide. A polynucleotide adapter may comprise DNA, RNA, modified DNA (such as abasic DNA), RNA, PNA, LNA, BNA and / or PEG. Usually, the or each adapter comprises single stranded and / or double stranded DNA or RNA.
[0315] Accordingly, in some embodiments of the disclosed methods, the leader is attached to a second polynucleotide-containing strand. In some embodiments this forms a double-stranded portion. In some embodiments the double stranded portion may define a linear adapter or a Y adapter as described herein.
[0316] In some embodiments, an adapter is a linear adapter. A linear adapter may be bound to either or both ends of a oligopeptide.
[0317] A linear adapter may comprise the leader as described herein. A linear adapter may comprise a portion for hybridisation with a tag (such as a pore tag) as described herein. A linear adapter may be 10 to 150 nucleotides in length, such as from 20 to 120, e.g. 30 to 100, for example 40 to 80 such as 50 to 70 nucleotides in length. A linear adapter may be single stranded. A linear adapter may be double stranded.
[0318] In some embodiments, an adapter may be a Y adapter. A Y adapter may comprise the leader as described herein. A Y adapter is typically a polynucleotide adapter. A Y adapter is typically double stranded and comprises (a) at one end, a region where the two strands are hybridised together and (b), at the other end, a region where the two strands are not complementary. The non-complementary parts of the strandstypically form overhangs. The presence of a non-complementary region in the Y adapter gives the adapter its Y shape since the two strands typically do not hybridise to each other unlike the double stranded portion. The two single-stranded portions of the Y adapter may be the same length, or may be different lengths. For example, one single-stranded portion of the Y adapter may be 10 to 150 nucleotides in length, such as from 20 to 120, e.g. 30 to 100, for example 40 to 80 such as 50 to 70 nucleotides in length and the other single stranded portion of the Y adapter may independently by 10 to 150 nucleotides in length, such as from 20 to 120, e.g. 30 to 100, for example 40 to 80 such as 50 to 70 nucleotides in length. The double-stranded “stem” portion of the Y adapter may be e.g. from 10 to 150 nucleotides in length, such as from 20 to 120, e.g.
[0319] 30 to 100, for example 40 to 80 such as 50 to 70 nucleotides in length. A Y adapter may be attached to either or both ends of a barcode as described herein.
[0320] Polynucleotide
[0321] As explained in more detail herein, in some embodiments a leader as described here may comprise or consist of a polynucleotide. In such embodiments, any suitable polynucleotide can be used.
[0322] In some embodiments the polynucleotide is secreted from cells. Alternatively, polynucleotide can be produced inside cells such that it must be extracted from cells for use in the disclosed methods.
[0323] A polynucleotide may be provided as an impure mixture of one or more polynucleotides and one or more impurities. Impurities may comprise truncated forms of polynucleotides which are distinct from the polynucleotide for use in the formation of the conjugate. For example the polynucleotide for use in the formation of the construct may be genomic DNA and impurities may comprise fractions of genomic DNA, plasmids, etc. The target polynucleotide may be a coding region of genomic DNA and undesired polynucleotides may comprise non-coding regions of DNA.
[0324] Examples of polynucleotides include DNA and RNA. The bases in DNA and RNA may be distinguished by their physical size.
[0325] A polynucleotide or nucleic acid may comprise any combination of any nucleotides. The nucleotides can be naturally occurring or artificial. One or more nucleotides in the polynucleotide can be oxidized or methylated. One or morenucleotides in the polynucleotide may be damaged. For instance, the polynucleotide may comprise a pyrimidine dimer. Such dimers are typically associated with damage by ultraviolet light and are the primary cause of skin melanomas.
[0326] One or more nucleotides in the polynucleotide may be modified, for instance with a label or a tag, for which suitable examples are known by a skilled person. The polynucleotide may comprise one or more spacers. An adapter, for example a sequencing adapter, may be comprised in the polynucleotide. Adapters, tags and spacers are described in more detail herein.
[0327] Examples of modified bases are disclosed herein and can be incorporated into the polynucleotide by means known in the art, e.g. by polymerase incorporation of modified nucleotide triphosphates during strand copying (e.g. in PCR) or by polymerase fill-in methods. In some embodiments one or more bases can be modified by chemical means using reagents known in the art.
[0328] A nucleotide typically contains a nucleobase, a sugar and at least one phosphate group. The nucleobase and sugar form a nucleoside. The nucleobase is typically heterocyclic. Nucleobases include, but are not limited to, purines and pyrimidines and more specifically adenine (A), guanine (G), thymine (T), uracil (U) and cytosine (C). The sugar is typically a pentose sugar. Nucleotide sugars include, but are not limited to, ribose and deoxyribose. The sugar is preferably a deoxyribose. The polynucleotide preferably comprises the following nucleosides: deoxyadenosine (dA), deoxyuridine (dU) and / or thymidine (dT), deoxyguanosine (dG) and deoxycytidine (dC). The nucleotide is typically a ribonucleotide or deoxyribonucleotide. The nucleotide typically contains a monophosphate, diphosphate or triphosphate. The nucleotide may comprise more than three phosphates, such as 4 or 5 phosphates. Phosphates may be attached on the 5’ or 3’ side of a nucleotide. The nucleotides in the polynucleotide may be attached to each other in any manner. The nucleotides are typically attached by their sugar and phosphate groups as in nucleic acids. The nucleotides may be connected via their nucleobases as in pyrimidine dimers.
[0329] A polynucleotide may be double stranded or single stranded.
[0330] In some embodiments the polynucleotide is single stranded DNA. In some embodiments the polynucleotide is single stranded RNA. In some embodiments the polynucleotide is a single-stranded DNA-RNA hybrid. DNA-RNA hybrids can beprepared by ligating single stranded DNA to RNA or vice versa. The polynucleotide is most typically single stranded deoxyribonucleic acid (DNA) or single stranded ribonucleic nucleic acid (RNA).
[0331] In some embodiments the polynucleotide is double stranded DNA. In some embodiments the polynucleotide is double stranded RNA. In some embodiments the polynucleotide is a double-stranded DNA-RNA hybrid. Double-stranded DNA-RNA hybrids can be prepared from single-stranded RNA by reverse transcribing the cDNA complement.
[0332] The polynucleotide can be any length. For example, the polynucleotide can be at least 10, at least 50, at least 100, at least 150, at least 200, at least 250, at least 300, at least 400 or at least 500 nucleotides or nucleotide pairs in length. The polynucleotide can be 1000 or more nucleotides or nucleotide pairs, 5000 or more nucleotides or nucleotide pairs in length or 100000 or more nucleotides or nucleotide pairs in length.
[0333] More typically, the polynucleotide has a length of from about 1 to about 10,000 nucleotides or nucleotide pairs, such as from about 1 to about 1000 nucleotides or nucleotide pairs (e.g. from about 10 to about 1000 nucleotides or nucleotide pairs), e.g. from about 5 to about 500 nucleotides or nucleotide pairs, such as from about 10 to about 100 nucleotides or nucleotide pairs, e.g. from about 20 to about 80 nucleotides or nucleotide pairs such as from about 30 to about 50 nucleotides or nucleotide pairs.
[0334] Any number of polynucleotides can be used in the disclosed methods. For instance, the method may comprise using 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 50, 100 or more polynucleotides. If two or more polynucleotides are used, they may be different polynucleotides or two instances of the same polynucleotide. The polynucleotide can be naturally occurring or artificial.
[0335] Nucleotides can have any identity, and include, but are not limited to, adenosine monophosphate (AMP), guanosine monophosphate (GMP), thymidine monophosphate (TMP), uridine monophosphate (UMP), 5-methylcytidine monophosphate, 5-hydroxymethylcytidine monophosphate, cytidine monophosphate (CMP), cyclic adenosine monophosphate (cAMP), cyclic guanosine monophosphate (cGMP), deoxyadenosine monophosphate (dAMP), deoxyguanosine monophosphate (dGMP), deoxythymidine monophosphate (dTMP), deoxyuridine monophosphate (dUMP), deoxycytidine monophosphate (dCMP) and deoxymethylcytidine monophosphate. Thenucleotides are preferably selected from AMP, TMP, GMP, CMP, UMP, dAMP, dTMP, dGMP, dCMP and dUMP. A nucleotide may be abasic (i.e. lack a nucleobase). A nucleotide may also lack a nucleobase and a sugar (i.e. is a C3 spacer).
[0336] The polynucleotide may comprise the products of a PCR reaction, genomic DNA, the products of an endonuclease digestion and / or a DNA library. The polynucleotide may be obtained from or extracted from any organism or microorganism. The polynucleotide may be obtained from a human or animal, e.g. from urine, lymph, saliva, mucus, seminal fluid or amniotic fluid, or from whole blood, plasma or serum. The polynucleotide may be obtained from a plant e.g. a cereal, legume, fruit or vegetable. The polynucleotide may comprise genomic DNA. The genomic DNA may be fragmented. The DNA may be fragmented by any suitable method. For example, methods of fragmenting DNA are known in the art, Such methods may use a transposase, such as a MuA transposase. Often the genomic DNA is not fragmented.
[0337] It is within the scope of the methods provided herein that the polynucleotide is labelled with a molecular label. A molecular label may be a modification to the polynucleotide which promotes the detection of the polynucleotide or construct in the methods provided herein. For example the label may be a modification to the polynucleotide which alters the signal obtained as construct is characterised. For example, the label may interfere with a flux of ions through the nanopore. In such a manner, the label may improve the sensitivity of the methods.
[0338] Protein translocase
[0339] In the disclosed methods, the movement of the construct with respect to the nanopore is controlled using a protein translocase. Thus, the protein translocase is capable of controlling the movement of the construct with respect to the nanopore. Suitable protein translocases are known in the art and some exemplary protein translocases are described in more detail below.
[0340] Protein translocases are protein-binding polypeptides which are able to control movement of a protein substrate, for example an enzyme, enzyme complex, or a part of an enzyme complex that operates on a protein substrate and moves it relative to the enzyme in a processive manner, i.e. as a function of enzymatic activity. Included withinthe term as used herein is a class of enzymes often referred to as ‘unfoldases’, in that they catalyse the unfolding of a native protein without affecting the primary structure, i.e. the primary sequence of the protein.
[0341] As explained here, in some embodiments the protein translocase is present on the construct prior to its contact with the nanopore. Accordingly, in some embodiments the method comprises, prior to contacting the construct with the nanopore, the step of loading the protein translocase onto the construct. In some embodiments the construct is contacted with the nanopore prior to loading of the protein translocase onto the construct. Accordingly, in some embodiments the method comprises loading the protein translocase onto the construct after contacting the construct with the nanopore.
[0342] In some embodiments the protein translocase is capable of remaining bound to the construct when the portion of the construct in contact with the active site (i.e. the polypeptide binding site) of the protein translocase comprises a polypeptide or a polynucleotide. In other words, in some embodiments the protein translocase does not dissociate from the construct. In some embodiments the protein translocase is designed, configured or selected to prevent it from disengaging from the polypeptide or a construct comprising the polypeptide (other than by passing off the end of the polypeptide or construct). Such protein translocases are particularly suitable for use in the disclosed methods. Thus, in some embodiments of such methods, the construct does not disengage from the protein translocase.
[0343] As used herein, the term “disengaging” refers to the dissociation of the protein translocase from the construct, e.g. into the reaction medium. It is important to distinguish potential “disengagement” of a protein translocase from “unbinding” of the protein translocase from the construct, e.g. from a polynucleotide portion of the construct. As used herein, “unbinding” refers to the transient release of the construct (e.g. a polynucleotide portion thereof) from the active site of the protein translocase (e.g. the polypeptide-binding active site of the protein translocase) but does not imply disengagement. Thus, for example, a protein translocase may unbind from the construct without disengaging from the construct, e.g. when it contacts the leader of the construct (e.g. a polynucleotide portion of the construct). When unbound, the protein translocase remains engaged with the construct. When the protein translocase is unbound from the construct it may be able to move on (e.g., along) the construct under an applied forceand may be capable of re-binding to the construct (e.g. at a polypeptide portion of the construct). When engaged on the construct but unbound from the construct (e.g. unbound from a polynucleotide portion of the construct), the protein translocase does not dissociate from the target polynucleotide.
[0344] In some embodiments the protein translocase is an unfoldase. In some embodiments the protein translocase is an NTP driven unfoldase (where NTP refers to a nucleoside 5 ’-triphosphate, for example ATP or GTP). NTP driven unfoldases are NTP-dependent enzymes that catalyze protein unfolding. NTP driven unfoldases include ATP-dependent proteases, such as proteasomal ATPases, AAA proteases, AAA+ enzymes; membrane fusion proteins, such as NSF (N-Ethylmal eimide-sensitive fusion protein) / Sacl8p (N-Ethylmaleimide-sensitive fusion protein homologue in yeast) or p97 / VCP / Cdc48p (97-kDa valosin-containing protein); Pexlp and Pex6p (peroxisomal ATPase); Katanin and SKD1 (Vps4p homolog in mouse) / Vps4p (Vacuolar protein sorting 4 homolog in yeast); Dynein (motor protein); DNA replication proteins, such as ORC (origin recognition complex), Cdc6 (cell division control protein 6), MCM (mini chromosome maintenance protein), DnaA, or RFC (replication factor C) / clamp-loader; RuvB (holliday junction ATP-dependent DNA helicase RuvB, EC=3.6.4.12); TIP49a / TIP49 and TIP49b / TIP48 (eukaryotic RuvB-like protein).
[0345] In some embodiments the protein translocase is an AAA+ enzyme, AAA+ enzymes are members of the AAA+ superfamily of enzymes. AAA+ is an abbreviation for ATPases Associated with diverse cellular Activities. They share a common conserved module of approximately 230 amino acid residues. This is a large, functionally diverse protein family belonging to the AAA+ superfamily of ring-shaped P-loop NTPases, which exert their activity through the energy-dependent remodeling or translocation of macromolecules. Examples include ClpAP, ClpXP, ClpCP, HslYU and Lon in bacteria and their homologues in mitochondria and chloroplasts. With the exception of Lon, AAA+ enzymes (sometimes referred to as unfoldases or proteases) consist of regulatory (ATPase) and proteolytic subunits, while Lon is a single polypeptide containing both regulatory and proteolytic domains. ClpX and ClpA dock with ClpP to form ClpXP and ClpAP proteases, whereas HslU docks with HslY to form another protease, HslVU. ClpA and ClpX form hexamers, in contrast to ClpP whichforms heptamers. HslU and HslY each form hexamers, although HslU heptamers have also been reported. The regulatory subunits ClpA, ClpX and HslU function as chaperones.
[0346] Examples of AAA+ enzymes include ATP-dependent Clp protease ATP -binding subunit ClpX (ClpX), ATP-dependent Clp protease ATP -binding subunit ClpA (ClpA), proteasome-activating nucleotidase (PAN), LON protease, VCP-like ATPase (VAT), AMA, 854, membrane-bound AAA (MBA), small archaeal ubiquitin-like modifier protein (SAMP), ATP-dependent Clp protease ATP -binding subunit ClpC (ClpC), ATP-dependent Clp protease ATP -binding subunit ClpE (ClpE), ATP-dependent protease ATPase subunit HslU (HslU), Caseinolytic mitochondrial matrix peptidase chaperone subunit Y (ClpY), LonA, LonB, ATP-dependent zinc metalloprotease FtsH (FtsH), Proteasome-associated ATPase (Mpa), Cell division cycle protein 48 (Cdc48, also called p97 and VCP) and Cdc48-like protein of actinobacteria (Cpa), Outer mitochondrial transmembrane helix translocase (Mspl), and Protein translocase subunit SecA (SecA),
[0347] AAA+ enzymes may also be referred to as AAA+ molecular motors.
[0348] HslU is a member of the HsplOO and Clp family of ATPase. It can also form complex with HslY to act as an unfoldase.
[0349] Lon proteases are ATP-dependent serine peptidases belonging to the MEROPS peptidase family S16 (Ion protease family, clan SF).
[0350] In some embodiments the protein translocase is ClpX or is a derivative thereof. Further details on ClpX may be found in Maillard et al. 2011 (Cell. 2011 Apr 29;145(3):459-469). ClpX is a member of the HSP (heat-shock protein) 100 family having the Uniprot designation clpX and having the 424 amino acid sequence given there, processed into mature form, as a subunit. ClpX subunits associate to form a sixmembered (homohexameric) ring that is stabilized by binding of ATP or nonhydrolysable analogs of ATP. The N-terminal domain of ClpX is a C4-type zinc binding domain (ZBD) involved in substrate recognition. ZBD forms a very stable dimer that is essential for promoting the degradation of some typical ClpXP substrates such as and MuA.
[0351] In some embodiments the protein translocase is E. coli ClpX. E. coli ClpX generates sufficient mechanical force (>20 pN) to denature stable protein folds, andtranslocates along proteins at a suitable rate for primary sequence analysis by nanopore sensors (up to 80 amino acids per second). ClpX is part of the ClpXP proteasome-like complex. ClpP is composed of a diheptameric cylinder-like protease that binds at one or both ends a regulatory hexameric ATP-dependent unfoldase / translocase complex (e.g. ClpX). ClpX acts as a gate that allows for tagged proteins to enter into the inner lumen of the ClpP protease complex for subsequent degradation. The ATP-dependent unfoldase / translocase activity of the hexameric protein complex, ClpX, is employed to unfold and thread proteins through a nanopore.
[0352] In some embodiments the protein translocase is a ClpX-deltaN subunit, lacking N-terminal amino acids 1-60, linked with a 20 amino acid long linker and prepared as a single polypeptide chain.
[0353] In some embodiments the protein translocase is a Clp / HsplOO ATPase.
[0354] Clp / HsplOO ATPases are responsible for selecting protein targets. For example, the two different bacterial ATPases ClpX and ClpA impart distinct substrate preferences to the ClpP peptidase.
[0355] In some embodiments the protein translocase is a mitochondrial protein translocase. Examples include TOM or TIM from human or eukaryotic cells, such as TOMM20 (translocase of outer mitochondrial membrane homolog), TOMM22 (mitochondrial import receptor subunit 22 homolog), TOMM40 (translocase of outer mitochondrial membrane 40 homolog), T0M7 (translocase of mitochondrial outer membrane 7), T0MM7 (translocase of outer mitochondrial membrane 7 homolog), TIMM8A (translocase of inner mitochondrial membrane 8 homolog A), TIMM50 (translocase of inner mitochondrial membrane 50 homolog).
[0356] Another alternative protein translocase may be prepared from the Sec family of translocases. These include SecB (chaperone protein), SecA (ATPase), SecY (internal membrane complex in prokaryotes), SecE (interal membrane complex in prokaryotes), SecG (internal membrane complex in prokaryotes) or Sec61 (internal membrane complex in eukaryotes), SecD (membrane protein), and SecF (membrane protein).
[0357] Another alternative protein translocase is Type III Secretion System (TTS) Translocase, such as HrcN and any of the subunits of the TTS translocases, or Secindependent periplasmic protein translocase TatC.Examples of suitable protein translocases, such as NTP driven unfoldases as described above, are described in WO 2013 / 123379, hereby incorporated by reference.
[0358] A protein translocase typically requires fuel in order to handle the processing of polynucleotides and / or polypeptides. Fuel is typically free nucleotides or free nucleotide analogues. The free nucleotides may be one or more of, but are not limited to, adenosine monophosphate (AMP), adenosine diphosphate (ADP), adenosine triphosphate (ATP), guanosine monophosphate (GMP), guanosine diphosphate (GDP), guanosine triphosphate (GTP), thymidine monophosphate (TMP), thymidine diphosphate (TDP), thymidine triphosphate (TTP), uridine monophosphate (UMP), uridine diphosphate (UDP), uridine triphosphate (UTP), cytidine monophosphate (CMP), cytidine diphosphate (CDP), cytidine triphosphate (CTP), cyclic adenosine monophosphate (cAMP), cyclic guanosine monophosphate (cGMP), deoxyadenosine monophosphate (dAMP), deoxyadenosine diphosphate (dADP), deoxyadenosine triphosphate (dATP), deoxyguanosine monophosphate (dGMP), deoxyguanosine diphosphate (dGDP), deoxyguanosine triphosphate (dGTP), deoxythymidine monophosphate (dTMP), deoxythymidine diphosphate (dTDP), deoxythymidine triphosphate (dTTP), deoxyuridine monophosphate (dUMP), deoxyuridine diphosphate (dUDP), deoxyuridine triphosphate (dUTP), deoxycytidine monophosphate (dCMP), deoxycytidine diphosphate (dCDP) and deoxycytidine triphosphate (dCTP). The free nucleotides are usually selected from AMP, TMP, GMP, CMP, UMP, dAMP, dTMP, dGMP or dCMP. The free nucleotides are typically adenosine triphosphate (ATP).
[0359] A cofactor for the protein translocase is a factor that allows the protein translocase to function. The cofactor is often a divalent metal cation. The divalent metal cation is often Mg2+, Mn2+, Ca2+or Co2+. The cofactor is most typically Mg2+.
[0360] In some embodiments the protein translocase is loaded onto the construct at a loading site for the protein translocase. A loading site may in some embodiments comprise a peptide recognition sequence for the protein translocase. In other words, the construct may comprise a peptide sequence selectively targeted by the protein translocase such that the protein translocase can engage with the construct by binding to the loading site. Known protein translocases such as those described herein typically have known properties in respect of loading sites. By way of non-limiting example,ClpX is known to target complementary ssrA sequence tags. For instance, the sequence of the E. coli ssrA tag is AANDENYALAA (SEQ ID NO: 1). Accordingly, in some embodiments the construct comprises an ssrA tag for binding ClpX.
[0361] The loading site (recognition sequence for the protein translocase) may be comprised in an adapter. For example, the loading site (recognition sequence for the protein translocase) may be attached to a leader as described herein. In some embodiments the adapter comprises the leader and the recognition sequence for the protein translocase. In some embodiments the polypeptide may be attached to the leader via the loading site (recognition sequence for the protein translocase). In some embodiments the loading site (recognition sequence for the protein translocase) may be comprised in the leader. Thus, in some embodiments the first end of the polypeptide is attached to the leader via the recognition sequence for the protein translocase. In some embodiments the disclosed methods comprise loading the protein translocase onto the recognition sequence prior to contacting the construct with the nanopore. In some embodiments the disclosed methods comprise, prior to conjugating the target polypeptide with the adapter, the step of loading the protein translocase onto the adapter. In some embodiments the methods comprises the step of loading the protein translocase onto the adapter after conjugating the target polypeptide with the adapter.
[0362] Detector
[0363] Embodiments described herein refer to movement of constructs and polypeptides comprised in such constructs as described herein with respect to a nanopore. However, whilst the disclosure provides nanopores as exemplary detectors, the methods provided herein are also amenable to other detectors including (i) a zero-mode waveguide, (ii) a field-effect transistor, optionally a nanowire field-effect transistor; (iii) an AFM tip; (iv) a nanotube, optionally a carbon nanotube and (v) a nanopore. The disclosed methods are particularly amenable to methods in which a polypeptide is moved through a detector or through a structure containing a detector, e.g. a well in a detector chip.
[0364] Nanopore
[0365] In the disclosed methods, any suitable nanopore can be used. In one embodiment a nanopore is a transmembrane pore.A transmembrane pore is a structure that crosses the membrane to some degree. It permits hydrated ions driven by an applied potential to flow across or within the membrane. The transmembrane pore typically crosses the entire membrane so that hydrated ions may flow from one side of the membrane to the other side of the membrane. However, the transmembrane pore does not have to cross the membrane. It may be closed at one end. For instance, the pore may be a well, gap, channel, trench or slit in the membrane along which or into which hydrated ions may flow.
[0366] Any suitable transmembrane pore may be used in the methods provided herein. The pore may be biological or artificial. Suitable pores include, but are not limited to, protein pores, polynucleotide pores, and solid state pores.
[0367] A solid state pore may, in one embodiment, comprise a nanochannel. In some embodiments the solid state pore is a pore disclosed in WO 2003 / 003446, WO 2009 / 020682 or WO 2016 / 187519, each of which is incorporated by reference in their entirety.
[0368] In one embodiment, the pore may be a DNA origami pore (Langecker et al.. Science, 2012; 338: 932-936). Suitable DNA origami pores are disclosed in WO2013 / 083983, WO 2018 / 011603 and WO 2020 / 025974, each of which is incorporated by reference in their entirety.
[0369] In one embodiment, the nanopore is a scaffolded polypeptide nanopore. In some embodiments the pore is a scaffolded polypeptide nanopore as disclosed in WO 2020 / 025909 or WO 2020 / 074399, each of which is incorporated by reference in their entirety.
[0370] In one embodiment, the nanopore is a transmembrane protein pore. A transmembrane protein pore is a polypeptide or a collection of polypeptides that permits hydrated ions, such as polynucleotides, to flow from one side of a membrane to the other side of the membrane. In the methods provided herein, the transmembrane protein pore is capable of forming a pore that permits hydrated ions driven by an applied potential to flow from one side of the membrane to the other. The transmembrane protein pore typically permits polynucleotides and polypeptides to flow from one side of the membrane, such as a polymer membrane, to the other. The transmembrane protein pore allows a polynucleotide or polypeptide to be moved through the pore.In one embodiment, the nanopore is a transmembrane protein pore which is a monomer or an oligomer. The pore is typically made up of several repeating subunits, such as at least 6, at least 7, at least 8, at least 9, at least 10, at least 11, at least 12, at least 13, at least 14, at least 15, or at least 16 subunits. The pore is typically a hexameric, heptameric, octameric or nonameric pore. The pore may be a homooligomer or a hetero-oligomer.
[0371] In one embodiment, the transmembrane protein pore comprises a barrel or channel through which the ions may flow. The subunits of the pore typically surround a central axis and contribute strands to a transmembrane P-barrel or channel or a transmembrane a -helix bundle or channel.
[0372] Typically, the barrel or channel of the transmembrane protein pore comprises amino acids that facilitate interaction with an analyte, such as a target polypeptide (as described herein). These amino acids are typically located near a constriction of the barrel or channel. The transmembrane protein pore typically comprises one or more positively charged amino acids, such as arginine, lysine or histidine, or aromatic amino acids, such as tyrosine or tryptophan. These amino acids typically facilitate the interaction between the pore and nucleotides, polynucleotides, nucleic acids and polypeptides.
[0373] The transmembrane protein pore may be from or derived from Wza, Iota toxin, Anthrax protective antigen, Vibrio cholerae cytolysin, Cytotoxin K (CytK), CELIII, CsgG, CsgF, CsgG-CsgF, Aerolysin, alpha hemolysin, MspA, MspB, MspC, PorARr, PorBRr, PorARc, PilQ, necrotic enteritis B-like toxin (NetB), FraC, portal proteins including G20c, P23_45, T4, SPP1, P22 and Phi29, gamma hemolysin, Monalysin, Lysenin, ClyA, an actinoporin, Clostridium perfringens beta toxin, parasporin-2, epsilon toxin, lectin from the parasitic mushroom Laetiporus sulphureus (LSL), volvatoxin, Cry toxins, Cytl Aa, Cyt2Aa, Complement component 9 (C9), Perfringolysin O, Pleurotolysin, Listeriolysin, Perforin-2, Gasdermin-A3, L-, P- and M-ring protein, Type II secretion system protein D, GspD, InvG, VirB7, SpoIIIAG, Cag8, Cag3, Cag or other proteins in the Type IV secretion system apparatus protein CagY, WzzB, Pentraxin, Afp2, Major vault protein, Thioredoxin-dependent peroxidase reductase, Arf-GAP, Respiratory syncytial virus ribonucleoprotein, Chikungunya virus nonstructural protein 1, PRC, YaxA, XaxA, HfaB, NfpAB, leukocidin and PrgH.Suitable transmembrane protein pores for use in the invention include those described in WO 2016 / 034591, WO 2017 / 149316, WO 2017 / 149317, WO 2017 / 149318, WO 2018 / 211241, WO 2019 / 002893, WO 2023 / 118404, WO 2023 / 198911, WO 2024 / 033421, WO 2024 / 033422, WO 2024 / 033443, and WO 2024 / 089270 (all incorporated by reference herein in their entirety).
[0374] The transmembrane protein pore may also be any of the CsgG pores described in WO 2023 / 060420, WO 2023 / 60418, WO 2023 / 60422, WO 2023 / 060421, WO 2023 / 019470, CN114957412, WO 2023 / 019471, W02023 / 060419 and WO 2023 / 050031 (all incorporated herein by reference in their entireties) or a variant thereof.
[0375] The transmembrane protein pore may be any of the pores described in WO 2023 / 123370, WO 2024 / 138470, WO 2024 / 138472, WO 2024 / 138424, WO 2024 / 138425, WO 2024 / 138512 and WO 2024 / 138565 (all incorporated herein by reference in their entireties) or a variant thereof.
[0376] The transmembrane pore may be formed from a chimeric pore monomer comprising two or more regions, wherein at least two of the two or more regions are from at least two different pores. The chimeric pore monomer may comprise any number of regions, such as three or more, four or more, five or more, six or more, seven or more, eight or more, nine or more or ten or more regions, from different pores. The chimeric pore monomer may comprise two or three regions. The regions are preferably selected from a cap region, a constriction region, and a transmembrane region. The regions may be a cap region and a constriction region. The regions may be a cap region, a constriction region, and a transmembrane region. The at least two different pores are typically at least two different pores that appear in nature. The at least two different pores are typically at least two different wild-type or naturally occurring pores. The at least two different pores are preferably different before any artificial or synthetic modifications, such as additions, deletions and / or substitutions, are made to them. The at least two different pores are preferably homologues, for example structural homologues. A structural homologue refers to a protein or molecule that shares a similar three-dimensional structure with another protein or molecule. This can be determined using standard methods in the art (e.g., AlphaFold or PSIPRED). Structural homologues typically have similar sequences. Structural homologues are normallyidentified in similar species. The at least two different pores may be selected from any of the pores listed above. The at least two different pores may be two different PorARc pores or three different PorARc pores. The at least two different pores may be two different CsgG pores or three different CsgG pores. The chimeric pore monomer may be any of those described in WO 2024 / 089270 (incorporated by reference herein in its entirety).
[0377] The transmembrane pore may be formed from a pore monomer comprising (a) a CsgG monomer and (b) a fusion polypeptide comprising a first portion comprising a CsgF peptide and a second portion comprising a helix-forming auxiliary protein, wherein the fusion protein is attached to the pore monomer. The pore monomer may be derived from a protein transmembrane pore complex comprising (a) a CsgG transmembrane pore comprising a lumen and (b) a fusion polypeptide comprising a first portion comprising a CsgF protein and a second portion comprising a helix-forming auxiliary protein, wherein the fusion protein is attached to the transmembrane pore. The auxiliary protein can be designed de novo using computer-based structural analysis tools to confer certain desirable features to the CsgG monomer (e.g., modulation of pore width, lengthening of pore lumen, formation of one or more additional constrictions, etc.). The de novo designed auxiliary protein may form one or more additional constrictions in the lumen of a CsgG pore formed from the monomer, and improve discrimination of polymer units as an analyte moves through the pore. The pore monomer may be any of the pore monomers described in WO 2024 / 033447 (incorporated by reference herein in its entirety).
[0378] In some embodiments of the disclosed methods an electroosmotic force is present across the nanopore. In some embodiments the methods comprise applying an electroosmotic force across the nanopore. In some embodiments an electroosmotic force across the nanopore facilitates translocation of the construct or part thereof across the nanopore. For example, in some embodiments an electroosmotic force across the nanopore maintains the polypeptide in a linearised (e.g. unfolded) state in which translocation of the polypeptide through the nanopore under the control of the protein translocase is facilitated.In some embodiments an electroosmotic force across the nanopore can be generated or enhanced by using asymmetric salt conditions and / or providing suitable charge in the environment of the polypeptide.
[0379] In some embodiments the nanopore is chosen or modified to increase the electroosmotic force across the nanopore. In some embodiments the nanopore is configured to generate an electroosmotic force across the nanopore. In some embodiments the nanopore is configured to modify (e.g. to increase or enhance) the electroosmotic force across the nanopore. For example, an electroosmotic force across the nanopore can be increased by increasing the charge in the channel of the nanopore. The charge in the channel of a protein nanopore can be altered e.g. by mutagenesis. The charge in the channel of a solid state nanopore can be altered e.g. by chemical modification of the substrate in which the solid state nanopore is produced. Altering the charge of a nanopore is well within the capacity of those skilled in the art. In some embodiments altering the charge of a nanopore generates electroosmotic forces from the unbalanced flow of cations and anions through the nanopore when an electrical potential is applied across the nanopore. Accordingly, in some embodiments the nanopore is modified to increase the electroosmotic force across the nanopore.
[0380] In some embodiments the electroosmotic force acts in the direction from the cis opening to the trans opening of the nanopore. The electroosmotic force arises from electroosmotic flow through the nanopore and thus in some embodiments the electroosmotic force arises from an electroosmotic flow through the nanopore in the direction from the cis opening to the trans opening of the nanopore.
[0381] In some embodiments the electroosmotic force acts in the direction from the trans opening to the cis opening of the nanopore. In some embodiments the electroosmotic force arises from an electroosmotic flow through the nanopore in the direction from the trans opening to the cis opening of the nanopore.
[0382] In some embodiments the electroosmotic force is in the same direction as an electrical potential applied across the nanopore. In other words, in some embodiments the electroosmotic force is in the same direction as an electrophoretic force across the nanopore. In some embodiments the electroosmotic force arises from an electroosmotic flow through the nanopore in the same direction as the electrophoretic force across the nanopore. In some embodiments the electroosmotic force arises from an electroosmoticflow through the nanopore in the same direction as the electrophoretic force arising from an applied electrical (voltage) potential applied across the nanopore. In some embodiments the electroosmotic force acts in the direction from the cis opening to the trans opening of the nanopore and an electrical (voltage) potential is applied across the nanopore such that an electrophoretic force acts on the target polypeptide in the direction from the cis opening to the trans opening of the nanopore. In some embodiments the electroosmotic force acts in the direction from the trans opening to the cis opening of the nanopore and an electrical (voltage) potential is applied across the nanopore such that an electrophoretic force acts on the target polypeptide in the direction from the trans opening to the cis opening of the nanopore. Thus, in some embodiments the electroosmotic force augments the electrophoretic force across the nanopore provided by an electrical potential applied across the nanopore. In some embodiments the or each force is the force acting on the polypeptide or construct during its movement with respect to the nanopore.
[0383] In some embodiments the electroosmotic force is in the opposite direction as an electrical potential applied across the nanopore. In other words, in some embodiments the electroosmotic force is in the opposite direction to an electrophoretic force across the nanopore. In some embodiments the electroosmotic force arises from an electroosmotic flow through the nanopore in the opposite direction to the electrophoretic force across the nanopore. In some embodiments the electroosmotic force arises from an electroosmotic flow through the nanopore in the opposite direction to the electrophoretic force arising from an applied electrical (voltage) potential applied across the nanopore. In some embodiments the electroosmotic force acts in the direction from the cis opening to the trans opening of the nanopore and an electrical (voltage) potential is applied across the nanopore such that an electrophoretic force acts on the target polypeptide in the direction from the trans opening to the cis opening of the nanopore. In some embodiments the electroosmotic force acts in the direction from the trans opening to the cis opening of the nanopore and an electrical (voltage) potential is applied across the nanopore such that an electrophoretic force acts on the target polypeptide in the direction from the cis opening to the trans opening of the nanopore. Thus, in some embodiments the electroosmotic force reduces the force across the nanopore provided by an electrical potential applied across the nanopore. In someembodiments the or each force is the force acting on the polypeptide or construct during its movement with respect to the nanopore.
[0384] Tags
[0385] In some embodiments of the methods provided herein, a tag on the nanopore can be used, e.g. to promote the capture of the polypeptide or a construct comprising the polypeptide.
[0386] The interaction between a tag on a nanopore and a binding site on the polypeptide or a construct comprising the polypeptide may be reversible. A strong non-covalent bond (e.g., biotin / avidin) is still reversible and can be useful in some embodiments of the methods described herein. For example, a pair of pore tag and binding site (e.g. comprised in a polypeptide, construct or adaptor) can be designed to provide a sufficient interaction with the nanopore such that the polypeptide or construct is held close to the nanopore (without detaching from the nanopore and diffusing away) but is able to release from the nanopore as it is processed.
[0387] A pore tag and adaptor can be configured such that the binding strength or affinity of a binding site on a construct comprising the polypeptide (e.g., a binding site provided by an anchor or a leader sequence of an adaptor or by a capture sequence within the duplex stem of an adaptor) to a tag on a nanopore is sufficient to maintain the coupling between the nanopore and construct until an applied force is placed on it to release the bound construct from the nanopore.
[0388] One or more molecules that attract or bind the polypeptide or construct, or an adapter attached thereto, may be linked to the nanopore. Any molecule that hybridizes to the polypeptide, construct and / or adaptor may be used. The molecule attached to the pore may be selected from a PNA tag, a PEG linker, a short oligonucleotide, a positively charged amino acid and an aptamer. Pores having such molecules linked to them are known in the art. For example, pores having short oligonucleotides attached thereto are disclosed in Howarka et al (2001) Nature Biotech. 19: 636-639 and WO 2010 / 086620, and pores comprising PEG attached within the lumen of the pore are disclosed in Howarka et al (2000) J. Am. Chem. Soc. 122(11): 2411-2416. In some embodiments, a tag or tether may be uncharged. This can ensure that the tags or tethers are not drawn into the nanopore under the influence of a potential difference if present.A short oligonucleotide attached to the nanopore, which comprises a sequence complementary to a sequence in a construct comprising the polypeptide (e.g. in a leader sequence or another single stranded sequence in an adaptor) may be used to enhance capture of the polypeptide, construct or adapter attached thereto in the methods described herein.
[0389] Anchors
[0390] In some embodiments of the methods provided herein, an anchor on the polypeptide or a construct comprising the polypeptide can be used, e.g. to promote the localisation of the polypeptide or construct to a membrane in which a nanopore may be present.
[0391] Thus, in one embodiment, a polypeptide, construct or adapter attached thereto may comprise a membrane anchor or a transmembrane pore anchor. The anchor may be a polypeptide anchor and / or a hydrophobic anchor that can be inserted into the membrane. In one embodiment, the hydrophobic anchor is a lipid, fatty acid, sterol, carbon nanotube, polypeptide, protein or amino acid, for example cholesterol, palmitate or tocopherol. The anchor may comprise thiol, biotin or a surfactant. In one aspect the anchor may be biotin (for binding to streptavidin), amylose (for binding to maltose binding protein or a fusion protein), Ni-NTA (for binding to poly-histidine or polyhistidine tagged proteins) or peptides (such as an antigen).
[0392] In one embodiment, the anchor is or comprises cholesterol or a fatty acyl chain. For example, any fatty acyl chain having a length of from 6 to 30 carbon atom, such as hexadecanoic acid, may be used. Examples of suitable anchors and methods of attaching anchors to adapters are disclosed in WO 2012 / 164270 and WO 2015 / 150786, hereby incorporated by reference in their entirety.
[0393] Blocking moiety
[0394] As explained herein, the protein translocase may be prevented from disengaging from the construct by using a blocking moiety. The blocking moiety may for example by attached to or comprised in the second end of the polypeptide of the construct.
[0395] Thus, in some embodiments the second end of the polypeptide comprises or is attachedto a blocking moiety to prevent the protein translocase from disengaging from the construct.
[0396] In some embodiments the disclosed methods comprise attaching the blocking moiety to the second end of the polypeptide prior to attaching the first end of the polypeptide to the leader. Thus, in some embodiments, the method comprises attaching the blocking moiety to the second end of the target polypeptide prior to conjugating the adapter to the first end of the target polypeptide.
[0397] In some embodiments the disclosed methods comprise attaching the first end of the polypeptide to the leader prior to attaching the second end of the polypeptide to the leader. Thus, in some embodiments the method comprises attaching the blocking moiety to the second end of the target polypeptide after conjugating the adapter to the first end of the target polypeptide.
[0398] The blocking moiety is typically too large to pass through the protein translocase (e.g. through the polypeptide-binding site of a protein translocase) and so when the movement of the construct with respect to the nanopore brings the blocking moiety into contact with the protein translocase, the movement of the construct through the protein translocase and therefore typically also through the nanopore is prevented.
[0399] Accordingly, in some embodiments, the blocking moiety limits the movement of the construct through the polypeptide binding site of the protein translocase and thereby limits the movement of the construct in the direction from the first opening of the nanopore towards the second opening of the nanopore. In some embodiments the blocking moiety limits the movement of the polypeptide with respect to the polypeptide binding site of the protein translocase and thereby limits the movement of the polypeptide with respect to the nanopore in the direction from the first opening to the second opening.
[0400] Any suitable blocking moiety can be used in the provided methods. For example, the construct may be modified with biotin and the blocking moiety may be e.g. traptavidin, streptavidin, avidin or neutravidin. The blocking moiety may be a large chemical group such as a dendrimer. The blocking moiety may be a nanoparticle or a bead. A blocking moiety may comprise one or more of:
[0401] a polynucleotide secondary structure, preferably a hairpin or G-quadruplex (TBA);a nucleic acid analog, preferably selected from peptide nucleic acid (PNA), glycerol nucleic acid (GNA), threose nucleic acid (TNA), locked nucleic acid (LNA), bridged nucleic acid (BNA) and abasic nucleotides; and fluorophores, avidins such as traptavidin, streptavidin and neutravidin, and / or biotin, cholesterol, methylene blue, dinitrophenols (DNPs), digoxigenin and / or anti-digoxigenin and dibenzylcyclooctyne groups.
[0402] The blocking moiety may be attached to the construct in any suitable manner. The blocking moiety may be attached directly to the polypeptide. The blocking moiety may be attached to the polypeptide via a linker. Any suitable linker may be used. Any suitable chemistry for attaching the blocking moiety to the polypeptide may be used. Any of the attachment methods described herein for attaching the leader to the polypeptide may be used to attach the blocking moiety to the polypeptide. For avoidance of doubt, in embodiments where the construct comprises both a leader attached to the polypeptide as described herein and also a blocking moiety attached to the polypeptide of the construct as described herein, the attachment means may be the same or different. In some embodiments a linker is used to attach the leader to the polypeptide and the blocking moiety is attached directly to the polypeptide. In some embodiments the leader is attached directly to the polypeptide and a linker is used to attach the blocking moiety to the polypeptide. In some embodiments the leader is attached directly to the polypeptide and the blocking moiety is attached directly to the polypeptide. In some embodiments a linker is used to attach the leader to the polypeptide and a linker is used to attach the blocking moiety to the polypeptide. When a linker is used to attach both the blocking moiety and the leader to the construct the linkers may be the same or different.
[0403] Membrane
[0404] The nanopore is typically present in a membrane. Any suitable membrane may be used.
[0405] The membrane is preferably an amphiphilic layer. An amphiphilic layer is a layer formed from amphiphilic molecules, such as phospholipids, which have both hydrophilic and lipophilic properties. The amphiphilic molecules may be synthetic or naturally occurring. Non-naturally occurring amphiphiles and amphiphiles which form amonolayer are known in the art and include, for example, block copolymers (Gonzalez-Perez et al., Langmuir, 2009, 25, 10447-10450). Block copolymers are polymeric materials in which two or more monomer sub-units that are polymerized together to create a single polymer chain. Block copolymers typically have properties that are contributed by each monomer sub-unit. However, a block copolymer may have unique properties that polymers formed from the individual sub-units do not possess. Block copolymers can be engineered such that one of the monomer sub-units is hydrophobic (i.e., lipophilic), whilst the other sub-unit(s) are hydrophilic whilst in aqueous media. In this case, the block copolymer may possess amphiphilic properties and may form a structure that mimics a biological membrane. The block copolymer may be a diblock (consisting of two monomer sub-units) but may also be constructed from more than two monomer sub-units to form more complex arrangements that behave as amphiphiles. The copolymer may be a triblock, tetrablock or pentablock copolymer.
[0406] The membrane may comprise just one type of amphiphile or may comprise more than one type of amphiphile. The membrane may comprise a mixture of naturally occurring amphiphilic molecules such as a mixture of phospholipids. The membrane may comprise a mixture of non-naturally occurring amphiphiles such as a mixture of block copolymers. The membrane may comprise a mixture of naturally occurring amphiphilic molecules such as a phospholipids and non-naturally occurring amphiphiles such as block copolymers. For example, in some embodiments the membrane comprises from about 1% to about 99% (e.g. wt% or mol%) of naturally occurring amphiphilic molecules such as one or more phospholipids and from about 1% to about 99% (e.g. wt% or mol%) of non-naturally occurring amphiphiles such as one or more block copolymers (wherein the sum of naturally occurring amphiphilic molecules such as a phospholipids and non-naturally occurring amphiphiles such as block copolymers does not exceed 100 % (e.g. wt% or mol%)).
[0407] The membrane may be one of the membranes disclosed in WO 2014 / 064443 or WO 2014 / 064444 (both of which are incorporated herein by reference in their entireties).
[0408] The amphiphilic molecules may be chemically modified or functionalised to facilitate coupling of the polynucleotide. The amphiphilic layer may be a monolayer ora bilayer. The amphiphilic layer is typically planar. The amphiphilic layer may be curved. The amphiphilic layer may be supported.
[0409] Amphiphilic membranes are typically naturally mobile, essentially acting as two-dimensional fluids with lipid diffusion rates of approximately 10'8cm s'1. This means that the pore and coupled polynucleotide can typically move within an amphiphilic membrane.
[0410] The membrane may be a lipid bilayer. Lipid bilayers are models of cell membranes and serve as excellent platforms for a range of experimental studies. For example, lipid bilayers can be used for in vitro investigation of membrane proteins by single-channel recording. Alternatively, lipid bilayers can be used as biosensors to detect the presence of a range of substances. The lipid bilayer may be any lipid bilayer. Suitable lipid bilayers include, but are not limited to, a planar lipid bilayer, a supported bilayer, or a liposome. The lipid bilayer is preferably a planar lipid bilayer. Suitable lipid bilayers are disclosed in WO 2008 / 102121, WO 2009 / 077734, and WO 2006 / 100484 (incorporated herein by reference in their entireties).
[0411] Methods for forming lipid bilayers are known in the art. Lipid bilayers are commonly formed by the method of Montal and Mueller (Proc. Natl. Acad. Sci. USA., 1972; 69: 3561-3566).
[0412] A lipid bilayer may be formed as described in WO 2009 / 077734 (incorporated herein by reference in its entirety). In this method, the lipid bilayer is formed from dried lipids. A lipid bilayer may be formed across an opening as described in WO 2009 / 077734.
[0413] The membrane may comprise a solid-state layer. Solid state layers can be formed from both organic and inorganic materials including, but not limited to, microelectronic materials, insulating materials such as SisN4, AI2O3, and SiO, organic and inorganic polymers such as polyamide, plastics such as Teflon® or elastomers such as two-component addition-cure silicone rubber, and glasses. The solid-state layer may be formed from graphene. Suitable graphene layers are disclosed in WO 2009 / 035647 (incorporated herein by reference in its entirety). If the membrane comprises a solid-state layer, the pore is typically present in an amphiphilic membrane or layer contained within the solid-state layer, for instance within a hole, well, gap, channel, trench or slit within the solid-state layer. The skilled person can prepare suitable solidstate / amphiphilic hybrid systems. Suitable systems are disclosed in WO 2009 / 020682 and WO 2012 / 005857 (incorporated herein by reference in their entireties). Any of the amphiphilic membranes or layers discussed above may be used.
[0414] The methods disclosed herein are typically carried out using (i) an artificial amphiphilic layer comprising a pore, (ii) an isolated, naturally occurring lipid bilayer comprising a pore, or (iii) a cell having a pore inserted therein. The methods are typically carried out using an artificial amphiphilic layer, such as an artificial block copolymer layer. The layer may comprise other transmembrane and / or intramembrane proteins as well as other molecules in addition to the pore. Suitable apparatus and conditions are discussed below. The method of the invention is typically carried out in vitro.
[0415] Characterisation
[0416] As discussed in more detail herein, some embodiments of the disclosed methods comprise taking one or more measurements characteristic of the polypeptide of the construct as the protein translocase controls the movement of the construct with respect to the nanopore, thereby characterising the target polypeptide.
[0417] Many different characteristics can be determined and any suitable measurements can be taken.
[0418] For example, in some embodiments the characterising the polypeptide comprises determining (i) the length of the polypeptide, (ii) the identity of the polypeptide, (iii) the sequence of the polypeptide, (iv) the secondary structure of the polypeptide; and / or (v) whether or not and / or to the extent to which the polypeptide is modified. In typical embodiments the measurements are characteristic of the sequence of the polypeptide or whether or not the polypeptide is modified, e.g. by one or more post-translational modifications. In some embodiments the measurements are characteristics of the sequence of the polypeptide.
[0419] Those skilled in the art will appreciate that characterising a polypeptide does not necessarily comprise determining any or all of these features. Many characterisation measurements can be made as a polypeptide moves with respect to a nanopore.
[0420] In some embodiments the measurements are characteristic of the sequence of the target polypeptide.Conditions
[0421] The disclosed methods may be carried out using any apparatus that is suitable for investigating a membrane / pore system in which a nanopore is inserted into a membrane. The methods may be carried out using any apparatus that is suitable for transmembrane pore sensing. For example, the apparatus may comprise a chamber comprising an aqueous solution and a barrier that separates the chamber into two sections. The barrier may have an aperture in which a membrane containing a transmembrane pore is formed. Transmembrane pores are described herein.
[0422] The methods may be carried out using the apparatus described in WO 2008 / 102120, WO 2010 / 122293 or WO 00 / 28312.
[0423] The methods may comprise optical measurements, for example such as described in WO 2016 / 009180 and WO 2021 / 198695.
[0424] The methods may involve measuring the ion current flow through the pore, typically by measurement of a current. Alternatively, the ion flow through the pore may be measured optically, such as disclosed by Heron et al: J. Am. Chem. Soc. 9 Vol. 131, No. 5, 2009. Therefore the apparatus may also comprise an electrical circuit capable of applying a potential and measuring an electrical signal across the membrane and pore. The characterisation methods may be carried out using a patch clamp or a voltage clamp. The characterisation methods typically involve the use of a voltage clamp.
[0425] The methods may be carried out on a silicon-based array of wells where each array comprises 128, 256, 512, 1024, 2000, 3000, 4000, 6000, 10000, 12000, 15000 or more wells.
[0426] The methods may involve the measuring of a current flowing through the pore. The method is typically carried out with a voltage applied across the membrane and pore. The voltage used is typically from +2 V to -2 V, typically -400 mV to +400mV. The voltage used is typically in a range having a lower limit selected from -400 mV, -300 mV, -200 mV, -150 mV, -100 mV, -50 mV, -20m V and 0 mV and an upper limit independently selected from +10 mV, + 20 mV, +50 mV, +100 mV, +150 mV, +200 mV, +300 mV and +400 mV. The voltage used is more typically in the range 100 mV to 240mV and most typically in the range of 120 mV to 220 mV. It is possible toincrease discrimination between different nucleotides by a pore by using an increased applied potential.
[0427] The methods are typically carried out in the presence of any charge carriers, such as metal salts, for example alkali metal salts, halide salts, for example chloride salts, such as alkali metal chloride salt. Charge carriers may include ionic liquids or organic salts, for example tetramethyl ammonium chloride, trimethylphenyl ammonium chloride, phenyltrimethyl ammonium chloride, or l-ethyl-3 -methyl imidazolium chloride. In the exemplary apparatus discussed above, the salt is present in the aqueous solution in the chamber. Potassium chloride (KC1), sodium chloride (NaCl) or caesium chloride (CsCl) is typically used. KC1 is typical. The salt may be an alkaline earth metal salt such as calcium chloride (CaCh). The salt concentration may be at saturation. The salt concentration may be 3M or lower and is typically from 0.1 to 2.5 M, from 0.3 to 1.9 M, from 0.5 to 1.8 M, from 0.7 to 1.7 M, from 0.9 to 1.6 M or from 1 M to 1.4 M. The salt concentration is typically from 150 mM to 1 M. The characterisation method may be carried out using a salt concentration of at least 0.3 M, such as at least 0.4 M, at least 0.5 M, at least 0.6 M, at least 0.8 M, at least 1.0 M, at least 1.5 M, at least 2.0 M, at least 2.5 M or at least 3.0 M. High salt concentrations provide a high signal to noise ratio and allow for currents indicative of binding / no binding to be identified against the background of normal current fluctuations.
[0428] The methods are typically carried out in the presence of a buffer. In the exemplary apparatus discussed above, the buffer is present in the aqueous solution in the chamber. Any suitable buffer may be used. Typically, the buffer is HEPES.
[0429] Another suitable buffer is Tris-HCl buffer. The methods are typically carried out at a pH of from 4.0 to 12.0, from 4.5 to 10.0, from 5.0 to 9.0, from 5.5 to 8.8, from 6.0 to 8.7 or from 7.0 to 8.8 or 7.5 to 8.5. The pH used may be about 7.5.
[0430] The methods may be carried out at from 0 °C to 100 °C, from 15 °C to 95 °C, from 16 °C to 90 °C, from 17 °C to 85 °C, from 18 °C to 80 °C, 19 °C to 70 °C, or from 20 °C to 60 °C. The methods are typically carried out at room temperature. The methods are optionally carried out at a temperature that supports enzyme function, such as about 37 °C.
[0431] Further aspectsAlso provided is a construct as described herein. In some embodiments provided herein is a construct comprising a polypeptide conjugated to an adapter comprising (a) a leader comprising a non-polypeptide component and optionally (b) a recognition sequence for a protein translocase, wherein the construct further comprises said protein translocase; wherein
[0432] the polypeptide comprises a first end and a second end;
[0433] the first end of the polypeptide is attached to the leader;
[0434] the second end of the polypeptide is attached to a blocking moiety capable of limiting the movement of the polypeptide with respect to the polypeptide binding site of the protein translocase; and
[0435] the protein translocase is oriented on the construct in an orientation for processing the polypeptide in a direction from the first end to the second end.
[0436] In some embodiments the polypeptide, leader, recognition sequence (if present), blocking moiety, and protein translocase are as described herein.
[0437] Also provided is a kit comprising:
[0438] a first adapter having a first end comprising a leader comprising a non- polypeptide component; an optional recognition sequence for a protein translocase; and a second end comprising an attachment point for attaching to a first end of a polypeptide analyte; and
[0439] a second adapter having a first end comprising an attachment point for attaching to a second end of the polypeptide analyte; and a second end comprising a blocking moiety;
[0440] wherein the first adapter comprises a protein translocase in an orientation for processing said adapter in a direction from the first end to the second end;
[0441] and wherein the blocking moiety of the second adapter is suitable for preventing the protein translocase from disengaging from the second end of the polypeptide analyte when the first and second adapters are attached to the polypeptide analyte.
[0442] In some embodiments, the leader, polypeptide, recognition sequence (if present), attachment points (and chemistry comprised therein), protein translocase and blocking moiety are as described herein.Also provided is a system for characterising a target polypeptide comprising: first and second adapters as defined in the kit;
[0443] a nanopore for characterising the target polypeptide as the target polypeptide moves with respect to the nanopore; and
[0444] a protein translocase for controlling the movement of the target polypeptide with respect to the nanopore. Typically, the nanopore is as described herein.
[0445] The system and kit disclosed herein may be configured for use with an algorithm, also provided herein, adapted to be run on a computer system. The algorithm may be adapted to detect information characteristic of a polypeptide (e.g. characteristic of the sequence of the polypeptide and / or whether the polypeptide is modified), and to selectively process the signal obtained as a construct comprising the polypeptide conjugated to a leader comprising a non-polypeptide component (e.g. a polynucleotide component) moves with respect to the nanopore. Also provided is a system comprising computing means configured to detect information characteristic of a polypeptide (e.g. characteristic of the sequence of the polypeptide and / or whether the polypeptide is modified) and to selectively process the signal obtained as a construct comprising the polypeptide conjugated to a leader comprising a non-polypeptide component (e.g. a polynucleotide component) moves with respect to the nanopore. In some embodiments the system comprises receiving means for receiving data from detection of the polypeptide, processing means for processing the signal obtained as the construct moves with respect to the nanopore, and output means for outputting the characterisation information thus obtained.
[0446] It is to be understood that although particular embodiments, specific configurations as well as materials and / or molecules, have been discussed herein for methods according to the present invention, various changes or modifications in form and detail may be made without departing from the scope and spirit of this invention. The preceding embodiments and subsequent examples are provided for illustration only, and should not be considered limiting the application. The application is limited only by the claims.EXAMPLES
[0447] Preparation of analytes for electrophysiology
[0448] Polypeptide analytes referred to as KSI (SEQ ID NO: 2) and DHFR (SEQ ID NO: 3) were prepared. Each analyte was labelled at the ybbR tag to introduce a linking moiety to facilitate leader attachment as described herein; labelling was carried out using the method of Yin et al., Nat. Protoc. 1, 280-285 (2006). A peptide corresponding to SEQ ID 4 was attached to analyte KSI, and an oligonucleotide corresponding to SEQ ID 5 was attached to analyte DHFR.
[0449] Leaders were prepared by annealing oligonucleotides corresponding to SEQ ID 6 and SEQ ID 7. The leader was attached to a peptide containing an unfoldase recognition sequence (SEQ ID 8) using click chemistry, and subsequently purified by SPRI using AMPure XP beads.
[0450] The leader with attached unfoldase recognition sequence was then modified by DNA ligation using T4 DNA ligase to attach an oligonucleotide corresponding to SEQ ID 9. This was then attached to the linker-modified KSI using click chemistry. The conjugate was subsequently purified by SPRI using AMPure XP beads. The leader with attached unfoldase recognition sequence was attached to DHFR using click chemistry and subsequently purified by SPRI using AMPure XP beads.
[0451] Modification of the charge of analyte DHFR
[0452] DHFR analyte protein to be modified was diluted into a buffer containing 6M Guanidine hydrochloride and DMSO. The sample was reduced through addition of TCEP and incubated for 15 mins at room temperature. The sample was then subjected to heating at 95°C for 30 mins. The protein was then exchanged into pH 9.0 borate buffer and labelled with 2500 equivalents of 4-sulfophenyl isothiocyanate for 1 hour at 37°C. Labelling in this manner modifies the sidechain of each lysine residue, removing the positive charge and introducing a negative charge. The protein was then exchanged into 25 mM HEPES-KOH pH 7.5, 50 mM NaCl.
[0453] Electrophysiology
[0454] Electrical data were collected using FLO-MIN004RA flow cells (Oxford Nanopore Technologies pic). Sequencing libraries consisting of 200 nM analyte and 40nM E. coli ClpX unfoldase enzyme were prepared in 75 pL Sequencing Buffer (Oxford Nanopore Technologies pic). Flow cells were first flushed with 1 mL of Sequencing Buffer. Sample (75 pL) was then introduced into the flow cell via the SpotON port. Electrical data were acquired with a sample rate of 4 kHz at 30°C and applied potential of 140 mV.
[0455] Results and discussion
[0456] KSI and DHFR polypeptides with attached oligonucleotide leader and unfoldase recognition sequence are depicted schematically in Figure 1.
[0457] Electrophysiology measurements were first collected with KSI; a depiction of the analyte delivery is shown in Figure 2. An example of an obtained trace is shown in Figure 5, demonstrating successful controlled translocation of the analyte through the nanopore.
[0458] Data collection was then attempted with unmodified DHFR. No signals could be collected, suggesting that the unmodified DHFR analyte could not be translocated through the nanopore.
[0459] A charge-modified version of DHFR was subsequently tested. In this version, the DHFR polypeptide had its negative charge increased - after ybbR tagging and prior to attachment of the recognition sequence-appended leader - by modification with 4-sulfophenyl isothiocyanate, as described above. Electrophysiology measurements were conducted and signals for translocation were collected. An example of an obtained trace is shown in Figure 6, demonstrating that modification of the analyte to increase its net negative charge led to successful controlled translocation through the nanopore.
[0460] The charge distribution for unmodified and modified DHFR was calculated. Shown in Figures 7A and 7B are plots representing charge distribution for unmodified and modified DHFR domains. A subsequence from SEQ ID 3 (85-242 inclusive) was taken, corresponding to the folded DHFR domain. A moving average of charge was calculated and plotted taking into account a 20 residue window for both unmodified and modified sequences. Y axis on the plot is the moving average charge, and the X axis is the position in the sequence. A horizontal dotted line is shown indicating neutral net charge. The process of charge modification increased the net negative charge, andreduced the positive charge; this is shown as a reduction in the total area under the curve between 1 and 0.
[0461] SEQUENCE LISTING
[0462]
Claims
CLAIMS1. A method of moving a target polypeptide with respect to a nanopore using a protein translocase;the nanopore having a first opening and a second opening;the method comprising:(i) modifying the target polypeptide to increase the net charge of the polypeptide and (ii) conjugating the target polypeptide with a leader comprising a non-polypeptide component, thereby forming a construct; wherein step (i) can be conducted prior to, simultaneously with or subsequent to step (ii);contacting the construct with the first opening of the nanopore under conditions such that the leader threads through the nanopore in the direction from the first opening to the second opening and the construct moves with respect to the nanopore in the direction from the first opening to the second opening; andcontrolling the movement of the construct in the direction from the first opening to the second opening of the nanopore using the protein translocase.
2. A method according to claim 1, wherein modifying the target polypeptide comprises modifying the side chain(s) of one or more amino acids in the target polypeptide.
3. A method according to claim 2, wherein modifying the target polypeptide comprises covalently attaching said one or more charge-modifying moieties to said side chain(s).
4. A method according to any one of the preceding claims, wherein modifying the target polypeptide comprises contacting the target polypeptide with one or more aminoacid modifying enzymes and / or with one or more chemical reagents.
5. A method according to any one of the preceding claims, wherein step (i) comprises modifying the target polypeptide to increase the net negative charge of the polypeptide.
6. A method according to any one of the preceding claims, wherein the protein translocase controls the movement of a portion of the target polypeptide with respect to the nanopore.
7. A method according to claim 6, wherein the portion of the target polypeptide has a length of at least 5, at least 10, at least 20, at least 30, at least 40, at least 50, or at least 100 amino acids.
8. A method according to any one of the preceding claims, wherein the target polypeptide has a length of at least 5, at least 10, at least 20, at least 30, at least 40, at least 50, at least 100, at least 200, at least 300, at least 400 or at least 500 amino acids.
9. A method according to any one of the preceding claims, wherein the target polypeptide has a first end and a second end; the first end of the target polypeptide is attached to the leader; and the protein translocase is oriented on the construct in an orientation to process the target polypeptide in a direction from the first end towards the second end.
10. A method according to any one of the preceding claims, wherein the leader is comprised in an adapter comprising a recognition sequence for the protein translocase.
11. A method according to claim 10, wherein the method comprises loading the protein translocase onto the recognition sequence prior to contacting the construct with the nanopore.
12. A method according to claim 10 or 11, wherein the method comprises, prior to conjugating the target polypeptide with the adapter, the step of loading the protein translocase onto the adapter.
13. A method according to claim 10 or 11, wherein the method comprises the step of loading the protein translocase onto the adapter after conjugating the target polypeptide with the adapter.
14. A method according to any one of the preceding claims, wherein the target polypeptide has a first end and a second end; the first end of the target polypeptide is conjugated to the leader; and the second end of the target polypeptide comprises or is attached to a blocking moiety to prevent the protein translocase from disengaging from the construct and / or to prevent the second end of the target polypeptide from translocating through the nanopore.
15. A method according to claim 14, wherein the blocking moiety limits the movement of the target polypeptide with respect to the polypeptide binding site of the protein translocase and thereby limits the movement of the target polypeptide with respect to the nanopore in the direction from the first opening to the second opening.
16. A method according to claim 14 or 15, wherein the method comprises attaching the blocking moiety to the second end of the target polypeptide prior to conjugating the leader to the first end of the target polypeptide.
17. A method according to claim 14 or 15, wherein the method comprises attaching the blocking moiety to the second end of the target polypeptide after conjugating the leader to the first end of the target polypeptide.
18. A method according to any one of the preceding claims, wherein the construct comprises the target polypeptide attached to the leader via a recognition sequence for the protein translocase.
19. A method according to any one of the preceding claims, wherein the leader comprises a polynucleotide.
20. A method according to any one of the preceding claims, wherein the leader is functionalised for localisation at the nanopore and / or at a membrane comprising the nanopore.
21. A method according to any one of the preceding claims, wherein the construct moves in the direction from the first opening of the nanopore to the second opening of the nanopore under an applied force.
22. A method according to any one of the preceding claims, wherein the construct moves in the direction from the first opening of the nanopore to the second opening of the nanopore with the applied force and wherein the protein translocase controls the movement of the construct in the direction from the first opening of the nanopore to the second opening of the nanopore with the applied force.
23. A method according to claim 21 or 22, wherein the applied force is an electrical or chemical potential applied across the nanopore;preferably wherein the applied force is a voltage potential.
24. A method according to any one of the preceding claims, comprising applying an electroosmotic force across the nanopore.
25. A method according to any one of the preceding claims, wherein the nanopore is configured to generate an electroosmotic force across the nanopore.
26. A method according to claim 24 or 25, wherein the construct moves with the electroosmotic force with respect to the nanopore.
27. A method according to any one of claims 24 to 26, wherein the electroosmotic force arises from an electroosmotic flow through the nanopore in the same direction as an electrophoretic force applied across the nanopore.
28. A method according to any one of the preceding claims, wherein the nanopore is a transmembrane nanopore spanning a membrane having a cis side and a trans side, and:(i) the first opening of the nanopore is at the cis side of the membrane and the second opening of the nanopore is at the trans side; and the protein translocase controls the movement of the construct through the nanopore from the cis side to the trans side of the membrane; or(ii) the first opening of the nanopore is at the trans side of the membrane and the second opening of the nanopore is at the cis side; and the protein translocase controls the movement of the construct through the nanopore from the trans side to the cis side of the membrane.
29. A method according to any one of the preceding claims, wherein the nanopore is comprised in a membrane separating a first volume from a second volume, wherein the first volume contacts the first opening of the nanopore and the second volume contacts the second opening of the nanopore; and wherein the protein translocase is retained in the first volume during said method.
30. A method according to any one of the preceding claims, wherein the protein translocase is an NTP-driven unfoldase.
31. A method of characterising a target polypeptide,the target polypeptide being comprised in a construct comprising the target polypeptide conjugated to a leader comprising a non-polypeptide component , wherein the construct further comprises said protein translocase;the method comprisingmoving the target polypeptide with respect to a nanopore as defined in any one of the preceding claims; andtaking one or more measurements characteristic of the target polypeptide as the protein translocase controls the movement of the construct with respect to the nanopore,thereby characterising the target polypeptide.
32. A method according to claim 31, wherein the one or more measurements are characteristic of one or more characteristics of the target polypeptide selected from (i) the length of the polypeptide, (ii) the identity of the polypeptide, (iii) the sequence of the polypeptide, (iv) the secondary structure of the polypeptide and (v) whether or not the polypeptide is modified.
33. A method of controlling the movement of a target polypeptide;the target polypeptide being comprised in a construct comprising the target polypeptide conjugated to a leader comprising a non-polypeptide component, wherein the construct further comprises said protein translocase;the method comprising moving the target polypeptide with respect to a nanopore as defined in any one of the preceding claims; and controlling the speed of the movement of the construct with respect to the nanopore using the protein translocase;thereby controlling the movement of the target polypeptide with respect to the nanopore.
34. A construct comprising a polypeptide conjugated to an adapter comprising (a) a leader comprising a non-polypeptide component and optionally (b) a recognition sequence for a protein translocase, wherein the construct further comprises said protein translocase; whereinthe polypeptide comprises a first end and a second end;the first end of the polypeptide is attached to the leader;the second end of the polypeptide is attached to a blocking moiety capable of limiting the movement of the polypeptide with respect to the polypeptide binding site of the protein translocase; andthe protein translocase is oriented on the construct in an orientation for processing the polypeptide in a direction from the first end to the second end.
35. A kit, comprisinga first adapter having a first end comprising a leader comprising a non- polypeptide component; an optional recognition sequence for a proteintranslocase; and a second end comprising an attachment point for attaching to a first end of a polypeptide analyte; anda second adapter having a first end comprising an attachment point for attaching to a second end of the polypeptide analyte; and a second end comprising a blocking moiety;wherein the first adapter comprises a protein translocase in an orientation for processing said adapter in a direction from the first end to the second end;and wherein the blocking moiety of the second adapter is suitable for preventing the protein translocase from disengaging from the second end of the polypeptide analyte when the first and second adapters are attached to the polypeptide analyte.
36. A system for characterising a target polypeptide comprising:first and second adapters as defined in claim 35;a nanopore for characterising the target polypeptide as the target polypeptide moves with respect to the nanopore; anda protein translocase for controlling the movement of the target polypeptide with respect to the nanopore.