method
By destabilizing polypeptide structure through charge disruption and using a motor protein, the method enhances nanopore-based polypeptide characterization, addressing sequence fidelity loss and fragmentation challenges.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- OXFORD NANOPORE TECH LTD
- Filing Date
- 2026-01-19
- Publication Date
- 2026-07-23
AI Technical Summary
Existing methods for characterizing polypeptides using nanopores are hindered by secondary and tertiary structure, leading to sequence fidelity loss and fragmentation issues, especially for long polypeptides.
Destabilize the secondary and tertiary structure of polypeptides by disrupting the charge of amino acids using charge-modifying moieties, followed by using a motor protein to control the movement through a nanopore for characterization.
Facilitates the movement and characterization of polypeptides with improved sequence fidelity, allowing for efficient determination of properties such as identity, length, and modification status.
Smart Images

Figure IMGF000001_0001 
Figure IMGF000001_0002 
Figure IMGF000007_0001
Abstract
Description
[0001] METHOD
[0002] Related
[0003]
[0004] The present application claims priority to United Kingdom Patent Application No.
[0005] 2500667.7 filed 17 January 2025, United Kingdom Patent Application No. 2513808.2 filed 22 August 2025, and United Kingdom Patent Application No. 2518554.7 filed 6 November 2025, the entire contents of each of which are hereby incorporated by reference in their entirety.
[0006] Field
[0007] The present disclosure relates to methods of moving a target polypeptide with respect to a nanopore, by destabilizing the secondary and / or tertiary structure of the target polypeptide. The present disclosure also relates to methods of characterising a target polypeptide using a nanopore. Also provided are related constructs, kits, and uses of certain reagents to facilitate translocation of a target polypeptide through a nanopore.
[0008]
[0009] The characterisation of biological molecules is of increasing importance in biomedical and biotechnological applications. For example, sequencing of nucleic acids allows the study of genomes and the proteins they encode and, for example, allows correlation between nucleic acid mutations and observable phenomena such as disease indications. Nucleic acid sequencing can be used in evolutionary biology to study the relationship between organisms. Metagenomics involves identifying organisms present in samples, for example microbes in a microbiome, with nucleic acid sequencing allowing the identification of such organisms.
[0010] Whilst techniques to characterise (e.g. sequence) polynucleotides have been extensively developed, techniques to characterise polypeptides are less advanced, despite being of very significant biotechnological importance. For example, knowledge of a protein sequence can allow structure-activity relationships to be established and has implications in rational drug development strategies for developing ligands for specific receptors. Identification of post-translational modifications is also key to understanding the functional properties of many proteins. For example, typically 30-50% of protein species are phosphorylated in eukaryotes. Some proteins may have multiple phosphorylation sites, serving to activate or inactivate a protein, promote its degradation, or modulate interactions with protein partners.Known methods of characterising polypeptides include mass spectrometry and Edman degradation.
[0011] Protein mass spectrometry involves characterising whole proteins or fragments thereof in an ionised form. Known methods of protein mass spectrometry include electrospray ionisation (ESI) and matrix-assisted laser desorption / ionisation (MALDI). Mass spectrometry has some benefits, but results obtained can be affected by the presence of contaminants and it can be difficult to process fragile molecules without their fragmentation. Moreover, mass spectrometry is not a single molecule technique and provides only bulk information about the sample interrogated. Mass spectrometry is unsuitable for characterising differences within a population of polypeptide samples and is unwieldy when seeking to distinguish neighbouring residues.
[0012] Edman degradation is an alternative to mass spectrometry which allows the residue-by-residue sequencing of polypeptides. Edman degradation sequences polypeptides by sequentially cleaving the N-terminal amino acid and then characterising the individually cleaved residues using chromatography or electrophoresis. However, Edman sequencing is slow, involves the use of costly reagents, and like mass spectrometry is not a single molecule technique.
[0013] As such, there remains a pressing need for new techniques to characterise polypeptides, especially at the single molecule level. Single molecule techniques for characterising biomolecules such as polynucleotides have proven to be particularly attractive due to their high fidelity and avoidance of amplification bias.
[0014] One attractive method of single molecule characterization of biomolecules such as polypeptides is nanopore sensing. Nanopore sensing is an approach to analyte detection and characterization that relies on the observation of individual binding or interaction events between the analyte molecules and an ion conducting channel. Nanopore sensors can be created by placing a single pore of nanometre dimensions in an electrically insulating membrane and measuring voltage-driven ion currents through the pore in the presence of analyte molecules. The presence of an analyte inside or near the nanopore will alter the ionic flow through the pore, resulting in altered ionic or electric currents being measured over the channel. The identity of an analyte is revealed through its distinctive current signature, notably the duration and extent of current blocks and the variance ofcurrent levels during its interaction time with the pore. Nanopore sensing has the potential to allow rapid and cheap polypeptide characterisation.
[0015] Nanopore sensing and characterisation of polypeptides has been proposed in the art. For example, WO 2013 / 123379 discloses the use of an NTP-driven protein processing unfoldase enzyme to process a protein to be translocated through a nanopore. WO 2021 / 111125 discloses methods in which a target polypeptide may be characterised as it moves through a nanopore using a polynucleotide-handling protein. WO 2021 / 133168 discloses protein and polypeptide fingerprinting and sequencing by nanopore translocation. WO 2024 / 094986 discloses methods of characterising a target polypeptide as it moves with relation to a nanopore. Each of these documents is incorporated by reference in their entireties.
[0016] These methods have provided useful techniques for characterising polypeptides using nanopores. However, some challenges remain. In particular, long polypeptides can have significant secondary and tertiary structure which can hamper characterisation methods. Random shearing of long polypeptides (e.g. by exposing the peptides to chemical reagents, e.g. by exposing the polypeptide to a change in pH, or to a chemical reagent) can generate random shorter oligopeptides which avoid or reduce such structures, and some known methods involve characterising such oligopeptides including using nanopores. However, shearing a longer polypeptide into shorter fragments can result in a loss of sequence fidelity as the fragments can become scrambled in order. In other words, whilst such methods can allow the characteristics of each individual fragment to be determined, information about the overall characteristics of the longer polypeptide from which the fragments are derived may be lost or degraded. For example, if a desired characteristic is the sequence of the longer polypeptide, then random fragmentation of the polypeptide and scrambling before it is characterised typically hinders obtaining this information, because the order of the fragments itself is typically scrambled such that the overall sequence of the longer polypeptide is no longer present. Whilst there are approaches to address this, such approaches may be costly, e.g. in terms of computer processing requirements, and / or require redundancy in the data acquisition to try to obtain overlapping fragments from which the original polypeptide can be reconstructed.
[0017] Accordingly, there remains a need for further methods.
[0018] The inventors have recognised that significant technical advantages would arise from facilitating the movement of polypeptides with respect to nanopores. The inventors have appreciated that secondary and tertiary structure in long polypeptides may contributeto challenges in characterising such polypeptides using nanopores. Accordingly, the inventors have developed methods which allow polypeptides to be moved with respect to nanopores by disrupting such secondary and tertiary structure. The provided methods facilitate the movement of the polypeptides with respect to nanopores, and thus can be beneficially used in characterising polypeptides in nanopore methods.
[0019] Accordingly, an aspect of the disclosure relates to methods of characterising a target polypeptide. The method comprises destabilizing the secondary and / or tertiary structure of the polypeptide which in turn facilitates the movement of the target polypeptide with respect to a nanopore. The inventors have appreciated that secondary and / or tertiary structure in polypeptides is typically stabilized by charge-mediated interactions between amino acid residues in the polypeptide, and therefore in the provided methods the secondary and / or tertiary structure is destabilized by disrupting the charge of one or more amino acids in the target polypeptide. The inventors have further appreciated that this can be effectively achieved by modifying the side chains of the amino acids in the polypeptide. Once the polypeptide is destabilized, then a motor protein may in some aspects be used to control the movement of the polypeptide with respect to a nanopore. In some aspects, the target polypeptide can be characterised by taking one or more measurements characteristic of the target polypeptide during the movement. The methods may allow properties of the target polypeptide to be determined, such as its identity, length, composition, amino acid sequence and / or whether and how the polypeptide is modified.
[0020] Accordingly, provided herein is a method of moving a target polypeptide with respect to a nanopore;
[0021] the method comprising destabilizing the secondary and / or tertiary structure of the target polypeptide by disrupting the charge of one or more amino acids in the target polypeptide; wherein disrupting the charge of one or more amino acids in the target polypeptide comprises modifying the side chains of said one or more amino acids with one or more charge-modifying moieties; and
[0022] contacting the destabilized polypeptide with a motor protein under conditions such that the motor protein controls the movement of the polypeptide with respect to the nanopore.
[0023] In some embodiments, disrupting the charge of one or more amino acids in the target polypeptide comprises modifying the side chains of a plurality of amino acids in thetarget polypeptide. In some embodiments, the amino acids in said plurality of amino acids may be the same or different.
[0024] In some embodiments, the one or more amino acids that are modified with one or more charge-modifying moieties comprise one or more charged amino acids. In some embodiments, the one or more amino acids that are modified with one or more chargemodifying moieties comprise one or more positively charged amino acids and disrupting the charge of said one or more amino acids comprises reducing the positive charge of said amino acids. In some embodiments, the one or more amino acids that are modified with one or more charge-modifying moieties comprise one or more negatively charged amino acids and disrupting the charge of said one or more amino acids comprises reducing the negative charge of said amino acids. In some embodiments, the one or more amino acids that are modified with one or more charge-modifying moieties are selected from lysine, arginine, glutamate and aspartate.
[0025] In some embodiments, modifying the side chains of said one or more amino acids with one or more charge-modifying moieties comprises covalently attaching said one or more charge-modifying moieties to said side chains.
[0026] In some embodiments, modifying the side chains of said one or more amino acids with one or more charge-modifying moieties comprises one or more of oxidation, esterification, O-glycosylation, alkylation, nitration, nitrosylation, succinylation, hydroxylation, amidation, deamidation, acylation, sulfhydration, carbamylation, glycation, metal coordination, Schiff-base formation, and ADP-ribosylation of said side chains.
[0027] In some embodiments, modifying the side chains of said one or more amino acids with one or more charge-modifying moieties comprises contacting said one or more amino acids with one or more amino-acid modifying enzymes. In some embodiments, said one or more amino-acid modifying enzymes are selected from kinases, sulfotransferases, O-GlcNAc transferases (OGTs), methyltransferases, peroxidases, palmitoyltransferases (PATs), nitrosylases, acetyltransferases, hydroxylases, deiminases, succinyltransferases, glutamylases, oligosaccharyltransferases (OSTs), glycosyltransferases, and glutaminase.
[0028] In some embodiments, modifying the side chains of said one or more amino acids with one or more charge-modifying moieties comprises contacting said one or more amino acids with one or more chemical reagents. In some embodiments, the one or more chemical reagents are selected from oxidising agents, halogenating agents, nitrating agents, alkylating agents, phosphoric acid derivatives, glycosyl donors, acids, bases, NO donors,sulphide donors, glycation agents, nucleotides, acetyl anhydride, glyoxal, formaldehyde, cyanate, metal ions, succinic anhydride, hydroxyl radicals, and fatty acids.
[0029] In some embodiments, the target polypeptide comprises a protein or fragment thereof. In some embodiments, the target polypeptide has a length of at least 50 amino acids.
[0030] In some embodiments, the target polypeptide is comprised in a construct comprising said target polypeptide and one or more of a sequencing adaptor; a linker; and a membrane anchor. In some embodiments, the method comprises attaching the destabilized polypeptide to one or more of a sequencing adaptor; a linker; and a membrane anchor, thereby forming a construct. In some embodiments, the construct comprises a plurality of polypeptides attached together via one or more linkers. In some embodiments, the or each linker independently comprises a polynucleotide, a polypeptide and / or a polysaccharide.
[0031] In some embodiments, the method comprises loading the motor protein onto the target polypeptide or onto a sequencing adapter or linker attached to the target polypeptide. In some embodiments, the motor protein is a helicase. In some embodiments, the motor protein is a NTP driven unfoldase.
[0032] In some embodiments, disrupting the charge of one or more amino acids in the target polypeptide wholly or partially linearizes the target polypeptide.
[0033] Also provided is a method of characterising a target polypeptide, the method comprising
[0034] destabilizing the secondary and / or tertiary structure of the target polypeptide by disrupting the charge of one or more amino acids in the target polypeptide; wherein disrupting the charge of one or more amino acids in the target polypeptide comprises modifying the side chains of said one or more amino acids with one or more chargemodifying moieties;
[0035] contacting the destabilized polypeptide with a motor protein;
[0036] contacting the destabilized polypeptide with a nanopore; and
[0037] taking one or more measurements characteristic of the destabilized polypeptide as the motor protein controls the movement of the destabilized polypeptide with respect to the nanopore; thereby characterising the target polypeptide.
[0038] In some embodiments, the target polypeptide is as defined herein, destabilizing the secondary and / or tertiary structure of the target polypeptide is carried out as defined herein; and / or the motor protein is as defined herein.Also provided is a kit for characterising a target polypeptide, comprising a chemical or enzymatic reagent for disrupting the charge of one or more amino acids in a target polypeptide by modifying the side chains of said one or more amino acids;
[0039] and
[0040] a nanopore capable of detecting one or more characteristics of a target polypeptide as the target polypeptide moves with respect to the nanopore; and / or
[0041] a motor protein capable of controlling the movement of a target polypeptide with respect to a nanopore.
[0042] In some embodiments, the kit comprises a sequencing adapter capable of selectively reacting with the target polypeptide.
[0043] Also provided is a construct comprising a linearized polypeptide having a length of at least 50 amino acids and comprising at least 5 charge-modified amino acids, and a motor protein.
[0044] In some embodiments the construct comprises a plurality of said linearized polypeptides and one or more of a sequencing adaptor; a linker; and a membrane anchor.
[0045] Also provided is the use of a charge-modifying reagent to facilitate translocation of a target polypeptide through a nanopore,
[0046] wherein said use comprises modifying the side chains of one or more amino acids of said target polypeptide with said charge-modifying reagent,
[0047] thereby disrupting the charge of said one or more amino acids in the target polypeptide and destabilizing the secondary and / or tertiary structure of the target polypeptide,
[0048] thereby facilitating the translocation of the target polypeptide through the nanopore.
[0049]
[0050] Figure 1 shows by way of non-limiting illustration some exemplary reactions of amino acids that can be made to disrupt the charge of one or more amino acids in accordance with the provided methods.
[0051] Figure 2 shows a non-limiting exemplary embodiment of an analyte suitable for use in the methods disclosed herein. In Figure 2, (1) is an unfoldase recognition sequence, (2) is an oligonucleotide leader, (3) is a linking moiety introduced by ybbR tagging, and (4) is the protein analyte.Figure 3 shows a non-limiting exemplary embodiment of the disclosed methods. In Figure 3, a protein unfoldase (1) interacts with an analyte with an oligonucleotide leader and recognition sequence (2); the leader promotes capture by a nanopore and initiates threading into the nanopore (3); and the unfoldase controls the translocation of the analyte through the nanopore (4).
[0052] Figure 4 shows exemplary electrophysiology data demonstrating successful controlled translocation of a polypeptide analyte through a nanopore, as described in Example 1.
[0053] Figure 5 shows exemplary electrophysiology data demonstrating that modification of the polypeptide analyte to increase its net negative charge led to successful controlled translocation through the nanopore, as described in Example 1.
[0054] Figure 6 shows plots of charge distribution for unmodified (A) and modified (B) polypeptide analytes as described in Example 1.
[0055] Figure 7 shows a schematic of peptide analytes as described in Example 2.
[0056] Figure 8 shows controlled translocation of unmodified peptides as described in Example 2.
[0057] Figure 9 shows controlled translocation of modified peptides as described in Example 2.
[0058] Detailed Description
[0059] The present invention will be described with respect to particular embodiments and with reference to certain drawings but the invention is not limited thereto but only by the claims. Any reference signs in the claims shall not be construed as limiting the scope. Of course, it is to be understood that not necessarily all aspects or advantages may be achieved in accordance with any particular embodiment of the invention. Thus, for example those skilled in the art will recognize that the invention may be embodied or carried out in a manner that achieves or optimizes one advantage or group of advantages as taught herein without necessarily achieving other aspects or advantages as may be taught or suggested herein.
[0060] The invention, both as to organization and method of operation, together with features and advantages thereof, may best be understood by reference to the following detailed description when read in conjunction with the accompanying drawings. The aspects and advantages of the invention will be apparent from and elucidated with reference to the embodiment(s) described hereinafter. Reference throughout thisspecification to "one embodiment" or "an embodiment" means that a particular feature, structure or characteristic described in connection with the embodiment is included in at least one embodiment of the present invention. Thus, appearances of the phrases "in one embodiment" or "in an embodiment" in various places throughout this specification are not necessarily all referring to the same embodiment, but may. Similarly, it should be appreciated that in the description of exemplary embodiments of the invention, various features of the invention are sometimes grouped together in a single embodiment, figure, or description thereof for the purpose of streamlining the disclosure and aiding in the understanding of one or more of the various inventive aspects. This method of disclosure, however, is not to be interpreted as reflecting an intention that the claimed invention requires more features than are expressly recited in each claim. Rather, as the following claims reflect, inventive aspects lie in less than all features of a single foregoing disclosed embodiment.
[0061] It should be appreciated that “embodiments” of the disclosure can be specifically combined together unless the context indicates otherwise. The specific combinations of all disclosed embodiments (unless implied otherwise by the context) are further disclosed embodiments of the claimed invention.
[0062] In addition as used in this specification and the appended claims, the singular forms “a”, “an”, and “the” include plural referents unless the content clearly dictates otherwise. Thus, for example, reference to “a polynucleotide” includes two or more polynucleotides, reference to “a motor protein” includes two or more such proteins, reference to “a helicase” includes two or more helicases, reference to “a monomer” refers to two or more monomers, reference to “a pore” includes two or more pores and the like.
[0063] All publications, patents and patent applications cited herein, whether supra or infra, are hereby incorporated by reference in their entirety.
[0064] Definitions
[0065] Where an indefinite or definite article is used when referring to a singular noun e.g. "a" or "an", "the", this includes a plural of that noun unless something else is specifically stated. Where the term "comprising" is used in the present description and claims, it does not exclude other elements or steps. Furthermore, the terms first, second, third and the like in the description and in the claims, are used for distinguishing between similar elements and not necessarily for describing a sequential or chronological order. It is to be understood that the terms so used are interchangeable under appropriate circumstances andthat the embodiments of the invention described herein are capable of operation in other sequences than described or illustrated herein. The following terms or definitions are provided solely to aid in the understanding of the invention. Unless specifically defined herein, all terms used herein have the same meaning as they would to one skilled in the art of the present invention. Practitioners are particularly directed to Sambrook et al., Molecular Cloning: A Laboratory Manual, 4thed., Cold Spring Harbor Press, Plainsview, New York (2012); and Ausubel et al., Current Protocols in Molecular Biology (Supplement 114), John Wiley & Sons, New York (2016), for definitions and terms of the art. The definitions provided herein should not be construed to have a scope less than understood by a person of ordinary skill in the art.
[0066] "About" as used herein when referring to a measurable value such as an amount, a temporal duration, and the like, is meant to encompass variations of ± 20 % or ± 10 %, more preferably ± 5 %, even more preferably ± 1 %, and still more preferably ± 0.1 % from the specified value, as such variations are appropriate to perform the disclosed methods.
[0067] “Nucleotide sequence”, “DNA sequence” or “nucleic acid molecule(s)” as used herein refers to a polymeric form of nucleotides of any length, either ribonucleotides or deoxyribonucleotides. This term refers only to the primary structure of the molecule. Thus, this term includes double- and single-stranded DNA, and RNA. The term “nucleic acid” as used herein, is a single or double stranded covalently-linked sequence of nucleotides in which the 3' and 5' ends on each nucleotide are joined by phosphodiester bonds. The polynucleotide may be made up of deoxyribonucleotide bases or ribonucleotide bases. Nucleic acids may be manufactured synthetically in vitro or isolated from natural sources. Nucleic acids may further include modified DNA or RNA, for example DNA or RNA that has been methylated, or RNA that has been subject to post-transcriptional modification, for example 5 ’-capping with 7-methylguanosine, 3 ’-processing such as cleavage and polyadenylation, and splicing. Nucleic acids may also include synthetic nucleic acids (XNA), such as hexitol nucleic acid (HNA), cyclohexene nucleic acid (CeNA), threose nucleic acid (TNA), glycerol nucleic acid (GNA), locked nucleic acid (LNA) and peptide nucleic acid (PNA). Sizes of nucleic acids, also referred to herein as “polynucleotides” are typically expressed as the number of base pairs (bp) for double stranded polynucleotides, or in the case of single stranded polynucleotides as the number of nucleotides (nt). One thousand bp or nt equal a kilobase (kb). Polynucleotides of less than around 40 nucleotidesin length are typically called “oligonucleotides” and may comprise primers for use in manipulation of DNA such as via polymerase chain reaction (PCR).
[0068] The term “amino acid” in the context of the present disclosure is used in its broadest sense and is meant to include organic compounds containing amine (NH2) and carboxyl (COOH) functional groups, along with a side chain (e.g., a R group) specific to each amino acid. In some embodiments, the amino acids refer to naturally occurring L a-amino acids or residues. The commonly used one and three letter abbreviations for naturally occurring amino acids are used herein: A=Ala; C=Cys; D=Asp; E=Glu; F=Phe; G=Gly; H=His; I=Ile; K=Lys; L=Leu; M=Met; N=Asn; P=Pro; Q=Gln; R=Arg; S=Ser; T=Thr; V=Val; W=Trp; and Y=Tyr (Lehninger, A. L., (1975) Biochemistry, 2d ed., pp. 71-92, Worth Publishers, New York). The general term “amino acid” further includes D-amino acids, retro-inverso amino acids as well as chemically modified amino acids such as amino acid analogues, naturally occurring amino acids that are not usually incorporated into proteins such as norleucine, and chemically synthesised compounds having properties known in the art to be characteristic of an amino acid, such as P-amino acids. For example, analogues or mimetics of phenylalanine or proline, which allow the same conformational restriction of the peptide compounds as do natural Phe or Pro, are included within the definition of amino acid. Such analogues and mimetics are referred to herein as "functional equivalents" of the respective amino acid. Other examples of amino acids are listed by Roberts and Vellaccio, The Peptides: Analysis, Synthesis, Biology, Gross and Meiehofer, eds., Vol. 5 p. 341, Academic Press, Inc., N.Y. 1983, which is incorporated herein by reference.
[0069] The terms “polypeptide”, and “peptide” are interchangeably used herein to refer to a polymer of amino acid residues and to variants and synthetic analogues of the same. Thus, these terms apply to amino acid polymers in which one or more amino acid residues is a synthetic non-naturally occurring amino acid, such as a chemical analogue of a corresponding naturally occurring amino acid, as well as to naturally-occurring amino acid polymers. Polypeptides can also undergo maturation or post-translational modification processes that may include, but are not limited to: glycosylation, proteolytic cleavage, lipidization, signal peptide cleavage, propeptide cleavage, phosphorylation, and such like. A peptide can be made using recombinant techniques, e.g., through the expression of a recombinant or synthetic polynucleotide. A recombinantly produced peptide is typically substantially free of culture medium, e.g., culture medium represents less than about 20 %,more typically less than about 10 %, and most typically less than about 5 % of the volume of the protein preparation.
[0070] The term “protein” is used to describe a folded polypeptide having a secondary, tertiary, or quaternary structure. The protein may be composed of a single polypeptide, or may comprise multiple polypeptides that are assembled to form a multimer. The multimer may be a homooligomer, or a heterooligmer. The protein may be a naturally occurring, or wild type protein, or a modified, or non-naturally, occurring protein. The protein may, for example, differ from a wild type protein by the addition, substitution or deletion of one or more amino acids.
[0071] A “variant” of a protein encompass peptides, oligopeptides, polypeptides, proteins and enzymes having amino acid substitutions, deletions and / or insertions relative to the unmodified or wild-type protein in question and having similar biological and functional activity as the unmodified protein from which they are derived. The term "amino acid identity" as used herein refers to the extent that sequences are identical on an amino acid-by-amino acid basis over a window of comparison. Thus, a "percentage of sequence identity" is calculated by comparing two optimally aligned sequences over the window of comparison, determining the number of positions at which the identical amino acid residue (e.g., Ala, Pro, Ser, Thr, Gly, Vai, Leu, He, Phe, Tyr, Trp, Lys, Arg, His, Asp, Glu, Asn, Gin, Cys and Met) occurs in both sequences to yield the number of matched positions, dividing the number of matched positions by the total number of positions in the window of comparison (i.e., the window size), and multiplying the result by 100 to yield the percentage of sequence identity.
[0072] For all aspects and embodiments of the present invention, a “variant” has at least 40%, 50%, 60%, 70%, 80%, 90%, 95% or 99% complete sequence identity to the amino acid sequence of the corresponding wild-type protein. Sequence identity can also be to a fragment or portion of the full length polynucleotide or polypeptide. Hence, a sequence may have only 50 % overall sequence identity with a full length reference sequence, but a sequence of a particular region, domain or subunit could share 80 %, 90 %, or as much as 99 % sequence identity with the reference sequence.
[0073] The term “wild-type” refers to a gene or gene product isolated from a naturally occurring source. A wild-type gene is that which is most frequently observed in a population and is thus arbitrarily designed the “normal” or “wild-type” form of the gene. In contrast, the term “modified”, “mutant” or “variant” refers to a gene or gene product that displays modifications in sequence (e.g., substitutions, truncations, or insertions), post-translational modifications and / or functional properties (e.g., altered characteristics) when compared to the wild-type gene or gene product. It is noted that naturally occurring mutants can be isolated; these are identified by the fact that they have altered characteristics when compared to the wild-type gene or gene product. Methods for introducing or substituting naturally-occurring amino acids are well known in the art. For instance, methionine (M) may be substituted with arginine (R) by replacing the codon for methionine (ATG) with a codon for arginine (CGT) at the relevant position in a polynucleotide encoding the mutant monomer. Methods for introducing or substituting non-naturally-occurring amino acids are also well known in the art. For instance, non-naturally-occurring amino acids may be introduced by including synthetic aminoacyl-tRNAs in the IVTT system used to express the mutant monomer. Alternatively, they may be introduced by expressing the mutant monomer in E. coli that are auxotrophic for specific amino acids in the presence of synthetic (i.e. non-naturally-occurring) analogues of those specific amino acids. They may also be produced by native chemical ligation if the mutant monomer is produced using partial peptide synthesis. Conservative substitutions replace amino acids with other amino acids of similar chemical structure, similar chemical properties or similar side-chain volume. The amino acids introduced may have similar polarity, hydrophilicity, hydrophobicity, basicity, acidity, neutrality or charge to the amino acids they replace. Alternatively, the conservative substitution may introduce another amino acid that is aromatic or aliphatic in the place of a pre-existing aromatic or aliphatic amino acid. Conservative amino acid changes are well-known in the art and may be selected in accordance with the properties of the 20 main amino acids as defined in Table 1 below. Where amino acids have similar polarity, this can also be determined by reference to the hydropathy scale for amino acid side chains in Table 2.Table 1 - Chemical properties of amino acids
[0074]
[0075] Table 2 - Hydropathy scale
[0076] Side Chain Hydropathy
[0077] He T5
[0078] Vai 4.2
[0079] Leu 3.8
[0080] Phe 2.8
[0081] Cys 2.5
[0082] Met 1.9
[0083] Ala 1.8
[0084] Gly -0.4
[0085] Thr -0.7
[0086] Ser -0.8
[0087] Trp -0.9
[0088] Tyr -1.3
[0089] Pro -1.6
[0090] His -3.2
[0091] Glu -3.5
[0092] Gin -3.5
[0093] Asp -3.5
[0094] Asn -3.5
[0095] Lys -3.9
[0096] Arg -4.5
[0097] A mutant or modified protein, monomer or peptide can also be chemically modified in any way and at any site. A mutant or modified monomer or peptide may be chemically modified by attachment of a molecule to one or more cysteines (cysteine linkage), attachment of a molecule to one or more lysines, attachment of a molecule to one or more non-natural amino acids, enzyme modification of an epitope or modification of a terminus. Suitable methods for carrying out such modifications are well-known in the art. Themutant of modified protein, monomer or peptide may be chemically modified by the attachment of any molecule. For instance, the mutant of modified protein, monomer or peptide may be chemically modified by attachment of a dye or a fluorophore.
[0098] As used herein, an alkylene group is a bidentate moiety derived by abstraction of two hydrogen atoms from a linear or branched alkyl group. Typically an alkylene group comprises from 1 to 10 carbon atoms and is referred to as a Ci-io alkylene group. A Ci-io alkylene group is often a Ci-4 alkylene group, or a C1-3 alkylene group. Examples of C1-4 alkylene groups include methylene, ethylene, n-propylene, iso-propylene, n-butylene, secbutylene, and tert-butylene.
[0099] As used herein, an alkenylene group is a bidentate moiety derived by abstraction of two hydrogen atoms from a linear or branched linear alkenyl group having one or more, e.g. one or two, typically one double bonds. Typically an alkenylene group comprises from 2 to 10 carbon atoms and is referred to as a C2-10 alkenylene group. A C2-10 alkenylene group is often a C2 to C4 alkenylene group or a C2 to C3 alkenylene group. Examples of C2 to C4 alkenylene groups include ethenylene, propenylene and butenylene.
[0100] As used herein, a alkynylene group is a bidentate moiety derived by abstraction of two hydrogen atoms from a linear or branched linear alkynyl group having one or more, e.g. one or two, typically one triple bonds. Typically an alkynylene group comprises from 2 to 10 carbon atoms and is referred to as a C2-10 alkynylene group. A C2-10 alkenylene group is often a C2 to C4 alkynylene group or a C2 to C3 alkynylene group. Examples of C2 to C4 alkynylene groups include ethynylene, propynylene and butynylene.
[0101] An arylene group is a bidentate moiety derived from an aryl group. As used herein, an aryl group is often a G, to C10 aryl group which may be a substituted or unsubstituted, monocyclic or fused polycyclic aromatic group containing from 6 to 10 carbon atoms in the ring portion. Examples include monocyclic groups such as phenyl and fused bicyclic groups such as naphthyl and indenyl.
[0102] A heteroarylene group is a bidentate moiety derived from a heteroaryl group. As used herein, an heteroaryl group is often a 5- to 10- membered heteroaryl group which may be a substituted or unsubstituted monocyclic or fused polycyclic aromatic group containing from 5 to 10 atoms in the ring portion, including at least one heteroatom, for example 1, 2 or 3 heteroatoms, typically selected from O, S and N. A heteroaryl group is typically a 5-or 6-membered heteroaryl group or a 9- or 10- membered heteroaryl group. Examples include imidazole, pyridine, pyrimidine and pyrazine.A carbocyclylene group is a bidentate moiety derived from a carbocyclyl group. As used herein, a carbocyclyl group is often a 4-10- or 4-6 membered carbocyclic group containing from 4 to 10 carbon atoms. A carbocyclic group may be saturated or partially unsaturated, but is typically saturated. Examples of carbocyclic groups include cyclobutyl, cyclopentyl and cyclohexyl groups.
[0103] A heterocyclylene group is a bidentate moiety derived from a heterocyclyl group. As used herein, a heterocyclyl group is often a 4-10- or 4-6 membered heterocyclic group containing from 4 to 10 atoms in the ring portion, including at least one heteroatom, for example 1, 2 or 3 heteroatoms, typically selected from O, S and N. A heterocyclic group may be saturated or partially unsaturated, but is typically saturated. Examples of heterocyclic groups include azetidine, morpholine, 1,4-oxazepane, octahydropyrrolo[3,4-c]pyrrole, piperazine, piperidine, and pyrrolidine.
[0104] Disclosed Methods
[0105] In an aspect, provided herein is a method of moving a target polypeptide with respect to a nanopore;
[0106] the method comprising destabilizing the secondary and / or tertiary structure of the target polypeptide by disrupting the charge of one or more amino acids in the target polypeptide; wherein disrupting the charge of one or more amino acids in the target polypeptide comprises modifying the side chains of said one or more amino acids with one or more charge-modifying moieties; and
[0107] contacting the destabilized polypeptide with a motor protein under conditions such that the motor protein controls the movement of the polypeptide with respect to the nanopore.
[0108] In another aspect, provided herein is a method of characterising a target polypeptide, the method comprising
[0109] destabilizing the secondary and / or tertiary structure of the target polypeptide by disrupting the charge of one or more amino acids in the target polypeptide; wherein disrupting the charge of one or more amino acids in the target polypeptide comprises modifying the side chains of said one or more amino acids with one or more chargemodifying moieties;
[0110] contacting the destabilized polypeptide with a motor protein;
[0111] contacting the destabilized polypeptide with a nanopore; andtaking one or more measurements characteristic of the destabilized polypeptide as the motor protein controls the movement of the destabilized polypeptide with respect to the nanopore; thereby characterising the target polypeptide.
[0112] In another aspect, provided herein is the use of a charge-modifying reagent to facilitate translocation of a target polypeptide through a nanopore,
[0113] wherein said use comprises modifying the side chains of one or more amino acids of said target polypeptide with said charge-modifying reagent,
[0114] thereby disrupting the charge of said one or more amino acids in the target polypeptide and destabilizing the secondary and / or tertiary structure of the target polypeptide,
[0115] thereby facilitating the translocation of the target polypeptide through the nanopore.
[0116] As used herein, the term “destabilized” refers to the disruption of the secondary and / or tertiary structure of the target polypeptide. The term “destabilized” may be used synonymously with the term “destructured” or “denatured”. Thus, the disclosed methods can be understood as comprising destructuring (or denaturing) the secondary and / or tertiary structure of the target polypeptide by disrupting the charge of one or more amino acids in the target polypeptide; wherein disrupting the charge of one or more amino acids in the target polypeptide comprises modifying the side chains of said one or more amino acids with one or more charge-modifying moieties.
[0117] As those skilled in the art will appreciate, peptides may in some embodiments have a defined secondary structure, and may comprise one or more secondary structural elements such as alpha-helices or beta-sheets. Similarly, peptides may have a defined tertiary structure, arising from the three-dimensional arrangement of the peptide. Such structures if maintained may prevent or hinder the facile interaction of the polypeptide with a nanopore, and so disrupting such structures thereby destabilizing the polypeptide can facilitate nanopore translocation. Both secondary and tertiary structure of polypeptides may be determined (at least in part) by the interactions between amino acids in the polypeptide. Disrupting the charge of one or more amino acids in the polypeptide may therefore “destabilize” (also referred to as “destructure” or “denature”) the secondary and / or tertiary structure of the target polypeptide.
[0118] Destabilizing / destructuring / denaturing the target polypeptide in the disclosed methods may also be understood in terms of “linearizing” or “unfolding” the target polypeptide. As those skilled in the art will appreciate, the term “linearizing” a targetpolypeptide does not imply that all atoms or amino acids in the peptide lie along a single rigid axis. Rather, a linearized peptide may be understood as a peptide that is substantially free of native secondary and / or tertiary structure, whilst retaining conformational flexibility. Methods of linearizing or unfolding peptides are well known in the art; for example through the use of high concentrations of chaotropic agents such as urea.
[0119] In some embodiments, the disclosed methods therefore wholly or partially linearize the target polypeptide. In some embodiments the disclosed methods wholly or partially denature the target polypeptide. In some embodiments the disclosed methods wholly or partially destructure the target polypeptide. In some embodiments at least 10%, at least 20%, at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, at least 80% or at least 90% of the target polypeptide is linearized. The destructuring of a target polypeptide’s native structure can be determined using any suitable method, such as circular dichroism and FTIP spectroscopy. The effect of one or more modifications as described herein on a target polypeptide can thus be determined and confirmed by experimentally by those skilled in the art. The effect of one or more modifications as described herein on a target polypeptide can also be predicted through the use of computer-aided modelling of protein structures. There are many structural prediction models available to the skilled person including AlphaFold (Jumper et al, Nature 2021 and Varadi et al, Nucleic Acids Research 2021) and RaptorX (Peng et al, Proteins, 2011).
[0120] The disclosed methods thus comprise disrupting the charge of one or more amino acids in a target polypeptide, by modifying the side chains of one or more amino acids of said target polypeptide with a charge-modifying reagent.
[0121] As explained in more detail herein, disrupting the charge of one or more amino acids in a target polypeptide refers to altering the native charge of the amino acids and thus altering the charge distribution in the target polypeptide. The charge may be disrupted by increasing or decreasing negative or positive charge, and thus may comprise making an amino acid more or less positively charged than natively in the target polypeptide, or more or less negatively charged than natively in the target polypeptide. More generally, disrupting the charge of one or more amino acids in a target polypeptide may comprise introducing or removing charge from the one or more amino acids. This is described in more detail herein.
[0122] In the disclosed methods, any suitable target polypeptide can be used. Target polypeptides are described in more detail herein.The target polypeptide comprises one or more amino acids which can be modified as disclosed herein. In some embodiments a plurality of amino acids in the target polypeptide are modified. The amino acids which are modified in the target polypeptide may in some embodiments be referred to as target amino acids. For avoidance of doubt, and as described in more detail herein, the amino acids which are modified (i.e. the target amino acids) may be located at any position within the target polypeptide. In some embodiments a plurality of target amino acids are distributed randomly or pseudo-randomly throughout the target polypeptide. In some embodiments target amino acids may be adjacent to each other
[0123] In some embodiments a plurality of amino acids in the target polypeptide which are the same are modified. Thus, for example, in some embodiments a specific type of amino acid may be chosen or selected for modification. The method may then comprise modification of a plurality of that amino acid present in the target polypeptide. In some embodiments a plurality of amino acids in the target polypeptide which are different are modified. Thus, for example, in some embodiments multiple specific types of amino acid may be chosen or selected for modification. The method may then comprise modification of one or more of a first type of amino acid present in the target polypeptide and modification of one or more of a second type of amino acid present in the target polypeptide. In some embodiments the amino acids in the target polypeptide which are modified in the provided methods are charged amino acids, but this is not essential to the invention. This is described in more detail herein.
[0124] The one or more amino acids are modified with charge-modifying moieties.
[0125] Suitable charge modifying moieties are described in more detail herein.
[0126] The one or more amino acids are modified with the charge-modifying moieties by modifying the side chains of the of the amino acids. As those skilled in the art will appreciate and as described in more detail herein, the side chain of an amino acid may also be known as the “R” group and constitutes the chemical moiety attached to the a-carbon of the amino acid.
[0127] In some embodiments the side chain is modified by covalently attaching said charge-modifying moieties to the side chain. In some embodiments the side chain is modified by non-covalently attaching said charge-modifying moieties to the side chain.
[0128] In some embodiments the side chain is modified by covalently attaching said charge-modifying moieties to the side chain. This can be done in a “pure” chemical reaction by contacting the target polypeptide with one or more chemical reagents, or bycontacting the target polypeptide with one or more amino-acid modifying enzymes.
[0129] Suitable amino-acid modifying enzymes and suitable chemical reagents are described in more detail herein.
[0130] In some embodiments the target polypeptide is comprised in a construct. In some embodiments the construct may comprise one or more of a sequencing adaptor; a linker; and a membrane anchor. Suitable sequencing adapters are described in more detail herein. In some embodiments the construct may comprise a plurality of polypeptides attached together via one or more linkers. In some embodiments the or each linker independently comprises a polynucleotide, a polypeptide and / or a polysaccharide. This is described in more detail herein.
[0131] In some embodiments the target polypeptide, once modified (i.e. destabilized) as described herein, is moved with respect to a nanopore. In some embodiments the movement of the destabilized polypeptide is controlled using a motor protein. Any suitable motor protein can be used as described herein.
[0132] The destabilized polypeptide may be provided in the form of a construct comprising a linearized polypeptide having a length of at least 50 amino acids and comprising at least 5 charge-modified amino acids, and a motor protein. The construct may comprise a plurality of said linearized polypeptides and one or more of a sequencing adaptor; a linker; and a membrane anchor. The construct has a variety of uses. The motor protein may be used to control the movement of the construct, for example with respect to a nanopore. If measurements characteristic of the construct are taken as the construct moves with respect to the nanopore, the construct can be characterised; and in doing so the target polypeptide may be characterised. However, the methods are not limited to generation of constructs for characterisation (whether or not using a nanopore).
[0133] When the target polypeptide is to be characterised, it can be characterised in any suitable method. Most generally, it can be characterised by contacting it (for example in the form of a construct as described herein) with a suitable detector. As further described herein, a nanopore is provided as an exemplary detector which can be used in the disclosed methods. Thus, whilst embodiments described herein refer to characterisation of an oligopeptide-adapter adduct using a nanopore, the methods provided herein are also amenable to other detectors including (i) a zero-mode waveguide, (ii) a field-effect transistor, optionally a nanowire field-effect transistor; (iii) an AFM tip; (iv) a nanotube, optionally a carbon nanotube and (v) a nanopore. The disclosed methods are particularlyamenable to methods in which a polypeptide is moved through a detector or through a structure containing a detector, e.g. a well in a detector chip.
[0134] Some disclosed methods involve taking one or more measurements characteristic of the destabilized (also referred to as destructured / denatured / linearized / unfolded) polypeptide as it moves with respect to a nanopore. Some suitable measurements are described in more detail herein.
[0135] Thus, the disclosed methods have several advantages. They allow for efficient processing of a target polypeptide with respect to a nanopore. This may be useful to characterise the target polypeptide, or the movement of the polypeptide with respect to a nanopore may be used in other applications such as to controllably provide the polypeptide to a process element in a broader reaction context. The disclosed methods avoid or reduce the problems associated with some prior methods which involve moving polypeptides with respect to nanopores.
[0136] Further details of the disclosed methods are described in more detail herein.
[0137] Target Polypeptide
[0138] As explained herein, the disclosed methods comprise destabilizing the secondary and / or tertiary structure of a target polypeptide by disrupting the charge of one or more amino acids in the target polypeptide. Any suitable target polypeptide can be addressed in such methods.
[0139] As used herein, the term polypeptide refers to a peptide, polypeptide or protein which may, for example, be intended for characterisation in accordance with the present methods. The term polypeptide and peptide, polypeptide or protein can be used interchangeably unless implied otherwise by the context, but typically a peptide is a portion of a target polypeptide which may be derived (for example) by cleavage of a polypeptide as described in more detail herein.
[0140] In some embodiments the target polypeptide is an unmodified protein or a portion or fragment thereof, or a naturally occurring polypeptide or a portion or fragment thereof.
[0141] In some embodiments the target polypeptide is a modified protein or a portion thereof.
[0142] In some embodiments the target polypeptide is a denatured protein or a portion thereof. In some embodiments a protein may be denatured by contacting it with one or more denaturing conditions. Suitable denaturing conditions include chaotropic agents such as urea (often used at a concentration of about 5 to about 10 M, e.g. about 6-8 M),guanidinium chloride (often used at a concentration of about 5 to about 10 M, e.g. about 5-7 M), lithium perchlorate (often used at a concentration of about 3 to about 6 M, e.g. about 4.5 M); and detergents such as sodium dodecyl sulfate and the like.
[0143] Those skilled in the art will appreciate that in some embodiments the methods comprise selectively disrupting the charge of one or more amino acids in the target polypeptide. In some embodiments selectively disrupting the charge of one or more amino acids or types of amino acid in the target polypeptide comprises selectively targeting one or more amino acids and selectively not disrupting the charge of one or more other types of amino acids. Thus, in some embodiments the reagent or condition used in such selective modification of the one or more amino acids in the target polypeptide (described in more detail herein) is selective for one or more first types of amino acid in the target polypeptide. In some embodiments contacting a target polypeptide with a chaotrope such as urea or guanidinium chloride does not comprise selectively disrupting the charge of one or more amino acids in the target polypeptide. For example, in some embodiments this may be because such contacting disrupts the charge of many amino acids in the polypeptide in a non-selective manner.
[0144] In some embodiments selectively disrupting the charge of one or more amino acids or types of amino acid in the target polypeptide comprises selectively targeting the target polypeptide and not targeting other proteins that may be present in a composition, such as a motor protein or a protein nanopore. Thus, in some embodiments the reagent or condition used in such selective modification of the one or more amino acids in the target polypeptide (described in more detail herein) is selective for the target polypeptide. In some embodiments contacting a sample comprising the target polypeptide and a further protein such as a nanopore or motor protein with a chaotrope such as urea or guanidinium chloride does not comprise selectively disrupting the charge of one or more amino acids in the target polypeptide. For example in some embodiments this may be because such contacting is non-selective for the polypeptide.
[0145] The amino acids that are modified in the disclosed methods are described in more detail herein.
[0146] In some embodiments the target polypeptide is secreted from cells. Alternatively, the polypeptide can be produced inside cells such that it must be extracted from cells for use in the disclosed methods. The target polypeptide may comprise the products of cellular expression of a plasmid, e.g. a plasmid used in cloning of proteins in accordancewith the methods described in Sambrook et al., Molecular Cloning: A Laboratory Manual, 4thed., Cold Spring Harbor Press, Plainsview, New York (2012); and Ausubel et al., Current Protocols in Molecular Biology (Supplement 114), John Wiley & Sons, New York (2016).
[0147] In some embodiments, for example, the expression is expression in a bacterial cell, a yeast cell, an insect cell or a mammalian cell; or may be a cell free expression method such as a translation system selected from rabbit reticulocyte lysate, wheat germ extract, and E. coli cell-free systems (available commercially, such as from the PURExpress® systems available from New England Biolabs (Ipswich, MA, USA)). In some embodiments the expression is from the genomic DNA of an organism.
[0148] The target polypeptide may be obtained from or extracted from any organism or microorganism. The target polypeptide may be obtained from a human or animal, e.g. from urine, lymph, saliva, mucus, seminal fluid or amniotic fluid, or from whole blood, plasma or serum. The target polypeptide may be obtained from a plant e.g. a cereal, legume, fruit or vegetable. The target polypeptide may be obtained from bacteria, protozoa, algae or fungi.
[0149] The target polypeptide can be provided as an impure mixture of one or more polypeptides and one or more impurities. Impurities may comprise truncated forms of the polypeptide which are distinct from the intended polypeptide for use in the disclosed methods. For example, the target polypeptide may be a full length protein and impurities may comprise fractions of the protein. Impurities may also comprise proteins other than the polypeptide e.g. which may be co-purified from a cell culture or obtained from a sample.
[0150] A target polypeptide may comprise any combination of any amino acids, amino acid analogs and modified amino acids (i.e. amino acid derivatives). Amino acids (and derivatives, analogs etc) in the polypeptide can be distinguished by their physical size and charge.
[0151] The amino acids / derivatives / analogs can be naturally occurring or artificial.
[0152] In some embodiments the target polypeptide may comprise any naturally occurring amino acid. Twenty amino acids are encoded by the universal genetic code. These are alanine (A), arginine (R), asparagine (N), aspartic acid (D), cysteine (C), glutamic acid / glutamate (E), glutamine (Q), glycine (G), histidine (H), isoleucine (I), leucine (L), lysine (K), methionine (M), phenylalanine (F), proline (P), serine (S), threonine (T),tryptophan (W), tyrosine (Y) and valine (V). Other naturally occurring amino acids include selenocysteine and pyrrolysine.
[0153] In some embodiments the target polypeptide is also modified (e.g. prior to disrupting the charge of one or more amino acids in the target polypeptide) by one or more further modifications. In some embodiments the disclosed methods are for characterising a target polypeptide that has been previously modified. In some embodiments the disclosed methods are for characterising modifications (e.g. such previous modifications) in the target polypeptide.
[0154] For example, in some embodiments the target polypeptide may be post-translationally modified prior to being addressed in the disclosed methods. As such, the methods disclosed herein can be used to detect the presence, absence, number of positions of post-translational modifications in a target polypeptide. The disclosed methods can be used to characterise the extent to which a target polypeptide has been post-translationally modified.
[0155] Any one or more post-translational modifications may be present in the target polypeptide. Typical post-translational modifications include modification with a hydrophobic group, modification with a cofactor, addition of a chemical group, glycation (the non-enzymatic attachment of a sugar), biotinylation and pegylation. Post-translational modifications can also be non-natural, such that they are chemical modifications done in the laboratory for biotechnological or biomedical purposes. This can allow monitoring the levels of the laboratory made polypeptide in contrast to the natural counterparts.
[0156] Examples of post-translational modification with a hydrophobic group include myristoylation, attachment of myristate, a Ci4 saturated acid; palmitoylation, attachment of palmitate, a Ci6 saturated acid; isoprenylation or prenylation, the attachment of an isoprenoid group; famesylation, the attachment of a farnesol group; geranylgeranylation, the attachment of a geranylgeraniol group; and glypiation, and glycosylphosphatidylinositol (GPI) anchor formation via an amide bond.
[0157] Examples of post-translational modification with a cofactor include lipoylation, attachment of a lipoate (Cs) functional group; flavination, attachment of a flavin moiety (e.g. flavin mononucleotide (FMN) or flavin adenine dinucleotide (FAD)); attachment of heme C, for instance via a thioether bond with cysteine; phosphopantetheinylation, the attachment of a 4'-phosphopantetheinyl group; and retinylidene Schiff base formation.
[0158] Examples of post-translational modification by addition of a chemical group include acylation, e.g. O-acylation (esters), N-acylation (amides) or S-acylation(thioesters); acetylation, the attachment of an acetyl group for instance to the N-terminus or to lysine; formylation; alkylation, the addition of an alkyl group, such as methyl or ethyl; methylation, the addition of a methyl group for instance to lysine or arginine; amidation; butyrylation; gamma-carboxylation; glycosylation, the enzymatic attachment of a glycosyl group for instance to arginine, asparagine, cysteine, hydroxylysine, serine, threonine, tyrosine or tryptophan; poly si alyl ati on, the attachment of polysialic acid; malonylation; hydroxylation; iodination; bromination; citrulination; nucleotide addition, the attachment of any nucleotide such as any of those discussed herein, ADP ribosylation; oxidation; phosphorylation, the attachment of a phosphate group for instance to serine, threonine or tyrosine (O-linked) or histidine (N-linked); adenylyl ati on, the attachment of an adenylyl moiety for instance to tyrosine (O-linked) or to histidine or lysine (N-linked); propionylation; pyroglutamate formation; S-glutathionylation; Sumoylation; S-nitrosylation; succinylation, the attachment of a succinyl group for instance to lysine; sei enoyl ati on, the incorporation of selenium; and ubiquitinilation, the addition of ubiquitin subunits (N-linked).
[0159] It is within the scope of the methods provided herein that the target polypeptide is labelled with a molecular label. A molecular label may be a modification to the target polypeptide which promotes the detection of the polypeptide in the methods provided herein. For example the label may be a modification to the polypeptide which alters the signal obtained as a conjugate comprising the polypeptide (e.g. a construct comprising a plurality of said polypeptides attached together via one or more linkers wherein said linkers in some embodiments comprise a polynucleotide, a polypeptide and / or a polysaccharide) is characterised. For example, the label may interfere with a flux of ions through the nanopore. In such a manner, the label may improve the sensitivity of the methods.
[0160] In some embodiments the target polypeptide contains one or more cross-linked sections, e.g. C-C bridges. In some embodiments the target polypeptide is not cross-linked prior to the disclosed methods.
[0161] In some embodiments the target polypeptide comprises sulphide-containing amino acids and thus has the potential to form disulphide bonds. Typically, in such embodiments, the target polypeptide is reduced using a reagent such as DTT (Dithiothreitol) or TCEP (tris(2-carboxyethyl)phosphine) prior to being characterised using the disclosed methods.
[0162] In some embodiments the target polypeptide is a full length protein or naturally occurring polypeptide.The target polypeptide can be a polypeptide of any suitable length. In some embodiments the target polypeptide has a length of from about 2 peptide units to about 100, about 200, about 300, about 400, about 500, about 1000, about 5,000, about 10,000, about 15,000, about 20,000, about 30,000 or about 40,000 peptide units. In some embodiments the target polypeptide has a length of at least 25, at least 30, at least 40, at least 50, at least 100, at least 150, at least 200, at least 300 or at least 400 amino acids. In some embodiments the target polypeptide has a length of from about 50 to about 40,000 peptide units. In some embodiments the target polypeptide has a length of from about 50 to about 35,000 peptide units. In some embodiments the target polypeptide has a length of from about 50 to about 20,000 peptide units. In some embodiments the target polypeptide has a length of from about 50 to about 10,000 peptide units. In some embodiments the target polypeptide has a length of from about 100 to about 5,000 peptide units, for example from about 200 to about 1000 peptide units, e.g. from about 300 to about 500 peptide units, such as from about 200 to about 400 peptide units.
[0163] Any number of target polypeptides can be used in the disclosed methods. For instance, the method may comprise processing 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 50, 100, 200, 500, 1000, 2000 or more target polypeptides or about 10, about 50, about 100, about 200, about 500, about 1000, or about 2000 target polypeptides. If two or more target polypeptides are used, they may be different target polypeptides or two or more instances of the same target polypeptide.
[0164] In some embodiments the target polypeptide to be processed in the disclosed methods is present in a sample. In some embodiments the sample comprises a plurality of different polypeptides. In some embodiments the plurality may comprise at least 2, at least 3, at least 4, at least 5, at least 10, at least 20, at least 50, at least 100, at least 500, at least 1000, at least 5000, at least 10000, at least 100000, at least 1000000 or more polypeptides.
[0165] Modifying amino acids
[0166] In the disclosed methods, a target polypeptide as described herein is destabilized by disrupting the charge of one or more amino acids in the target polypeptide. As described herein in more detail, the charge of the one or more amino acids is disrupted by modifying the side chains of said one or more amino acids with one or more charge-modifying moieties. The amino acids which are modified in the disclosed methods may in some embodiments be referred to as target amino acids.As described herein, in some embodiments a plurality of amino acids in the target polypeptide are modified. In some embodiments the disclosed methods comprise modifying at least one, at least two, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 20, at least 50, at least 100, at least 500, at least 1000, at least 5000, at least 10000, or more amino acids in the or each target polypeptide.
[0167] The one or more target amino acids which are modified in the disclosed methods may be located at any position within the target polypeptide.
[0168] In some embodiments the target amino acids which are modified in the disclosed methods may be identified by consideration of the native (unmodified) structure of the target polypeptide. For example, in some embodiments the target amino acids can be inferred from the charges of the amino acids in the target polypeptide: for instance, it can be inferred that a positively charged amino acid may interact with a negatively charged amino acid if the amino acids are located proximate to each other in the sequence.
[0169] Alternatively, more detailed studies can be made. For example, x-ray crystallography can be used to determined the amino acids primarily responsible for stabilizing the secondary and / or tertiary structure of the target polypeptide. The identified amino acids can then be modified in accordance with the disclosed methods in order to destabilize the secondary and / or tertiary structure of the target polypeptide.
[0170] In some embodiments a plurality of amino acids for modifying in the disclosed methods are distributed randomly or pseudo-randomly throughout the target polypeptide. In some embodiments the amino acids in the target polypeptide that are modified are separated from each other by between about 1 and about 20 amino acids, such as between about 2 and about 15 amino acids e.g. between about 3, about 4, or about 5 amino acids and about 8, about 9 or about 10 amino acids. In some embodiments the amino acids that are modified are adjacent to each other.
[0171] In some embodiments the amino acids in the in the target polypeptide that are modified are separated from each other by more than about 20, such as more than about 30, e.g. more than about 40 or more than about 50 amino acids. For example, in some embodiments two or more amino acids which interact in the target polypeptide may be modified, wherein the two or more amino acids are spatially proximate in the target polypeptide before its modification. In some embodiments this can occur despite the two or more amino acids being separate from each other in the amino acid sequence of the target polypeptide.In some embodiments one or more amino acids which are modified in the disclosed methods may interact via one or more charge-mediated interactions with one or more other amino acids in the target polypeptide, thereby stabilizing the secondary and / or tertiary structure of the target polypeptide. Thus, in some embodiments the method comprises modifying one or more amino acids in the target polypeptide which may interact via one or more charge-mediated interactions with one or more further amino acids in the target polypeptide, thereby destabilizing the secondary and / or tertiary structure of the target polypeptide. In some embodiments the method comprises modifying two or more amino acids in the target polypeptide which interact via one or more charge-mediated interactions in the target polypeptide, thereby destabilizing the secondary and / or tertiary structure of the target polypeptide. In some embodiments amino acids suitable for modification in such methods are charged amino acids. For example, in some embodiments the secondary and / or tertiary structure of the target polypeptide is stabilized by interaction of one or more positively charged amino acids with one or more negatively charged amino acids, and the provided methods comprise modifying the one or more positively charged amino acids thereby destabilizing the secondary and / or tertiary structure of the target polypeptide. In some embodiments the secondary and / or tertiary structure of the target polypeptide is stabilized by interaction of one or more positively charged amino acids with one or more negatively charged amino acids, and the provided methods comprise modifying the one or more negatively charged amino acids thereby destabilizing the secondary and / or tertiary structure of the target polypeptide. In some embodiments the secondary and / or tertiary structure of the target polypeptide is stabilized by interaction of one or more positively charged amino acids with one or more negatively charged amino acids, and the provided methods comprise modifying the one or more positively charged amino acids and the one or more negatively charged amino acids, thereby destabilizing the secondary and / or tertiary structure of the target polypeptide.
[0172] In some embodiments one or more amino acids which are modified in the disclosed methods may interact via one or more hydrophobic interactions with one or more other amino acids in the target polypeptide, thereby stabilizing the secondary and / or tertiary structure of the target polypeptide. Thus, in some embodiments the method comprises modifying one or more amino acids in the target polypeptide which may interact via one or more hydrophobic interactions with one or more further amino acids in the target polypeptide, thereby destabilizing the secondary and / or tertiary structure of the target polypeptide. In some embodiments the method comprises modifying two or more aminoacids in the target polypeptide which interact via one or more hydrophobic interactions in the target polypeptide, thereby destabilizing the secondary and / or tertiary structure of the target polypeptide. In some embodiments amino acids suitable for modification in such methods are uncharged and / or aromatic amino acids. For example, in some embodiments the secondary and / or tertiary structure of the target polypeptide is stabilized by interaction of two or more uncharged and / or aromatic amino acids, and the provided methods comprise modifying one or more of the uncharged and / or aromatic amino acids thereby destabilizing the secondary and / or tertiary structure of the target polypeptide. In some embodiments the secondary and / or tertiary structure of the target polypeptide is stabilized by interaction of two or more uncharged and / or aromatic amino acids, and the provided methods comprise modifying two or more of the uncharged and / or aromatic amino acids thereby destabilizing the secondary and / or tertiary structure of the target polypeptide.
[0173] In some embodiments one or more amino acids which are modified in the disclosed methods may interact via hydrogen-bonding with one or more other amino acids in the target polypeptide, thereby stabilizing the secondary and / or tertiary structure of the target polypeptide. Thus, in some embodiments the method comprises modifying one or more amino acids in the target polypeptide which may interact via hydrogen bonding with one or more further amino acids in the target polypeptide, thereby destabilizing the secondary and / or tertiary structure of the target polypeptide. In some embodiments the method comprises modifying two or more amino acids in the target polypeptide which interact via hydrogen bonding in the target polypeptide, thereby destabilizing the secondary and / or tertiary structure of the target polypeptide. In some embodiments amino acids suitable for modification in such methods are polar amino acids. In some embodiments amino acids suitable for modification in such methods have side chains comprising hydrogen-bond donor groups or hydrogen bond acceptor groups. For example, in some embodiments the secondary and / or tertiary structure of the target polypeptide is stabilized by interaction of one or more amino acids comprising one or more hydrogen-bond donor groups and of one or more amino acids comprising one or more hydrogen-bond acceptor groups, and the methods comprise modifying one or more of the amino acids comprising one or more hydrogen-bond donor groups thereby destabilizing the secondary and / or tertiary structure of the target polypeptide. In some embodiments the secondary and / or tertiary structure of the target polypeptide is stabilized by interaction of one or more amino acids comprising one or more hydrogen-bond donor groups and of one or more amino acids comprising one or more hydrogen-bond acceptor groups, and the methods comprise modifying one or moreof the amino acids comprising one or more hydrogen-bond acceptor groups thereby destabilizing the secondary and / or tertiary structure of the target polypeptide. In some embodiments the secondary and / or tertiary structure of the target polypeptide is stabilized by interaction of one or more amino acids comprising one or more hydrogen -bond donor groups and of one or more amino acids comprising one or more hydrogen-bond acceptor groups, and the methods comprise modifying one or more of the amino acids comprising one or more hydrogen-bond donor groups and one or more of the amino acids comprising one or more hydrogen-bond acceptor groups thereby destabilizing the secondary and / or tertiary structure of the target polypeptide.
[0174] In some embodiments one or more amino acids which are modified in the disclosed methods may interact via one or more covalent interactions (e.g. disulphide bonds) with one or more other amino acids in the target polypeptide, thereby stabilizing the secondary and / or tertiary structure of the target polypeptide. Thus, in some embodiments the method comprises modifying one or more amino acids in the target polypeptide which may interact via one or more covalent interactions (e.g. disulphide bonds) with one or more further amino acids in the target polypeptide, thereby destabilizing the secondary and / or tertiary structure of the target polypeptide. In some embodiments the method comprises modifying two or more amino acids in the target polypeptide which interact via one or more covalent interactions (e.g. disulphide bonds) in the target polypeptide, thereby destabilizing the secondary and / or tertiary structure of the target polypeptide. In some embodiments amino acids suitable for modification in such methods are include cysteine and non-natural thiol-containing amino acids. For example, in some embodiments the secondary and / or tertiary structure of the target polypeptide is stabilized by disulphide bond formation between two or more cysteine amino acids, and the provided methods comprise modifying one or more of the cysteine amino acids thereby destabilizing the secondary and / or tertiary structure of the target polypeptide. In some embodiments the secondary and / or tertiary structure of the target polypeptide is stabilized by disulphide bond formation between two or more cysteine amino acids, and the provided methods comprise modifying two or more cysteine amino acids thereby destabilizing the secondary and / or tertiary structure of the target polypeptide.
[0175] In some embodiments each occurrence of a target amino acid in a polypeptide may be modified in the disclosed methods. For example, as described in more detail herein, in some embodiments chemical reagents may be used to modify the side chains of each occurrence of an amino acid in the target polypeptide. By way of illustration, in some embodiments the chemical reagent is an amine-reactive reagent and the method maycomprise modifying of each occurrence of lysine in the target amino acid. In some embodiments the chemical reagent is a carboxylic acid- (or carboxylate)-reactive reagent and the method may comprise modifying of each occurrence of glutamic acid (glutamate) and aspartic acid (aspartate) in the target amino acid.
[0176] In some embodiments the one or more amino acids that are modified in the disclosed methods are each comprised in a target motif. In some embodiments the target motif comprises from about 2, 3, 4, 5, 6, 7, 8, 9 or 10 amino acids comprising the amino acid to be modified. For example, in some embodiments the target motif may comprise or consist of a recognition sequence for an amino-acid modifying enzyme which may be used to modify the target polypeptide. Through appropriate selection of reagents, the skilled operator of the methods is capable of modifying specific amino acids in the target polypeptide, and thus can specifically target amino acids which contribute to the secondary and / or tertiary structure of the target polypeptide, and therefore can selectively modify such amino acids thereby destabilizing the secondary and / or tertiary structure of the target polypeptide.
[0177] The choice or selection of amino acids to be modified in the disclosed methods is not particularly limited. In some embodiments any amino acid involved in stabilizing the secondary and / or tertiary structure of the target polypeptide can be modified in the provided methods.
[0178] In some embodiments, the or each amino acid modified in the provided methods is selected from Ser, Thr, Tyr, Cys, Lys, Arg, His, Asp, Glu, Pro, Asn, Gin, Met, Trp, Phe, Ala, Vai, Leu and He. In some embodiments the or each amino acid modified in the provided methods is selected from Ser, Thr, Tyr, Cys, Lys, Arg, His, Asp, Glu, Asn, Gin, Met, and Trp. In some embodiments the or each amino acid modified in the provided methods is selected from Ser, Thr, Tyr, Cys, Lys, Arg, Asp, Glu, Asn, and Gin. In some embodiments the or each amino acid modified in the provided methods is selected from Cys, Lys, Arg, Asp, Glu, Asn, and Gin.
[0179] In some embodiments the one or more amino acids that are modified with one or more charge-modifying moieties comprise one or more charged amino acids. In some embodiments the one or more amino acids that are modified with one or more chargemodifying moieties are selected from histidine (His), lysine (Lys), arginine (Arg), glutamic acid / glutamate (Glu) and aspartic acid / aspartate (Asp). In some embodiments the one ormore amino acids that are modified with one or more charge-modifying moieties are selected from Lys, Arg, Glu and Asp.
[0180] In some embodiments the one or more amino acids that are modified with one or more charge-modifying moieties comprise one or more positively charged amino acids and disrupting the charge of said one or more amino acids comprises reducing the positive charge of said amino acids. Thus, in some embodiments the one or more amino acids that are modified with one or more charge-modifying moieties are selected from His, Lys and Arg, and the method comprises modifying the side chain of said one or more amino acids to reduce the positive charge of the said amino acids. In some embodiments the one or more amino acids are one or more lysines. In some embodiments modifying said positively charged amino acids may comprise neutralising the positive charge of said amino acids. In some embodiments the modifying said positively charged amino acids may comprise introducing negatively charged groups to said amino acids thereby making said amino acids negatively charged.
[0181] In some embodiments the one or more amino acids that are modified with one or more charge-modifying moieties comprise one or more negatively charged amino acids and disrupting the charge of said one or more amino acids comprises reducing the negative charge of said amino acids. Thus, in some embodiments the one or more amino acids that are modified with one or more charge-modifying moieties are selected from Asp and Glu, and the method comprises modifying the side chain of said one or more amino acids to reduce the negative charge of the said amino acids. In some embodiments modifying said negatively charged amino acids may comprise neutralising the negative charge of said amino acids. In some embodiments the modifying said negatively charged amino acids may comprise introducing positively charged groups to said amino acids thereby making said amino acids positively charged.
[0182] In some embodiments both one or more positively charged amino acids (e.g. one or more amino acids selected from His, Lys and Arg, particularly Lys) and one or more negatively charged amino acids (e.g. one or more amino acids selected from Asp and Glu) are modified as described herein.
[0183] In some embodiments the one or more amino acids that are modified with one or more charge-modifying moieties comprise one or more amino acids that stabilize the secondary and / or tertiary structure of the target polypeptide by hydrogen bonding. In some embodiments the one or more amino acids that are modified with one or more chargemodifying moieties are selected from Ser, Thr, Tyr, Asn and Gin and the methodcomprises disrupting the charge of said amino acids. In some embodiments modifying said amino acids comprises increasing the positive charge of said amino acids. In some embodiments modifying said amino acids may comprise introducing positively charged groups to said amino acids thereby making said amino acids positively charged. In some embodiments modifying said amino acids comprises increasing the negative charge of said amino acids. In some embodiments modifying said amino acids may comprise introducing negatively charged groups to said amino acids thereby making said amino acids negatively charged.
[0184] In some embodiments the one or more amino acids that are modified with one or more charge-modifying moieties comprise one or more amino acids that stabilize the secondary and / or tertiary structure of the target polypeptide by hydrophobic interactions. In some embodiments the one or more amino acids that are modified with one or more charge-modifying moieties are selected from Ala, Vai, Leu, He, Met, Trp and Phe, particularly Trp and Phe which can engage in pi-pi interactions. In some embodiments modifying said amino acids comprises introducing positive charge to said amino acids. In some embodiments modifying said amino acids may comprise introducing positively charged groups to said amino acids thereby making said amino acids positively charged. In some embodiments modifying said amino acids comprises introducing negative charge to said amino acids. In some embodiments modifying said amino acids may comprise introducing negatively charged groups to said amino acids thereby making said amino acids negatively charged.
[0185] In the provided methods, the one or more amino acids that are modified thereby destabilizing the secondary and / or tertiary structure of the target polypeptide are modified by modifying the side chains of said amino acids with one or more charge-modifying moieties.
[0186] Any suitable charge-modifying moiety can be used. The charge-modifying moiety alters the native charge of the one or more amino acids as described herein (e.g. to introduce charge to a neutral amino acid; to decrease the positive charge of a positively charged amino acid or to decrease the negative charge of a negatively charged amino acid).
[0187] Suitable charge-modifying moieties are thus typically capable of reacting with functional groups present in the side chain of the one or more amino acids.
[0188] In some embodiments a charge-modifying moiety may comprise a reactive group for reacting with the side chain of the amino acid being modified. The reactive group canbe chosen or determined by the user of the disclosed methods to ensure selective reaction with the or each amino acid.
[0189] In some embodiments the or each amino acid natively comprises a reactive side chain for reacting with a charge-modifying moiety. In some embodiments the or each amino acid is modified to comprise a reactive side chain for reacting with a chargemodifying moiety. In some embodiments the methods comprise activating the side chain of the or each first amino acid for reaction with a charge-modifying moiety.
[0190] For example, in some embodiments the side chain of the amino acid comprises a nucleophilic group and the charge-modifying moiety comprises an electrophilic group. In some embodiments the side chain of the amino acid comprises an electrophilic group and the charge-modifying moiety comprises a nucleophilic group.
[0191] In some embodiments modifying the side chains of said one or more amino acids with one or more charge-modifying moieties comprises covalently attaching said one or more charge-modifying moieties to said side chains. For example, in some embodiments modifying the side chains of said one or more amino acids with one or more chargemodifying moieties comprises one or more of oxidation, esterification, O-glycosylation, alkylation, nitration, nitrosylation, succinylation, hydroxylation, amidation, deamidation, acylation, sulfhydration, carb amyl ati on, glycation, metal coordination, Schiff-base formation, and ADP-ribosylation of said side chains. Reagents to achieve these modifications are well known to those skilled in the art. Some exemplary reagents are described herein.
[0192] In some embodiments, modifying the side chains of said one or more amino acids with one or more charge-modifying moieties comprises contacting said one or more amino acids with one or more chemical reagents. The attachment chemistry between the side chain of the amino acid and the charge-modifying moiety is not particularly limited and any suitable chemistry can be used. For example, in some embodiments the one or more chemical reagents are selected from oxidising agents, halogenating agents, nitrating agents, alkylating agents, phosphoric acid derivatives, glycosyl donors, acids, bases, NO donors, sulphide donors, glycation agents, nucleotides, acetyl anhydride, glyoxal, formaldehyde, cyanate, metal ions, succinic anhydride, hydroxyl radicals, and fatty acids. Practitioners are directed to texts such as March's Advanced Organic Chemistry: Reactions, Mechanisms, and Structure (2019), ed. Smith, Wiley; and to G. Hermanson, Bioconjugate Techniques, 3rd Edition (2013), each of which are hereby incorporated by reference intheir entirety. Practitioners are directed particularly to discussion in those texts of reactions of amines and guanidines, carboxylic acids, alcohols, and thiols, especially amines and guanidines.
[0193] Reactive groups and their corresponding targets include aryl azides which may react with amine, carbodiimides which may react with amines and carboxyl groups, hydrazides which may react with carbohydrates, hydroxmethyl phosphines which may react with amines, imidoesters which may react with amines, isocyanates which may react with hydroxyl groups, carbonyls which may react with hydrazines, maleimides which may react with sulfhydryl groups, NHS-esters which may react with amines, PFP-esters which may react with amines, psoralens which may react with thymine, pyridyl disulfides which may react with sulfhydryl groups, vinyl sulfones which may react with sulfhydryl amines and hydroxyl groups, vinylsulfonamides, and the like. In some embodiments a reactive group (e.g. a first and / or second reactive group) is selected from NHS esters, maleimides, imido esters, aldehydes, hydrazides, epoxides, isocyanates, activated carboxylic acids, azides, thiols, iodoacetamides, bromoacetamides, diazirines, photoreactive benzophenones, alkyl halides, succinimides, pyridyldisulfides, amines, carbodiimides, and oxiranes.
[0194] Other suitable chemistry for attaching a charge-modifying moiety to the side chain of an amino acid includes click chemistry. Many suitable click chemistry reagents are known in the art. Suitable examples of click chemistry include, but are not limited to, the following:
[0195] (a) copper(I)-catalyzed azide-alkyne cycloadditions (azide alkyne Huisgen cycloadditions);
[0196] (b) strain-promoted azide-alkyne cycloadditions; including alkene and azide [3+2] cycloadditions; alkene and tetrazine inverse-demand Diels- Alder reactions; and alkene and tetrazole photoclick reactions;
[0197] (c) copper-free variant of the 1,3 dipolar cycloaddition reaction, where an azide reacts with an alkyne under strain, for example in a cyclooctane ring such as in bicycle[6.1.0]nonyne (BCN);
[0198] (d) the reaction of an oxygen nucleophile on one linker with an epoxide or aziridine reactive moiety on the other; and
[0199] (e) the Staudinger ligation, where the alkyne moiety can be replaced by an aryl phosphine, resulting in a specific reaction with the azide to give an amide bond. Thus, any reactive group may be used in the reaction between the side chain of the amino acid and the charge-modifying moiety. Some suitable reactive groups include [1, 4-Bis[3-(2-pyridyldithio)propionamido]butane; 1,1 1-bis-maleimidotriethyleneglycol; 3,3’-dithiodipropionic acid di(N-hydroxysuccinimide ester); ethylene glycol-bis(succinic acid N-hydroxysuccinimide ester); 4,4’ -diisothiocyanatostilbene-2, 2’ -disulfonic acid disodium salt; Bis[2-(4-azidosalicylamido)ethyl] disulphide; 3-(2-pyridyldithio)propionic acidN-hydroxysuccinimide ester; 4-maleimidobutyric acid N-hydroxysuccinimide ester; lodoacetic acid N-hydroxysuccinimide ester; S-acetylthioglycolic acid N-hydroxysuccinimide ester; azide-PEG-maleimide; and alkyne-PEG-maleimide. The reactive group may be any of those disclosed in WO 2010 / 086602, particularly in Table 3 of that application.
[0200] In some embodiments the side chain of the amino acid to be modified comprises an amine group. Examples of such amino acids include lysine. In some embodiments such an amine groups may be reacted with a charge-modifying moieties which comprises an amine-reactive group.
[0201] Examples of amine-reactive groups include: carboxylic acids and activated derivatives thereof (e.g. NHS-esters), which can form amide bonds with amine groups; thiols or activated derivatives thereof, which can form thioether bonds by reaction with amine groups; squaramates, which react with amines to form squaramides; amine coupling agents (e.g. N,N’-Disuccinimidyl carbonate, l,l'-Carbonyldiimidazole) which react with amines to form a urea linkage; and aldehydes and ketones, which react with amines via reductive amination (e.g. via reduction using a reducing agent such as NaBEE or NaBH3CN).
[0202] Further examples of amine-reactive groups are isothiocyanates (e.g. 4-sulfophenyl isothiocyanate (SPITC)), which react with amines to form thiourea linkages. In some embodiments the disclosed methods comprise modifying the side chains of one or more amino acids in the target polypeptide wherein the one or more amino acids comprise amine groups (e.g. wherein the one or more amino acids comprise one or more lysine residues of the target polypeptide) using an amine-reactive group as disclosed herein, such as an isothiocyanate group. In some embodiments the disclosed methods comprise modifying amine groups in the target polypeptide (e.g. comprised in lysine amino acid residues of the target polypeptide) using 4-sulfophenyl isothiocyanate (SPITC)). Modifying lysine residues with 4-sulfophenyl isothiocyanate removes the positive charge at each reacted lysine residue and introduces a negative charge.In one embodiment an amine group comprised in the side chain of an amino acid may be activated for reaction with the charge-modifying moiety. For example, an amine group (e.g. comprised in the side chain of a lysine amino acid) may be modified by reaction with a maleimide-containing compound such as a maleimide-NHS-ester (e.g. 3-maleimido-propionic NHS ester), with the amine group reacting with the NHS-ester to form an amide, optionally followed by reaction of the free maleimide group with a further moiety to alter the charge of the residue, for example, with a thiol group, or with a diene such as a furan group. The charge-modifying moiety may be neutral and thus disrupt the charge of the lysine by reducing the positive charge of the side chain of the lysine, or may be negatively charged (e.g. may comprise a carboxylic acid moiety) and thus disrupt the charge of the lysine by making the side chain of the lysine negatively charged.
[0203] In another embodiment, an amine group (e.g. comprised in the side chain of a lysine amino acid) may be modified for subsequent reaction (e.g. with the chargemodifying moiety) in a click chemistry reaction. Examples include copper-catalysed azide / alkyne cycloaddition (CuAAC) reactions. For example, in one embodiment an amine may be activated by reaction with an azide-containing compound, such as an azidoacetic acid NHS-ester, followed by reaction of the free azide group with the chargemodifying moiety, e.g. with an alkyne group of the charge-modifying moiety. In one embodiment an amine may be activated by reaction with an alkyne-containing compound, followed by reaction of the alkyne group with the charge-modifying moiety, for example with an azide group of the charge-modifying moiety, such as an azidoacetic acid NHS-ester. In one embodiment the click chemistry reaction is a strain-promoted azide-alkyne reaction. In one embodiment an amine may be activated by reaction with an azide-containing compound, such as an azidoacetic acid NHS-ester, followed by reaction of the free azide group with the charge-modifying moiety, for example with a cyclooctynyl group of the charge-modifying moiety (e.g. a bicyclononyne, BCN, group). In one embodiment an amine may be activated by reaction with an cyclooctynyl-containing compound, such as a bicyclononyne (BCN) group, followed by reaction of the alkyne group with the chargemodifying moiety, for example with an azide group of the charge-modifying moiety, such as an azidoacetic acid NHS-ester. In one embodiment the click chemistry reaction is an inverse electron demand Diels-Alder reaction (iEDDA). In one embodiment an amine may be activated by reaction with a tetrazine-containing compound, such as a methyltetrazinecontaining compound, followed by reaction of the free tetrazine group with the chargemodifying moiety, for example with a trans-cyclooctene group of the charge-modifyingmoiety. In one embodiment an amine may be activated by reaction with a trans-cycloocten-containing compound, followed by reaction of the TCO group with the chargemodifying moiety, for example with a tetrazine group (e.g. a methyltetrazine group) of the charge-modifying moiety.
[0204] In some embodiments the side chain of the amino acid to be modified comprises a guanidine group. Examples of such amino acids include arginine. In some embodiments such a guanidine group may be reacted with a charge-modifying moiety which comprises a guanidine-reactive group. Examples of guanidine-reactive groups include diketones such as a 1,2-diketone or 1,3-diketone; and NHS-esters, which can form amide bonds with guanidine groups. In one embodiment a guanidine group may be modified by citrullination via arginine deaminase to form a carbamide group; by using an aldehyde such as formaldehyde; or by reaction with a glyoxal-containing compound. The resulting activated groups may optionally be further modified by reaction with further groups to modify the charge of the residue. Analogous strategies can be used to those described above in the context of activating amine groups.
[0205] In some embodiments the side chain of the amino acid to be modified comprises a carboxyl group. Examples of such amino acids include Asp and Glu. In some embodiments such a carboxyl group may be reacted with a charge-modifying moiety which comprises a carboxyl -reactive group. Examples of carboxyl -reactive groups include amines, which can form amide bonds with carboxyl groups; and alcohols, which can form esters with carboxyl groups. In one embodiment a carboxyl group may be modified using a reagent such as a carbodiimide, followed by reaction with a nucleophilic group (e.g. an amine). In one embodiment a compound comprising a nucleophilic group such as an amine group coupled to a click chemistry group such as an azide or alkyne group may be used, with the carboxyl group being activated by the carbodiimide, the amine group reacting with the activated carboxyl group; and followed by reaction of the free click chemistry group with a charge modifying moiety; for example with a moiety comprising an azide or alkyne group. Analogous strategies can be used to those described above in the context of activating amine groups.
[0206] In some embodiments the side chain of the amino acid to be modified comprises a hydroxyl group. Examples of such amino acids include Ser, Thr and Tyr. In some embodiments such a hydroxyl group may be reacted with a charge-modifying moiety which comprises a hydroxyl-reactive group. Examples of hydroxyl -reactive groups include carboxyl groups which can react with hydroxyl groups to form esters; isocyanateswhich react with hydroxyl groups to form urethanes; and vinyl sulfones. A hydroxyl group may be activated for further reaction, e.g. by using a compound comprising a carboxyl, isocyanate or vinyl sulphone group coupled to a click chemistry group such as an azide or alkyne group, with the hydroxyl group being activated, followed by reaction of the free click chemistry group; for example with an azide or alkyne group. Analogous strategies can be used to those described above in the context of activating amine groups.
[0207] In some embodiments the side chain of the amino acid to be modified comprises a thiol group. Examples of such amino acids include Cys. In some embodiments such a thiol group may be reacted with a charge-modifying moiety which comprises a thiolreactive group. Examples of thiol -reactive groups include maleimides, haloacetamides, pyridyl disulfides and vinyl sulfones. In one embodiment a thiol group may be activated for further reaction, e.g. using a compound comprising a maleimide group, a haloacetamide or a pyridyl disulfides or vinyl sulfone group coupled to a click chemistry group such as an azide or alkyne group, with the thiol group being activated, followed by reaction of the free click chemistry group; for example with an azide or alkyne group of the linker. Analogous strategies can be used to those described above in the context of activating amine groups.
[0208] In some embodiments the target polypeptide is modified to prevent cross-reaction of functional groups in the target polypeptide apart from the target amino acid(s) the charge of which is altered in the disclosed methods. For example, in some embodiments the N-terminal amine group of the target polypeptide is modified, e.g. by being capped. In some embodiments the N-terminal amine group of the target polypeptide is modified by being acetylated. In some embodiments the C-terminal amine group of the target polypeptide is modified, e.g. by being capped. In some embodiments the C-terminal amine group of the target polypeptide is modified by being amidated. Capping of the N- and / or C -terminals of the target polypeptide can prevent reaction of such groups, which can be useful e.g in embodiments where it may be desirable for the N- and / or C-terminals of the target polypeptide to remain unmodified so that they can be modified with one or more sequencing adapters as described herein.
[0209] In some embodiments, modifying the side chains of said one or more amino acids with one or more charge-modifying moieties comprises contacting said one or more amino acids with one or more amino-acid modifying enzymes.
[0210] Any suitable such enzymes can be used. As those skilled in the art will appreciate, choice of a suitable amino-acid modifying enzyme is an operational parameter of thedisclosed methods which can be made by the user of the method according to the polypeptide to be processed.
[0211] In some embodiments an amino-acid modifying enzyme is selective for a recognition sequence in the target polypeptide. Choice of a suitable enzyme will include consideration of any recognition sequence required. In brief, however, and by way of nonlimiting illustration, the statistical distribution of a given recognition site in a polypeptide comprising a random or pseudo-random arrangement of amino acids (e.g. in a full length protein or other polypeptide as described herein) can be calculated based on the length and sequence of the recognition sequence and the composition of the polypeptide. This means that a skilled person can choose appropriate amino-acid modifying enzymes to use according to the frequency of the amino acid that it is desired to modify.
[0212] Suitable amino-acid modifying enzymes include kinases, sulfotransferases, O-GlcNAc transferases (OGTs), methyltransferases, peroxidases, palmitoyltransferases (PATs), nitrosylases, acetyltransferases, hydroxylases, deiminases, succinyltransferases, glutamylases, oligosaccharyltransferases (OSTs), glycosyltransferases, and glutaminase.
[0213] In some embodiments modifying the side chains of said one or more amino acids with one or more charge-modifying moieties comprises conducting a reaction as set out in the table in Figure 1.
[0214] In some embodiments an amino acid comprising an amine group in the side chain (e.g. a lysine amino acid) may be modified by acetylation, alkylation (e.g. methylation), ubiquitination, or hydroxylation. Suitable reagents include acetyl anhydride, alkyl halides (e.g. methyl iodide), glyoxal, aldehydes (e.g. formaldehyde). Suitable enzymes include methyltransferases, acetyltransferases, hydroxylases, and the like.
[0215] In some particular embodiments an amino acid comprising a hydroxyl group in the side chain (e.g. Ser, Thr, Tyr) may be modified by phosphorylation, O-GlcNAcylation, sulfation, alkylation (e.g. methylation), or nitration. Suitable reagents include phosphoric acid derivatives (e.g. ATP), glycosyl donors, oxidising agents (e.g. peroxides such as hydrogen peroxide, hypocholorites such as sodium hypochlorite, permanganates such as potassium permanganate), halogen sources (such as Ch and h), nitrating agents (such as peroxynitrites and nitryl chloride), and Fenton reagents. Suitable enzymes include kinases, sulfotransferases, O-GlcNAc transferases (OGTs), methyltransferases, peroxidases, and the like.
[0216] In some particular embodiments an amino acid comprising a carboxyl group in the side chain (e.g. Asp, Glu) may be modified by succinylation or alkylation (e.g.methylation). Suitable reagents include succinic anhydride, alkyl halides such as methyl iodide, and metal ion salts. Suitable enzymes include methyltransferases and succinyltransferases.
[0217] In some particular embodiments an amino acid comprising an imidazole group in the side chain (e.g. His) may be modified by phosphorylation or alkylation (e.g. methylation). Suitable reagents include nitrating agents and transition metals (e.g. copper and iron salts). Suitable enzymes include methyltransferases and kinases.
[0218] In some particular embodiments an amino acid comprising an guanidine group in the side chain (e.g. Arg) may be modified by alkylation (e.g. methylation) or citrullination. Suitable reagents include cyanates, nitric oxide donors (e.g. NO2, ONO?' etc) and alkyl halides such as methyl iodide. Suitable enzymes include methyltransferases, nitrosylases and deiminases.
[0219] In some particular embodiments an amino acid comprising a thiol group in the side chain (e.g. Cys) may be modified by S-palmitoylation, S-nitrosylation, disulphide bond formation, or sulfydration. Suitable reagents include alkyl halides (e.g. iodoacetamide) NO donors, oxidising agents (e.g. peroxides such as hydrogen peroxide, 02, etc), and sulphide donors such as NaHS. Suitable enzymes include palmitoyltransferases (PATs), nitrosylases, peroxidases, and the like.
[0220] In some particular embodiments an amino acid comprising a secondary amide group in the side chain (e.g. Asn, Gin) may be modified by N-linked glycosylation, ADP-ribosylation, or deamidation. Suitable reagents include aldehydes (e.g. formaldehyde), glycation agents such as glucose and fructose, NAD+, and elevated heat and pH conditions (e.g. treatment with alkali such as sodium hydroxide). Suitable enzymes include glycosyltransferases, glutaminase, ribotransferases, etc.
[0221] Construct
[0222] In some embodiments of the provided methods, the target polypeptide is comprised in a construct. In some embodiments the construct comprises the target polypeptide and one or more of a sequencing adapter, a linker and a membrane anchor.
[0223] In some embodiments the target polypeptide is comprised in a construct comprising a linker. In some embodiments the construct comprises a plurality of polypeptides attached together via one or more linkers.When the target polypeptide is comprised in a construct comprising a linker, then any suitable linker can be used.
[0224] Generally speaking, in some embodiments a linker may comprise a multifunctional molecule. As used herein, a multi-functional molecule is a molecule having at least two reactive functional groups and is thus capable of binding to at least two polypeptides. In some embodiments the multi-functional molecule has a first reactive group capable of attaching to a first amino acid in a first polypeptide and a second reactive group capable of attaching to a second amino acid in a second polypeptide.
[0225] A linker may comprise a linear or branched, unsubstituted or substituted alkylene, alkenylene, alkynylene, arylene, heteroarylene, carbocyclylene or heterocyclylene moiety. A linker may comprise an unsubstituted or substituted alkylene, alkenylene, or alkynylene moiety. A linker may comprise an unsubstituted or substituted alkylene or alkenylene moiety. A linker may comprise an unsubstituted or substituted alkylene moiety.
[0226] Typically, an alkylene group is a Ci-io alkylene group. Typically, an alkenylene group is a C2-10 alkenylene group. Typically, an alkynylene group is a C2-10 alkynylene group.
[0227] Typically, an arylene group is a Ce-12 arylene group. Typically, a heteroarylene group is a 5- to 12- membered heteroarylene group. Typically, a carbocyclylene group is a C5-12 carbocyclylene group. Typically, a heterocyclylene group is a 5- to 12- membered heterocyclylene group.
[0228] An alkylene, alkenylene, or alkynylene moiety may be uninterrupted or interrupted by or terminate in one or more atoms or groups selected from maleimide, O, N(R), S, C(O), C(O)NR, C(O)O, phosphate, thiophosphate, unsubstituted or substituted arylene, unsubstituted or substituted heteroarylene, unsubstituted or substituted carbocyclylene and unsubstituted or substituted heterocyclylene; wherein R is selected from H, unsubstituted or substituted alkyl, and unsubstituted or substituted aryl.
[0229] In some embodiments a multifunctional molecule which can be used as a linker in the disclosed methods may comprise a linear or branched, unsubstituted or substituted alkylene, alkenylene, alkynylene, arylene, heteroarylene, carbocyclylene or heterocyclylene moiety comprising at least two reactive groups. In some embodiments a first reactive group may be an amine reactive group capable of reacting with the N-terminal amine group of a first polypeptide. In some embodiments a second reactive group may be a carboxyl reactive group capable of reacting with the C-terminal carboxyl group of a second polypeptide. In some embodiments the first and second reactive groups areeach independently selected from maleimide groups, carboxyl groups, amine groups, hydroxy groups, azide groups and alkyne groups
[0230] In some embodiments a linker may be comprised from a multifunctional molecule such as glutaraldehyde, disuccinimidyl suberate (DSS), bismaleimidoethane (BMOE), sulfo-SMCC (sulfosuccinimidyl 4-(N-maleimidomethyl)cyclohexane- 1 -carboxylate), sulfo-EMCS (N-s-maleimidocaproyl-oxysulfosuccinimide ester), EDC (l-ethyl-3-carbodiimide), DMP (Dess-Martin periodinane; 1, 1,1 -Tris(acetyloxy)- 1,1 -dihydro- 1,2-benziodoxol-3-(lH)-one), bis(sulfosuccinimidyl) suberate (BS3), DTME (dithiobismaleimidoethane), SPDP (succinimidyl 3-(2-pyridyldithio)propionate), BMH (bismaleimidohexane), sulfo-GMBS (N-y-maleimidobutyryl-oxysulfosuccinimide ester), MPBH (4-(4-N-maleimidophenyl)butyric acid hydrazide), NHS-PC (e.g., PC Biotin-PEG-NHS ester, e.g. CAS number 2353409-93-3), and NHS-PEG-maleimide (e.g. CAS number 1325208-25-0).
[0231] In some embodiments, a linker is or comprises a polymer. In some embodiments, a linker is or comprises a biopolymer. Suitable polymers include polynucleotides, polypeptides, polysaccharides, poly(ethylene glycols) and the like. In some embodiments the or each linker independently comprises a polynucleotide, a polypeptide and / or a polysaccharide. In some embodiments the or each linker independently comprises a polynucleotide, and / or a polypeptide. In some embodiments the or each linker independently comprises or consists of a polynucleotide. In some embodiments the or each linker independently comprises or consists of a polypeptide.
[0232] In some embodiments a linker comprises a charged polymer. In some embodiments a linker comprises a polynucleotide, or a charged polypeptide. In some embodiments a linker comprises a polynucleotide or a polypeptide comprising polyglutamic acid, polyaspartic acid, polyarginine, and / or polylysine.
[0233] In some embodiments a linker comprises a synthetic polymer. In some embodiments a linker comprises or consists of a polymer such as polyethylene glycol (PEG), poly(vinyl alcohol) (PVA), poly(ethylene oxide) (PEO), poly(acrylic acid) (PAA), polypropylene glycol) (PPG), poly(caprolactone) (PCL), polydimethylsiloxane (PDMS), poly(methacrylate) and derivatives thereof, polyurethane, poly(2-oxazoline), poly(N-isopropylacrylamide) (PNIPAM), and / or polyethyleneimine (PEI). In some embodiments a linker comprises a dendrimer.
[0234] A linker for use in the disclosed methods can be chosen or selected according to the applications for which the construct is intended.For example, in some embodiments the construct is to be characterised by taking one or more measurements characteristic of the construct as the construct moves with respect to a nanopore. In some embodiments the construct moves with respect to the nanopore under the control of a motor protein. In some embodiments the motor protein is a polynucleotide motor protein (also known as a polynucleotide-handling protein) as described in more detail herein. In such embodiments, the linker may comprise or consist of a polynucleotide. For example, this may be useful as a polynucleotide-containing linker may provide a loading site or binding site for the polynucleotide motor protein.
[0235] In some embodiments the motor protein is a polypeptide motor protein. In some such embodiments, the linker may comprise or consist of a polypeptide. For example, this may be useful as a polypeptide-containing linker may provide a loading site or binding site for the polypeptide motor protein.
[0236] When the linker is a polymer, the length of the linker is not particularly limited and can be chosen or selected as required by the operator of the disclosed methods. The choice of linker is an operational parameter within the control of the skilled user.
[0237] In some embodiments, however, the linker is a polymer (e.g. a polynucleotide and / or a polypeptide) comprising from about 2 to about 1000 monomer units, such as from about 5 to about 500 monomer units, e.g. from about 10 to about 100 monomer units, e.g. from about 10 to about 50 monomer units, e.g. about 10, about 15, about 20, about 30, about 40 or about 50 monomer units. In some embodiments the linker comprises or consists of a polynucleotide comprising from about 2 to about 1000 nucleotides, such as from about 5 to about 500 nucleotides, e.g. from about 10 to about 100 nucleotides, e.g. from about 10 to about 50 nucleotides, e.g. about 10, about 15, about 20, about 30, about 40 or about 50 nucleotides. In some embodiments the linker comprises or consists of a polypeptide comprising from about 2 to about 1000 amino acids (optionally including amino acid analogs), such as from about 5 to about 500 amino acids, e.g. from about 10 to about 100 amino acids, e.g. from about 10 to about 50 amino acids, e.g. about 10, about 15, about 20, about 30, about 40 or about 50 amino acids.
[0238] In some embodiments the linker comprises one or more cleavable moieties. In some embodiments the linker comprises one or more photocleavable moieties, In some embodiments the linker comprises one or more enzyme-cleavable moieties, such as one or more protease recognition sites (e.g. when the linker comprises a polypeptide) and / or one or more restriction sites (e.g. when the linker comprises a polynucleotide). Protease recognition sequences and restriction enzyme binding sequences are well known in the artand are described in standard reference texts and in resources such as the Alphabetized List of Recognition Sequences accessible at https: / / www.neb.com / en-gb / tools-and-resources / selection-charts / alphabetized-list-of-recognition-specificities and the MEROPS database of proteolytic enzymes (Rawlings et al, Nucleic Acids Research, 46, DI, 2018, D624-632), each incorporated by reference.
[0239] Sequencing adapter
[0240] In some embodiments, the disclosed methods comprise attaching a sequencing adapter to the polypeptide in a construct as described in more detail herein.
[0241] In some embodiments the attachment between the target polypeptide (or a construct comprising the target polypeptide) and a sequencing adapter is covalent. In some embodiments the attachment is non-covalent. In some embodiments the attachment comprises ligating the sequencing adapter to the polypeptide.
[0242] As will be apparent from the discussion herein, the attachment between a sequencing adapter and the target polypeptide or construct is not especially limited and any suitable attachment means can be used.
[0243] In some embodiments, a sequencing adapter may be attached at the N-terminal of the target polypeptide or construct. In some embodiments a sequencing adapter may be attached at the C-terminal of the target polypeptide or construct. In some embodiments, sequencing adapters may be attached at the N-terminal and the C-terminal of the target polypeptide or construct.
[0244] Any suitable sequencing adapter may be used. For example, if the construct is for characterisation using a nanopore, the choice of the sequencing adapters may be dependent on the nanopore characterisation methods that are envisaged.
[0245] Typically, if sequencing adapters are attached at the N-terminal and the C-terminal of the target polypeptide or construct, the sequencing adapter attached at the N-terminal is different to the sequencing adapter attached at the C-terminal of the construct.
[0246] In some embodiments a sequencing adapter comprises a polynucleotide and / or a polypeptide.
[0247] In one embodiment, the or each adapter is synthetic or artificial. Typically, the or each adapter comprises a polymer as described herein. In some embodiments, the or each adapter comprises a spacer as described herein. In some embodiments, the or each adapter comprises a polynucleotide. The or each polynucleotide adapter may comprise DNA, RNA, modified DNA (such as abasic DNA), RNA, PNA, LNA, BNA and / or PEG.Usually, the or each adapter comprises single stranded and / or double stranded DNA or RNA.
[0248] In some embodiments, an adapter is a linear adapter. A linear adapter may be bound to either or both ends of a oligopeptide.
[0249] A linear adapter may comprise a leader sequence as described herein. A linear adapter may comprise a portion for hybridisation with a tag (such as a pore tag) as described herein. A linear adapter may be 10 to 150 nucleotides in length, such as from 20 to 120, e.g. 30 to 100, for example 40 to 80 such as 50 to 70 nucleotides in length. A linear adapter may be single stranded. A linear adapter may be double stranded.
[0250] In some embodiments, an adapter may be a Y adapter. A Y adapter is typically a polynucleotide adapter. A Y adapter is typically double stranded and comprises (a) at one end, a region where the two strands are hybridised together and (b), at the other end, a region where the two strands are not complementary. The non-complementary parts of the strands typically form overhangs. The presence of a non-complementary region in the Y adapter gives the adapter its Y shape since the two strands typically do not hybridise to each other unlike the double stranded portion. The two single-stranded portions of the Y adapter may be the same length, or may be different lengths. For example, one singlestranded portion of the Y adapter may be 10 to 150 nucleotides in length, such as from 20 to 120, e.g. 30 to 100, for example 40 to 80 such as 50 to 70 nucleotides in length and the other single stranded portion of the Y adapter may independently by 10 to 150 nucleotides in length, such as from 20 to 120, e.g. 30 to 100, for example 40 to 80 such as 50 to 70 nucleotides in length. The double-stranded “stem” portion of the Y adapter may be e.g. from 10 to 150 nucleotides in length, such as from 20 to 120, e.g. 30 to 100, for example 40 to 80 such as 50 to 70 nucleotides in length. A Y adapter may be attached to either or both ends of a barcode as described herein.
[0251] An adapter may be linked to the target polypeptide or construct by any suitable means known in the art, according to the present disclosure. The adapter may be synthesized separately and chemically attached or enzymatically ligated to the strand as described herein.
[0252] An adapter suitable for use in the described methods may in some embodiments comprise a leader. A leader may be useful to assist the capture of the adapter and thus of the barcode by a nanopore as described herein.
[0253] In some embodiments the leader may be from about 10 to 150 nucleotides (e.g. DNA and / or RNA nucleotides) in length, such as from 20 to 120, e.g. 30 to 100, forexample 40 to 80 such as 50 to 70 nucleotides in length, or from about 10 to about 60 nucleotides in length, e.g. from about 20 to about 50, such as from about 20 to about 40, e.g. about 30 nucleotides in length.
[0254] In some embodiments the leader is a charged polymer, e.g. a negatively charged polymer. In some embodiments the leader comprises a polymer such as PEG or a polysaccharide. In such embodiments the leader may be from 10 to 150 monomer units (e.g. ethylene glycol or saccharide units) in length, such as from 20 to 120, e.g. 30 to 100, for example 40 to 80 such as 50 to 70 monomer units (e.g. ethylene glycol or saccharide units) in length.
[0255] An adapter may comprise a polypeptide. An adapter may comprise a leader comprising a polypeptide. The polypeptide may in some embodiments have a net negative charge or it may have a specific recognition sequence, e.g. a sequence specific for a motor protein such as an unfoldase enzyme. In such embodiments the leader may be from 10 to 150 amino acids in length, such as from 20 to 120, e.g. 30 to 100, for example 40 to 80 such as 50 to 70 amino acids in length.
[0256] Motor proteins
[0257] Aspects of the disclosed methods comprise contacting the polypeptide once destabilized as described herein with a motor protein under conditions such that the motor protein controls the movement of the polypeptide with respect to the nanopore.
[0258] The movement of the polypeptide (or a construct comprising the polypeptide) with respect to the nanopore may be driven by any suitable means. In some embodiments, the movement of the polypeptide (or a construct comprising the polypeptide) is driven by a physical or chemical force (potential). In some embodiments the physical force is provided by an electrical (e.g. voltage) potential or a temperature gradient, etc.
[0259] In some embodiments, the movement of the polypeptide (or a construct comprising the polypeptide) comprises mechanically manipulating the polypeptide (or a construct comprising the polypeptide) thereby moving said polypeptide or construct with respect to the nanopore. In some embodiments, movement of the polypeptide or construct by mechanical manipulation does not comprise using a polynucleotide-handling protein.
[0260] In some embodiments the polypeptide (or a construct comprising the polypeptide) is moved by mechanical manipulation in a direction opposite to a potential applied across said nanopore. In some embodiments, the potential is a voltage potential applied across said nanopore. In some embodiments, the construct is moved with respect to the nanoporeas described in WO 2020 / 128517, the entire contents of which are hereby incorporated by reference, particularly in regards to discussion in that document of movements of polynucleotides with respect to nanoreactors.
[0261] In some embodiments, the polypeptide (or a construct comprising the polypeptide) moves with respect to the nanopore as an electrical potential is applied across the nanopore. In some embodiments the polypeptide or construct is charged (e.g. negatively charged), and so applying a voltage potential across a nanopore will cause the construct to move with respect to the nanopore under the influence of the applied voltage potential. For example, if a positive voltage potential is applied to the trans side of the nanopore relative to the cis side of the nanopore, then this will induce a negatively charged polypeptide or construct to move from the cis side of the nanopore to the trans side of the nanopore.
[0262] Similarly, if a positive voltage potential is applied to the trans side of the nanopore relative to the cis side of the nanopore then this will impede the movement of a negatively charged polypeptide or construct from the trans side of the nanopore to the cis side of the nanopore. The opposite will occur if a negative voltage potential is applied to the trans side of the nanopore relative to the cis side of the nanopore. Apparatuses and methods of applying appropriate voltages are described in more detail herein.
[0263] In some embodiments the chemical force is provided by a concentration (e.g. pH) gradient.
[0264] In some embodiments the movement of the polypeptide or construct with respect to the nanopore is controlled using a method as described in WO 2020 / 016573, the entire contents of which are incorporated herein by reference.
[0265] In some embodiments the movement of the polypeptide or construct is controlled using a method as disclosed in any of WO 2021 / 111125, WO 2021 / 133168, or PCT / GB2023 / 052838, the entire contents of which are incorporated herein by reference.
[0266] In some embodiments the movement of the polypeptide or construct with respect to the nanopore is controlled using a motor protein.
[0267] In some embodiments a motor protein is present (e.g. prior to the contact of the polypeptide or construct with the nanopore) on a sequencing adapter comprised in or attached to the polypeptide or construct, as described in more detail herein. In some embodiments a motor protein is present on the polypeptide, or on a polypeptide portion of a construct comprising the polypeptide. In some embodiments a motor protein is present on a linker portion of a construct comprising the polypeptide.In some embodiments a motor protein controls the movement of the polypeptide or construct in the same direction as the physical or chemical force (potential). For example, in some embodiments a positive voltage potential is applied to the trans side of the nanopore relative to the cis side of the nanopore, and a motor protein controls the movement of the polypeptide or construct from the cis side of the nanopore to the trans side of the nanopore. In some embodiments a positive voltage potential is applied to the cis side of the nanopore relative to the trans side of the nanopore, and a motor protein controls the movement of the polypeptide or construct from the trans side of the nanopore to the cis side of the nanopore.
[0268] In some embodiments a motor protein controls the movement of the polypeptide or construct in the opposite direction to the physical or chemical force (potential). For example, in some embodiments a positive voltage potential is applied to the trans side of the nanopore relative to the cis side of the nanopore, and the motor protein controls the movement of the polypeptide or construct from the trans side of the nanopore to the cis side of the nanopore. In some embodiments a positive voltage potential is applied to the cis side of the nanopore relative to the trans side of the nanopore, and the motor protein controls the movement of the polypeptide or construct from the cis side of the nanopore to the trans side of the nanopore.
[0269] In some embodiments the movement of the polypeptide or construct is driven by the motor protein in the absence of an applied potential.
[0270] In embodiments of the disclosed methods which comprise the use of a motor protein, the motor protein is typically capable of controlling the movement of the polypeptide or construct with respect to a nanopore. In other words, the motor protein is capable of controlling the movement of the polypeptide or construct .
[0271] Suitable motor proteins are in some embodiments also known as polynucleotide-handling proteins or polynucleotide-handling enzymes, or polypeptide-handling proteins or polypeptide-handling enzymes. Suitable proteins are known in the art and some exemplary motor proteins are described in more detail below.
[0272] In one embodiment, a motor protein is or is derived from a polynucleotide handling enzyme. A polynucleotide handling enzyme is a polypeptide that is capable of interacting with and modifying at least one property of a polynucleotide. The enzyme may modify the polynucleotide by cleaving it to form individual nucleotides or shorter chains of nucleotides, such as di- or trinucleotides. The enzyme may modify the polynucleotide by orienting it or moving it to a specific position.In some embodiments, a motor protein can be present on a construct or an adapter attached thereto prior to its contact with a nanopore. For example, a motor protein can be present on a polynucleotide portion of an adapter.
[0273] In one embodiment, the motor protein is derived from a member of any of the Enzyme Classification (EC) groups 3.1.11, 3.1.13, 3.1.14, 3.1.15, 3.1.16, 3.1.21, 3.1.22, 3.1.25, 3.1.26, 3.1.27, 3.1.30, 3.1.31 and 3.4.21.
[0274] In some embodiments of the claimed methods, the motor protein is a helicase, a polymerase, an exonuclease, a topoisomerase, or a variant thereof.
[0275] In some embodiments the motor protein is designed, configured or selected to prevent the motor protein disengaging from the polypeptide or a construct comprising the polypeptide (other than by passing off the end of the polypeptide or construct). Such modified polynucleotide-handling proteins are particularly suitable for use in the disclosed methods. This is particularly useful in the methods disclosed herein which comprise characterising the polypeptide or construct using a nanopore. Thus, in some embodiments of such methods, the polypeptide or construct does not disengage from the motor protein.
[0276] As used herein, the term “disengaging” refers to the dissociation of the motor protein from the polypeptide or construct. Thus, a motor protein may be modified to prevent it from dissociating from the polypeptide or construct, e.g. into the reaction medium. It is important to distinguish potential “disengagement” of a motor protein from “unbinding” of a motor protein from a polypeptide or construct. As used herein, “unbinding” refers to the transient release of the polypeptide or construct from the active site of the motor protein but does not imply disengagement. Thus, for example, a motor protein may be modified to prevent the motor protein from disengaging from a polypeptide or construct, but without preventing the motor protein from unbinding from the polypeptide or construct. When unbound, the motor protein remains engaged with the polypeptide or construct. For example, the motor protein may remain engaged with the polypeptide or construct (i.e. it may be prevented from disengaging from the polypeptide or construct) because it is topologically closed around the polypeptide or construct. The polynucleotide binding site may remain free to bind or unbind the polypeptide or construct such that the motor protein may bind or unbind to the polypeptide or construct, whilst the motor protein remains engaged with the polypeptide or construct. When the motor protein is unbound from the polypeptide or construct it may be able to move on (e.g., along) the polypeptide or construct under an applied force and may be capable of re-binding to thepolypeptide or construct. When engaged on the polypeptide or construct but unbound from the polypeptide or construct, the motor protein is not capable of dissociating from the polypeptide or construct.
[0277] The motor protein can be adapted to prevent disengagement in any suitable way. For example, the motor protein can be loaded on the polypeptide or construct and then modified in order to prevent it from disengaging from the polypeptide or construct.
[0278] Alternatively, the motor protein can be modified to prevent it from disengaging from the polypeptide or construct before it is loaded onto the polypeptide or construct. Modification of a motor protein in order to prevent it from disengaging from a polypeptide or construct can be achieved using methods known in the art, such as those discussed in WO 2014 / 013260 and WO 2021 / 255476, each of which is hereby incorporated by reference in its entirety, and with particular reference to passages describing the modification of motor proteins such as helicases in order to prevent them from disengaging from constructs, conjugates and polynucleotide strands. For example, a motor protein can be modified by treating with tetramethylazodicarboxamide (TMAD) or various other closing moieties.
[0279] When a polynucleotide motor protein is used, it may have a polynucleotide-unbinding opening; e.g. a cavity, cleft or void through which a polynucleotide strand may pass when the motor protein disengages from the strand. In some embodiments, the polynucleotide-unbinding opening is the opening through which a polynucleotide may pass when the motor protein disengages from the polynucleotide. In some embodiments, the polynucleotide-unbinding opening for a given motor protein can be determined by reference to its structure, e.g. by reference to its X-ray crystal structure. The X-ray crystal structure may be obtained in the presence and / or the absence of a polynucleotide substrate. In some embodiments, the location of a polynucleotide-unbinding opening in a given motor protein may be deduced or confirmed by molecular modelling using standard packages known in the art. In some embodiments, the polynucleotide-unbinding opening may be transiently produced by movement of one or more parts e.g. one or more domains of the motor protein.
[0280] The motor protein may be modified by closing the polynucleotide-unbinding opening. The polynucleotide-unbinding opening may be closed with a closing moiety. Closing the polynucleotide-unbinding opening may therefore prevent the motor protein from disengaging from the construct. For example, the motor protein may be modified by covalently closing the polynucleotide-unbinding opening. However, as explained above closing the polynucleotide-unbinding opening does not necessarily prevent the constructfrom unbinding from the polynucleotide binding site of the motor protein. Accordingly, in some embodiments of the disclosed methods, the motor protein is modified to wholly or partially close an opening existing in at least one conformation state of the unmodified protein through which a polynucleotide or polypeptide strand can unbind. In some embodiments, a preferred protein for addressing in this way is a helicase.
[0281] In embodiments in which the motor protein is a polypeptide motor protein, the motor protein may similarly have a polypeptide-unbinding opening. This can be similarly closed to prevent disengagement of the motor protein from the construct in the same way as for a polynucleotide motor protein.
[0282] In one embodiment, the motor protein is an exonuclease. Suitable enzymes include, but are not limited to, exonuclease I from E. coli. exonuclease III enzyme from E. coli, Red from T. thermophilus and bacteriophage lambda exonuclease, TatD exonuclease and variants thereof.
[0283] In one embodiment, the motor protein is a polymerase. The polymerase may be PyroPhage® 3173 DNA Polymerase (which is commercially available from Lucigen® Corporation), SD Polymerase (commercially available from Bioron®), Klenow from NEB or variants thereof. In one embodiment, the enzyme is Phi29 DNA polymerase or a variant thereof. Modified versions of Phi29 polymerase that may be used in the disclosed methods are disclosed in US Patent No. 5,576,204.
[0284] In some embodiments the motor protein is a polymerase, e.g. a polymerase as described herein.
[0285] In one embodiment the motor protein is a topoisomerase. In one embodiment, the topoisomerase is a member of any of the Moiety Classification (EC) groups 5.99.1.2 and 5.99.1.3. The topoisomerase may be a reverse transcriptase, which are enzymes capable of catalysing the formation of cDNA from a RNA template. They are commercially available from, for instance, New England Biolabs® and Invitrogen®.
[0286] In one embodiment the motor protein is a translocase. Examples include translocases in the FtsK and SpoIII families.
[0287] In one embodiment, the motor protein is a helicase. Any suitable helicase can be used in accordance with the methods provided herein. For example, the or each motor protein used in accordance with the present disclosure may be independently selected from a Hel308 helicase, a RecD helicase, a Tral helicase, a TrwC helicase, an XPD helicase, and a Dda helicase, or a variant thereof. Monomeric helicases may comprise several domains attached together. For instance, Tral helicases and Tral subgroup helicases may containtwo RecD helicase domains, a relaxase domain and a C-terminal domain. The domains typically form a monomeric helicase that is capable of functioning without forming oligomers. Particular examples of suitable helicases include Hel308, NS3, Dda, UvrD, Rep, PcrA, Pifl and Tral. These helicases typically work on single stranded DNA.
[0288] Examples of helicases that can move along both strands of a double stranded DNA include FtsK and hexameric enzyme complexes, or multisubunit complexes such as RecBCD, and are particularly suited to some embodiments disclosed herein. NS3 helicases are particularly suitable for use in the disclosed methods as they are capable of processing both DNA and RNA and so can be used in embodiments of the disclosed methods in which the target double stranded nucleic acid is a DNA-RNA hybrid.
[0289] Hel308 helicases are described in publications such as WO 2013 / 057495, the entire contents of which are incorporated by reference. RecD helicases are described in publications such as WO 2013 / 098562, the entire contents of which are incorporated by reference. XPD helicases are described in publications such as WO 2013 / 098561, the entire contents of which are incorporated by reference. Dda helicases are described in publications such as WO 2015 / 055981 and WO 2016 / 055777, the entire contents of each of which are incorporated by reference.
[0290] In some embodiments a motor protein (e.g. a helicase) can control the movement of a strand in at least two active modes of operation (when the motor protein is provided with all the necessary components to facilitate movement, e.g. fuel and cofactors such as ATP and Mg2+discussed herein) and one inactive mode of operation (when the motor protein is not provided with the necessary components to facilitate movement).
[0291] When provided with all the necessary components to facilitate movement (i.e. in the active modes), the motor protein (e.g. helicase) moves along the construct in a 5’ to 3’ or a 3’ to 5’ direction (depending on the motor protein). The motor protein can be used to either move the construct away from (e.g. out of) the pore (e.g. against an applied force) or the strand towards (e.g. into) the pore (e.g. with an applied force). For example, when the end of the construct towards which the motor protein moves is captured by a pore, the motor protein works against the direction of the force and pulls the threaded construct out of the pore (e.g. into the cis chamber). However, when the end away from which the motor protein moves is captured in the pore, the motor protein works with the direction of the force and pushes the threaded construct into the pore (e.g. into the trans chamber).
[0292] When the motor protein (e.g. helicase) is not provided with the necessary components to facilitate movement (i.e. in the inactive mode) it can bind to the constructand act as a brake slowing the movement of the construct when it is moved with respect to a nanopore, e.g. by being pulled into the pore by a force. In the inactive mode, it does not matter which end of the construct is captured, it is the applied force which determines the movement with respect to the pore, and the motor protein acts as a brake. When in the inactive mode, the movement control by the motor protein can be described in a number of ways including ratcheting, sliding and braking.
[0293] In another embodiment the motor protein is a protein translocase. Protein translocases are protein-binding polypeptides which are able to control movement of a protein substrate, for example an enzyme, enzyme complex, or a part of an enzyme complex that operates on a protein substrate and moves it relative to the enzyme in a processive manner, i.e. as a function of enzymatic activity.
[0294] In some embodiments the motor protein is a NTP driven unfoldase. NTP driven unfoldases are NTP-dependent enzymes that catalyze protein unfolding. NTP driven unfoldases include ATP-dependent proteases, such as proteasomal ATPases, AAA proteases, AAA+ enzymes; membrane fusion proteins, such as NSF (N-Ethylmal eimidesensitive fusion protein) / Sacl8p (N-Ethylmaleimide-sensitive fusion protein homologue in yeast) or p97 / VCP / Cdc48p (97-kDa valosin-containing protein); Pexlp and Pex6p (peroxisomal ATPase); Katanin and SKD1 (Vps4p homolog in mouse) / Vps4p (Vacuolar protein sorting 4 homolog in yeast); Dynein (motor protein); DNA replication proteins, such as ORC (origin recognition complex), Cdc6 (cell division control protein 6), MCM (minichromosome maintenance protein), DnaA, or RFC (replication factor C) / clamp-loader; RuvB (holliday junction ATP-dependent DNA helicase RuvB, EC=3.6.4.12); TIP49a / TIP49 and TIP49b / TIP48 (eukaryotic RuvB-like protein).
[0295] In some embodiments the motor protein is an AAA+ enzyme, AAA+ enzymes are members of the AAA+ superfamily of enzymes. AAA+ is an abbreviation for ATPases Associated with diverse cellular Activities. They share a common conserved module of approximately 230 amino acid residues. This is a large, functionally diverse protein family belonging to the AAA+ superfamily of ring-shaped P-loop NTPases, which exert their activity through the energy-dependent remodeling or translocation of macromolecules. Examples include ClpAP, ClpXP, ClpCP, HslYU and Lon in bacteria and their homologues in mitochondria and chloroplasts. With the exception of Lon, AAA+ enzymes (sometimes referred to as unfoldases or proteases) consist of regulatory (ATPase) and proteolytic subunits, while Lon is a single polypeptide containing both regulatory and proteolytic domains. ClpX and ClpA dock with ClpP to form ClpXP and ClpAP proteases,whereas HslU docks with HslY to form another protease, HslVU. ClpA and ClpX form hexamers, in contrast to ClpP which forms heptamers. HslU and HslY each form hexamers, although HslU heptamers have also been reported. The regulatory subunits ClpA, ClpX and HslU function as chaperones.
[0296] AAA+ enzymes may also be referred to as AAA+ molecular motors.
[0297] HslU is a member of the HsplOO and Clp family of ATPase. It can also form complex with HslY to act as an unfoldase.
[0298] Lon proteases are ATP-dependent serine peptidases belonging to the MEROPS peptidase family S16 (Ion protease family, clan SF).
[0299] In some embodiments the motor protein is ClpX or is a derivative thereof. ClpX is a member of the HSP (heat-shock protein) 100 family having the Uniprot designation clpX and having the 424 amino acid sequence given there, processed into mature form, as a subunit. ClpX subunits associate to form a six-membered (homohexameric) ring that is stabilized by binding of ATP or nonhydrolysable analogs of ATP. The N-terminal domain of ClpX is a C4-type zinc binding domain (ZBD) involved in substrate recognition. ZBD forms a very stable dimer that is essential for promoting the degradation of some typical ClpXP substrates such as and Mu A.
[0300] In some embodiments the motor protein is E. coli ClpX. E. coli ClpX generates sufficient mechanical force (>20 pN) to denature stable protein folds, and translocates along proteins at a suitable rate for primary sequence analysis by nanopore sensors (up to 80 amino acids per second). ClpX is part of the ClpXP proteasome-like complex. ClpP is composed of a diheptameric cylinder-like protease that binds at one or both ends a regulatory hexameric ATP-dependent unfoldase / translocase complex (e.g. ClpX). ClpX acts as a gate that allows for tagged proteins to enter into the inner lumen of the ClpP protease complex for subsequent degradation. The ATP-dependent unfoldase / translocase activity of the hexameric protein complex, ClpX, is employed to unfold and thread proteins through a nanopore.
[0301] In some embodiments the motor protein is a ClpX-deltaN subunits, lacking N-terminal amino acids 1-60, linked with a 20 amino acid long linker and prepared as a single polypeptide chain.
[0302] In some embodiments the motor protein is a Clp / HsplOO ATPase. Clp / HsplOO ATPases are responsible for selecting protein targets. For example, the two different bacterial ATPases ClpX and ClpA impart distinct substrate preferences to the ClpP peptidase.In some embodiments the motor protein is a mitochondrial protein translocase. Examples include TOM or TIM from human or eukaryotic cells, such as TOMM20 (translocase of outer mitochondrial membrane homolog), TOMM22 (mitochondrial import receptor subunit 22 homolog), TOMM40 (translocase of outer mitochondrial membrane 40 homolog), TOM7 (translocase of mitochondrial outer membrane 7), T0MM7 (translocase of outer mitochondrial membrane 7 homolog), TIMM8A (translocase of inner mitochondrial membrane 8 homolog A), TIMM50 (translocase of inner mitochondrial membrane 50 homolog).
[0303] Another alternative protein translocase may be prepared from the Sec family of translocases. These include SecB (chaperone protein), SecA (ATPase), SecY (internal membrane complex in prokaryotes), SecE (interal membrane complex in prokaryotes), SecG (internal membrane complex in prokaryotes) or Sec61 (internal membrane complex in eukaryotes), SecD (membrane protein), and SecF (membrane protein).
[0304] Another alternative protein translocase is Type III Secretion System (TTS) Translocase, such as HrcN and any of the subunits of the TTS translocases, or Secindependent periplasmic protein translocase TatC.
[0305] Examples of suitable protein translocases, such as NTP driven unfoldases as described above, are described in WO 2013 / 123379, hereby incorporated by reference.
[0306] A motor protein typically requires fuel in order to handle the processing of polynucleotides and / or polypeptides. Fuel is typically free nucleotides or free nucleotide analogues. The free nucleotides may be one or more of, but are not limited to, adenosine monophosphate (AMP), adenosine diphosphate (ADP), adenosine triphosphate (ATP), guanosine monophosphate (GMP), guanosine diphosphate (GDP), guanosine triphosphate (GTP), thymidine monophosphate (TMP), thymidine diphosphate (TDP), thymidine triphosphate (TTP), uridine monophosphate (UMP), uridine diphosphate (UDP), uridine triphosphate (UTP), cytidine monophosphate (CMP), cytidine diphosphate (CDP), cytidine triphosphate (CTP), cyclic adenosine monophosphate (cAMP), cyclic guanosine monophosphate (cGMP), deoxyadenosine monophosphate (dAMP), deoxyadenosine diphosphate (dADP), deoxyadenosine triphosphate (dATP), deoxyguanosine monophosphate (dGMP), deoxyguanosine diphosphate (dGDP), deoxyguanosine triphosphate (dGTP), deoxythymidine monophosphate (dTMP), deoxythymidine diphosphate (dTDP), deoxythymidine triphosphate (dTTP), deoxyuridine monophosphate (dUMP), deoxyuridine diphosphate (dUDP), deoxyuridine triphosphate (dUTP), deoxy cytidine monophosphate (dCMP), deoxy cytidine diphosphate (dCDP) anddeoxycytidine triphosphate (dCTP). The free nucleotides are usually selected from AMP, TMP, GMP, CMP, UMP, dAMP, dTMP, dGMP or dCMP. The free nucleotides are typically adenosine triphosphate (ATP).
[0307] A cofactor for the motor protein is a factor that allows the motor protein to function. The cofactor is often a divalent metal cation. The divalent metal cation is often Mg2+, Mn2+, Ca2+or Co2+. The cofactor is most typically Mg2+.
[0308] Detector
[0309] Embodiments described herein refer to movement of an analyte such as a target polypeptide or construct comprising a target polypeptide as described herein with respect to a nanopore. However, whilst the disclosure provides nanopores as exemplary detectors, the methods provided herein are also amenable to other detectors including (i) a zero-mode waveguide, (ii) a field-effect transistor, optionally a nanowire field-effect transistor; (iii) an AFM tip; (iv) a nanotube, optionally a carbon nanotube and (v) a nanopore. The disclosed methods are particularly amenable to methods in which a polypeptide is moved through a detector or through a structure containing a detector, e.g. a well in a detector chip.
[0310] Nanopore
[0311] In the disclosed methods, any suitable nanopore can be used. In one embodiment a nanopore is a transmembrane pore.
[0312] A transmembrane pore is a structure that crosses the membrane to some degree. It permits hydrated ions driven by an applied potential to flow across or within the membrane. The transmembrane pore typically crosses the entire membrane so that hydrated ions may flow from one side of the membrane to the other side of the membrane. However, the transmembrane pore does not have to cross the membrane. It may be closed at one end. For instance, the pore may be a well, gap, channel, trench or slit in the membrane along which or into which hydrated ions may flow.
[0313] Any suitable transmembrane pore may be used in the methods provided herein. The pore may be biological or artificial. Suitable pores include, but are not limited to, protein pores, polynucleotide pores, and solid state pores.
[0314] A solid state pore may, in one embodiment, comprise a nanochannel. In some embodiments the solid state pore is a pore disclosed in WO 2003 / 003446, WO 2009 / 020682 or WO 2016 / 187519, each of which is incorporated by reference in their entirety.In one embodiment, the pore may be a DNA origami pore (Langecker et aL, Science, 2012; 338: 932-936). Suitable DNA origami pores are disclosed in WO2013 / 083983, WO 2018 / 011603 and WO 2020 / 025974, each of which is incorporated by reference in their entirety.
[0315] In one embodiment, the nanopore is a scaffolded polypeptide nanopore. In some embodiments the pore is a scaffolded polypeptide nanopore as disclosed in WO 2020 / 025909 or WO 2020 / 074399, each of which is incorporated by reference in their entirety.
[0316] In one embodiment, the nanopore is a transmembrane protein pore. A transmembrane protein pore is a polypeptide or a collection of polypeptides that permits hydrated ions, such as polynucleotides, to flow from one side of a membrane to the other side of the membrane. In the methods provided herein, the transmembrane protein pore is capable of forming a pore that permits hydrated ions driven by an applied potential to flow from one side of the membrane to the other. The transmembrane protein pore typically permits polynucleotides and polypeptides to flow from one side of the membrane, such as a polymer membrane, to the other. The transmembrane protein pore allows a polynucleotide or polypeptide to be moved through the pore.
[0317] In one embodiment, the nanopore is a transmembrane protein pore which is a monomer or an oligomer. The pore is typically made up of several repeating subunits, such as at least 6, at least 7, at least 8, at least 9, at least 10, at least 11, at least 12, at least 13, at least 14, at least 15, or at least 16 subunits. The pore is typically a hexameric, heptameric, octameric or nonameric pore. The pore may be a homo-oligomer or a heterooligomer.
[0318] In one embodiment, the transmembrane protein pore comprises a barrel or channel through which the ions may flow. The subunits of the pore typically surround a central axis and contribute strands to a transmembrane P-barrel or channel or a transmembrane a-helix bundle or channel.
[0319] Typically, the barrel or channel of the transmembrane protein pore comprises amino acids that facilitate interaction with an analyte, such as a target polypeptide (as described herein). These amino acids are typically located near a constriction of the barrel or channel. The transmembrane protein pore typically comprises one or more positively charged amino acids, such as arginine, lysine or histidine, or aromatic amino acids, such as tyrosine or tryptophan. These amino acids typically facilitate the interaction between the pore and nucleotides, polynucleotides, nucleic acids and polypeptides.The transmembrane protein pore may be from or derived from Wza, Iota toxin, Anthrax protective antigen, Vibrio cholerae cytolysin, Cytotoxin K (CytK), CELIII, CsgG, CsgF, CsgG-CsgF, Aerolysin, alpha hemolysin, MspA, MspB, MspC, PorARr, PorBRr, PorARc, PilQ, necrotic enteritis B-like toxin (NetB), FraC, portal proteins including G20c, P23_45, T4, SPP1, P22 and Phi29, gamma hemolysin, Monalysin, Lysenin, ClyA, an actinoporin, Clostridium perfringens beta toxin, parasporin-2, epsilon toxin, lectin from the parasitic mushroom Laetiporus sulphureus (LSL), volvatoxin, Cry toxins, Cytl Aa, Cyt2Aa, Complement component 9 (C9), Perfringolysin O, Pleurotolysin, Listeriolysin, Perforin-2, Gasdermin- A3, L-, P- and M-ring protein, Type II secretion system protein D, GspD, InvG, VirB7, SpoIIIAG, Cag8, Cag3, Cag or other proteins in the Type IV secretion system apparatus protein CagY, WzzB, Pentraxin, Afp2, Major vault protein, Thioredoxindependent peroxidase reductase, Arf-GAP, Respiratory syncytial virus ribonucleoprotein, Chikungunya virus nonstructural protein 1, PRC, YaxA, XaxA, HfaB, NfpAB, leukocidin and PrgH.
[0320] Suitable transmembrane protein pores for use in the invention include those described in WO 2016 / 034591, WO 2017 / 149316, WO 2017 / 149317, WO 2017 / 149318, WO 2018 / 211241, WO 2019 / 002893, WO 2023 / 118404, WO 2023 / 198911, WO 2024 / 033421, WO 2024 / 033422, WO 2024 / 033443, and WO 2024 / 089270 (all incorporated by reference herein in their entirety).
[0321] The transmembrane protein pore may also be any of the CsgG pores described in WO 2023 / 060420, WO 2023 / 60418, WO 2023 / 60422, WO 2023 / 060421, WO 2023 / 019470, CN114957412, WO 2023 / 019471, W02023 / 060419 and WO 2023 / 050031 (all incorporated herein by reference in their entireties) or a variant thereof.
[0322] The transmembrane protein pore may be any of the pores described in WO 2023 / 123370, WO 2024 / 138470, WO 2024 / 138472, WO 2024 / 138424, WO 2024 / 138425, WO 2024 / 138512 and WO 2024 / 138565 or a variant thereof.
[0323] The transmembrane pore may be formed from a chimeric pore monomer comprising two or more regions, wherein at least two of the two or more regions are from at least two different pores. The chimeric pore monomer may comprise any number of regions, such as three or more, four or more, five or more, six or more, seven or more, eight or more, nine or more or ten or more regions, from different pores. The chimeric pore monomer may comprise two or three regions. The regions are preferably selected from a cap region, a constriction region, and a transmembrane region. The regions may be a cap region and a constriction region. The regions may be a cap region, a constriction region,and a transmembrane region. The at least two different pores are typically at least two different pores that appear in nature. The at least two different pores are typically at least two different wild-type or naturally occurring pores. The at least two different pores are preferably different before any artificial or synthetic modifications, such as additions, deletions and / or substitutions, are made to them. The at least two different pores are preferably homologues, for example structural homologues. A structural homologue refers to a protein or molecule that shares a similar three-dimensional structure with another protein or molecule. This can be determined using standard methods in the art (e.g., AlphaFold or PSIPRED). Structural homologues typically have similar sequences.
[0324] Structural homologues are normally identified in similar species. The at least two different pores may be selected from any of the pores listed above. The at least two different pores may be two different PorARc pores or three different PorARc pores. The at least two different pores may be two different CsgG pores or three different CsgG pores. The chimeric pore monomer may be any of those described in PCT / EP2023 / 080135 (incorporated by reference herein in its entirety).
[0325] The transmembrane pore may be formed from a pore monomer comprising (a) a CsgG monomer and (b) a fusion polypeptide comprising a first portion comprising a CsgF peptide and a second portion comprising a helix-forming auxiliary protein, wherein the fusion protein is attached to the pore monomer. The pore monomer may be derived from a protein transmembrane pore complex comprising (a) a CsgG transmembrane pore comprising a lumen and (b) a fusion polypeptide comprising a first portion comprising a CsgF protein and a second portion comprising a helix-forming auxiliary protein, wherein the fusion protein is attached to the transmembrane pore. The auxiliary protein can be designed de novo using computer-based structural analysis tools to confer certain desirable features to the CsgG monomer (e.g., modulation of pore width, lengthening of pore lumen, formation of one or more additional constrictions, etc.). The de novo designed auxiliary protein may form one or more additional constrictions in the lumen of a CsgG pore formed from the monomer, and improve discrimination of polymer units as an analyte moves through the pore. The pore monomer may be any of the pore monomers described in WO 2024 / 033447 (incorporated by reference herein in its entirety).
[0326] TagsIn some embodiments of the methods provided herein, a tag on the nanopore can be used, e.g. to promote the capture of the polypeptide or a construct comprising the polypeptide.
[0327] The interaction between a tag on a nanopore and a binding site on the polypeptide or a construct comprising the polypeptide may be reversible. A strong non-covalent bond (e.g., biotin / avidin) is still reversible and can be useful in some embodiments of the methods described herein. For example, a pair of pore tag and binding site (e.g. comprised in a polypeptide, construct or adaptor) can be designed to provide a sufficient interaction with the nanopore such that the polypeptide or construct is held close to the nanopore (without detaching from the nanopore and diffusing away) but is able to release from the nanopore as it is processed.
[0328] A pore tag and adaptor can be configured such that the binding strength or affinity of a binding site on a construct comprising the polypeptide (e.g., a binding site provided by an anchor or a leader sequence of an adaptor or by a capture sequence within the duplex stem of an adaptor) to a tag on a nanopore is sufficient to maintain the coupling between the nanopore and construct until an applied force is placed on it to release the bound construct from the nanopore.
[0329] One or more molecules that attract or bind the polypeptide, construct or an adapter attached thereto may be linked to the nanopore. Any molecule that hybridizes to the polypeptide, construct and / or adaptor may be used. The molecule attached to the pore may be selected from a PNA tag, a PEG linker, a short oligonucleotide, a positively charged amino acid and an aptamer. Pores having such molecules linked to them are known in the art. For example, pores having short oligonucleotides attached thereto are disclosed in Howarka et al (2001) Nature Biotech. 19: 636-639 and WO 2010 / 086620, and pores comprising PEG attached within the lumen of the pore are disclosed in Howarka et al (2000) J. Am. Chem. Soc. 122(11): 2411-2416. In some embodiments, a tag or tether may be uncharged. This can ensure that the tags or tethers are not drawn into the nanopore under the influence of a potential difference if present.
[0330] A short oligonucleotide attached to the nanopore, which comprises a sequence complementary to a sequence in a construct comprising the polypeptide (e.g. in a leader sequence or another single stranded sequence in an adaptor) may be used to enhance capture of the polypeptide, construct or adapter attached thereto in the methods described herein.Anchors
[0331] In some embodiments of the methods provided herein, an anchor on the polypeptide or a construct comprising the polypeptide can be used, e.g. to promote the localisation of the polypeptide or construct to a membrane in which a nanopore may be present.
[0332] Thus, in one embodiment, a polypeptide, construct or adapter attached thereto may comprise a membrane anchor or a transmembrane pore anchor. The anchor may be a polypeptide anchor and / or a hydrophobic anchor that can be inserted into the membrane. In one embodiment, the hydrophobic anchor is a lipid, fatty acid, sterol, carbon nanotube, polypeptide, protein or amino acid, for example cholesterol, palmitate or tocopherol. The anchor may comprise thiol, biotin or a surfactant. In one aspect the anchor may be biotin (for binding to streptavidin), amylose (for binding to maltose binding protein or a fusion protein), Ni-NTA (for binding to poly-histidine or poly-histidine tagged proteins) or peptides (such as an antigen).
[0333] In one embodiment, the anchor is or comprises cholesterol or a fatty acyl chain. For example, any fatty acyl chain having a length of from 6 to 30 carbon atom, such as hexadecanoic acid, may be used. Examples of suitable anchors and methods of attaching anchors to adapters are disclosed in WO 2012 / 164270 and WO 2015 / 150786.
[0334] Membrane
[0335] The detector or nanopore is typically present in a membrane. Any suitable membrane may be used.
[0336] The membrane is preferably an amphiphilic layer. An amphiphilic layer is a layer formed from amphiphilic molecules, such as phospholipids, which have both hydrophilic and lipophilic properties. The amphiphilic molecules may be synthetic or naturally occurring. Non-naturally occurring amphiphiles and amphiphiles which form a monolayer are known in the art and include, for example, block copolymers (Gonzalez-Perez et al., Langmuir, 2009, 25, 10447-10450). Block copolymers are polymeric materials in which two or more monomer sub-units that are polymerized together to create a single polymer chain. Block copolymers typically have properties that are contributed by each monomer sub-unit. However, a block copolymer may have unique properties that polymers formed from the individual sub-units do not possess. Block copolymers can be engineered such that one of the monomer sub-units is hydrophobic (z.e., lipophilic), whilst the other subunits) are hydrophilic whilst in aqueous media. In this case, the block copolymer maypossess amphiphilic properties and may form a structure that mimics a biological membrane. The block copolymer may be a diblock (consisting of two monomer sub-units) but may also be constructed from more than two monomer sub-units to form more complex arrangements that behave as amphipiles. The copolymer may be a triblock, tetrablock or pentablock copolymer. The membrane may be a triblock copolymer membrane.
[0337] Archaebacterial bipolar tetraether lipids are naturally occurring lipids that are constructed such that the lipid forms a monolayer membrane. These lipids are generally found in extremophiles that survive in harsh biological environments, thermophiles, halophiles and acidophiles. Their stability is believed to derive from the fused nature of the final bilayer. It is straightforward to construct block copolymer materials that mimic these biological entities by creating a triblock polymer that has the general motif hydrophilic-hydrophobic-hydrophilic. This material may form monomeric membranes that behave similarly to lipid bilayers and encompass a range of phase behaviours from vesicles through to laminar membranes. Membranes formed from these triblock copolymers hold several advantages over biological lipid membranes. Because the triblock copolymer is synthesised, the exact construction can be carefully controlled to provide the correct chain lengths and properties required to form membranes and to interact with pores and other proteins.
[0338] Block copolymers may also be constructed from sub-units that are not classed as lipid sub-materials; for example, a hydrophobic polymer may be made from siloxane or other non-hydrocarbon-based monomers. The hydrophilic sub-section of block copolymer can also possess low protein binding properties, which allows the creation of a membrane that is highly resistant when exposed to raw biological samples. This head group unit may also be derived from non-classical lipid head-groups.
[0339] Triblock copolymer membranes also have increased mechanical and environmental stability compared with biological lipid membranes, for example a much higher operational temperature or pH range. The synthetic nature of the block copolymers provides a platform to customise polymer-based membranes for a wide range of applications.
[0340] The membrane may comprise just one type of amphiphile or may comprise more than one type of amphiphile. The membrane may comprise a mixture of naturally occurring amphiphilic molecules such as a mixture of phospholipids. The membrane may comprise a mixture of non-naturally occurring amphiphiles such as a mixture of block copolymers. The membrane may comprise a mixture of naturally occurring amphiphilicmolecules such as a phospholipids and non-naturally occurring amphiphiles such as block copolymers. For example, in some embodiments the membrane comprises from about 1% to about 99% (e.g. wt% or mol%) of naturally occurring amphiphilic molecules such as one or more phospholipids and from about 1% to about 99% (e.g. wt% or mol%) of non-naturally occurring amphiphiles such as one or more block copolymers (wherein the sum of naturally occurring amphiphilic molecules such as a phospholipids and non-naturally occurring amphiphiles such as block copolymers does not exceed 100 % (e.g. wt% or mol%)).
[0341] The membrane may be one of the membranes disclosed in International Application No. WO2014 / 064443 or WO2014 / 064444 (both of which are incorporated herein by reference in their entireties).
[0342] The amphiphilic molecules may be chemically modified or functionalised to facilitate coupling of the polynucleotide. The amphiphilic layer may be a monolayer or a bilayer. The amphiphilic layer is typically planar. The amphiphilic layer may be curved. The amphiphilic layer may be supported.
[0343] Amphiphilic membranes are typically naturally mobile, essentially acting as two-dimensional fluids with lipid diffusion rates of approximately 10'8cm s'1. This means that the pore and coupled polynucleotide can typically move within an amphiphilic membrane.
[0344] The membrane may be a lipid bilayer. Lipid bilayers are models of cell membranes and serve as excellent platforms for a range of experimental studies. For example, lipid bilayers can be used for in vitro investigation of membrane proteins by single-channel recording. Alternatively, lipid bilayers can be used as biosensors to detect the presence of a range of substances. The lipid bilayer may be any lipid bilayer. Suitable lipid bilayers include, but are not limited to, a planar lipid bilayer, a supported bilayer, or a liposome. The lipid bilayer is preferably a planar lipid bilayer. Suitable lipid bilayers are disclosed in WO 2008 / 102121, WO 2009 / 077734, and WO 2006 / 100484 (incorporated herein by reference in their entireties).
[0345] Methods for forming lipid bilayers are known in the art. Lipid bilayers are commonly formed by the method of Montal and Mueller (Proc. Natl. Acad. Sci. USA., 1972; 69: 3561-3566).
[0346] A lipid bilayer may be formed as described in WO 2009 / 077734 (incorporated herein by reference in its entirety). In this method, the lipid bilayer is formed from dried lipids. A lipid bilayer may be formed across an opening as described in W02009 / 077734.The membrane may comprise a solid-state layer. Solid state layers can be formed from both organic and inorganic materials including, but not limited to, microelectronic materials, insulating materials such as SislS , AI2O3, and SiO, organic and inorganic polymers such as polyamide, plastics such as Teflon® or elastomers such as two-component addition-cure silicone rubber, and glasses. The solid-state layer may be formed from graphene. Suitable graphene layers are disclosed in WO 2009 / 035647 (incorporated herein by reference in its entirety). If the membrane comprises a solid-state layer, the pore is typically present in an amphiphilic membrane or layer contained within the solid-state layer, for instance within a hole, well, gap, channel, trench or slit within the solid-state layer. The skilled person can prepare suitable solid state / amphiphilic hybrid systems.
[0347] Suitable systems are disclosed in WO 2009 / 020682 and WO 2012 / 005857 (incorporated herein by reference in their entireties). Any of the amphiphilic membranes or layers discussed above may be used.
[0348] The methods disclosed herein are typically carried out using (i) an artificial amphiphilic layer comprising a pore, (ii) an isolated, naturally occurring lipid bilayer comprising a pore, or (iii) a cell having a pore inserted therein. The methods are typically carried out using an artificial amphiphilic layer, such as an artificial triblock copolymer layer. The layer may comprise other transmembrane and / or intramembrane proteins as well as other molecules in addition to the pore. Suitable apparatus and conditions are discussed below. The method of the invention is typically carried out in vitro.
[0349] Characterisation
[0350] As discussed in more detail herein, in some embodiments the disclosed methods comprise taking one or more measurements as the destabilized polypeptide generated by modifying the target polypeptide as described herein, or a construct comprising said polypeptide, moves with respect to a nanopore.
[0351] The one or more measurements are typically one or more measurements characteristic of the destabilized polypeptide. By taking one or more measurements characteristic of the destabilized polypeptide, characteristics of the target polypeptide can be determined.
[0352] Many different characteristics can be determined. Any suitable measurements can be taken.
[0353] For example, in some embodiments characterising the destabilized polypeptide comprises determining (i) the length of the destabilized polypeptide, (ii) the identity of thedestabilized polypeptide, (iii) the sequence of the destabilized polypeptide, (iv) the secondary structure of the destabilized polypeptide; (v) whether or not and / or to the extent to which the destabilized polypeptide is modified; and / or (vi) the presence, absence, concentration or relative abundance of the destabilized polypeptide in a sample comprising multiple polypeptides.
[0354] Because the destabilized polypeptide construct is derived from the target polypeptide, the characteristics of the destabilized polypeptide can inform on the characteristics of the target polypeptide. Thus, the disclosed methods can be used to inform regarding (i) the length of the target polypeptide, (ii) the identity of the target polypeptide, (iii) the sequence of the target polypeptide, (iv) the secondary structure of the target polypeptide; (v) whether or not and / or to the extent to which the target polypeptide is modified, e.g. by one or more post-translational modifications.; (vi) the presence, absence, concentration or relative abundance of the target polypeptide in a sample comprising multiple polypeptides.
[0355] Those skilled in the art will appreciate that characterising a polypeptide does not necessarily comprise determining any or all of these features. Many characterisation measurements can be made as a polypeptide moves with respect to a nanopore.
[0356] In some embodiments the measurements are characteristic of the sequence of the destabilized polypeptide. In some embodiments the measurements are characteristic of the sequence of the target polypeptide.
[0357] Conditions
[0358] The disclosed methods may be carried out using any apparatus that is suitable for investigating a membrane / pore system in which a nanopore is inserted into a membrane. The methods may be carried out using any apparatus that is suitable for transmembrane pore sensing. For example, the apparatus may comprise a chamber comprising an aqueous solution and a barrier that separates the chamber into two sections. The barrier may have an aperture in which a membrane containing a transmembrane pore is formed.
[0359] Transmembrane pores are described herein.
[0360] The methods may be carried out using the apparatus described in WO 2008 / 102120, WO 2010 / 122293 or WO 00 / 28312.
[0361] The methods may comprise optical measurements, for example such as described in WO 2016 / 009180 and WO 2021 / 198695.The methods may involve measuring the ion current flow through the pore, typically by measurement of a current. Alternatively, the ion flow through the pore may be measured optically, such as disclosed by Heron et al: J. Am. Chem. Soc. 9 Vol. 131, No. 5, 2009. Therefore the apparatus may also comprise an electrical circuit capable of applying a potential and measuring an electrical signal across the membrane and pore. The characterisation methods may be carried out using a patch clamp or a voltage clamp. The characterisation methods typically involve the use of a voltage clamp.
[0362] The methods may be carried out on a silicon-based array of wells where each array comprises 128, 256, 512, 1024, 2000, 3000, 4000, 6000, 10000, 12000, 15000 or more wells.
[0363] The methods may involve the measuring of a current flowing through the pore. The method is typically carried out with a voltage applied across the membrane and pore. The voltage used is typically from +2 V to -2 V, typically -400 mV to +400mV. The voltage used is typically in a range having a lower limit selected from -400 mV, -300 mV, -200 mV, -150 mV, -100 mV, -50 mV, -20mV and 0 mV and an upper limit independently selected from +10 mV, + 20 mV, +50 mV, +100 mV, +150 mV, +200 mV, +300 mV and +400 mV. The voltage used is more typically in the range 100 mV to 240mV and most typically in the range of 120 mV to 220 mV. It is possible to increase discrimination between different nucleotides by a pore by using an increased applied potential.
[0364] The methods are typically carried out in the presence of any charge carriers, such as metal salts, for example alkali metal salts, halide salts, for example chloride salts, such as alkali metal chloride salt. Charge carriers may include ionic liquids or organic salts, for example tetramethyl ammonium chloride, trimethylphenyl ammonium chloride, phenyltrimethyl ammonium chloride, or l-ethyl-3 -methyl imidazolium chloride. In the exemplary apparatus discussed above, the salt is present in the aqueous solution in the chamber. Potassium chloride (KC1), sodium chloride (NaCl) or caesium chloride (CsCl) is typically used. KC1 is typical. The salt may be an alkaline earth metal salt such as calcium chloride (CaCh). The salt concentration may be at saturation. The salt concentration may be 3M or lower and is typically from 0.1 to 2.5 M, from 0.3 to 1.9 M, from 0.5 to 1.8 M, from 0.7 to 1.7 M, from 0.9 to 1.6 M or from 1 M to 1.4 M. The salt concentration is typically from 150 mM to 1 M. The characterisation method may be carried out using a salt concentration of at least 0.3 M, such as at least 0.4 M, at least 0.5 M, at least 0.6 M, at least 0.8 M, at least 1.0 M, at least 1.5 M, at least 2.0 M, at least 2.5 M or at least 3.0 M. High salt concentrations provide a high signal to noise ratio and allow for currentsindicative of binding / no binding to be identified against the background of normal current fluctuations.
[0365] The methods are typically carried out in the presence of a buffer. In the exemplary apparatus discussed above, the buffer is present in the aqueous solution in the chamber. Any suitable buffer may be used. Typically, the buffer is HEPES. Another suitable buffer is Tris-HCl buffer. The methods are typically carried out at a pH of from 4.0 to 12.0, from 4.5 to 10.0, from 5.0 to 9.0, from 5.5 to 8.8, from 6.0 to 8.7 or from 7.0 to 8.8 or 7.5 to 8.5. The pH used may be about 7.5.
[0366] The methods may be carried out at from 0 °C to 100 °C, from 15 °C to 95 °C, from 16 °C to 90 °C, from 17 °C to 85 °C, from 18 °C to 80 °C, 19 °C to 70 °C, or from 20 °C to 60 °C. The methods are typically carried out at room temperature. The methods are optionally carried out at a temperature that supports enzyme function, such as about 37 °C.
[0367] Further aspects
[0368] Also provided herein is a method of characterising a target polypeptide, the method comprising
[0369] destabilizing the secondary and / or tertiary structure of the target polypeptide by disrupting the charge of one or more amino acids in the target polypeptide; wherein disrupting the charge of one or more amino acids in the target polypeptide comprises modifying the side chains of said one or more amino acids with one or more chargemodifying moieties;
[0370] contacting the destabilized polypeptide with a motor protein;
[0371] contacting the destabilized polypeptide with a nanopore; and
[0372] taking one or more measurements characteristic of the destabilized polypeptide as the motor protein controls the movement of the destabilized polypeptide with respect to the nanopore; thereby characterising the target polypeptide.
[0373] In such methods, the target polypeptide is typically a target polypeptide as defined herein. Destabilizing the secondary and / or tertiary structure of the target polypeptide is typically carried out as defined herein. The destabilized polypeptide may be moved with respect to a nanopore which may be a nanopore as described herein using methods as described herein, for example under the control of a motor protein as described herein.
[0374] The measurements characteristic of the polypeptide are typically as described herein. In some embodiments the one or more measurements are one or more measurements characteristic of the destabilized polypeptide. In some embodiments the oneor more measurements are one or more measurements characteristic of the target polypeptide from which the destabilized polypeptide is derived. In some embodiments the one or more measurements are one or more measurements as described herein. In some embodiments the one or more measurements are one or more electrical or optical measurements. In some embodiments the one or more measurements are one or more electrical measurements. In some embodiments the one or more measurements are one or more optical measurements. In some embodiments the one or more measurements comprise measuring the ion current flow through the nanopore.
[0375] Also provided is a kit for characterising a target polypeptide, comprising a chemical or enzymatic reagent for disrupting the charge of one or more amino acids in a target polypeptide by modifying the side chains of said one or more amino acids;
[0376] and
[0377] a nanopore capable of detecting one or more characteristics of a target polypeptide as the target polypeptide moves with respect to the nanopore; and / or
[0378] a motor protein capable of controlling the movement of a target polypeptide with respect to a nanopore.
[0379] In some embodiments the kit further comprises a sequencing adapter capable of selectively reacting with the target polypeptide. In some embodiments the target polypeptide is provided in the form of a construct as described in more detail herein, and / or the kit includes components for the production of such a construct from the target polypeptide, such as one or more linker for attaching a plurality of polypeptides together. In some embodiments the target polypeptide is a target polypeptide as described herein. In some embodiments the reagent is as described herein. In some embodiments the nanopore is a nanopore as described herein. In some embodiments the motor protein is a motor protein as described herein.
[0380] The kit may comprise instructions for preparing a peptide linker construct from a target polypeptide as described herein.
[0381] Also provided is a construct comprising a linearized polypeptide having a length of at least 50 amino acids and comprising at least 5 charge-modified amino acids, and a motor protein, in some embodiments the construct comprises a linearized polypeptide as described herein; for example, the linearized polypeptide may be derived from a targetpolypeptide as described herein. In some embodiments the linearized polypeptide has a length of at least 50, at least 100, at least 150, at least 200, at least 300, at least 400, at least 500, at least 1000, at least 5,000, at least 10,000 or more peptide units. In some embodiments the construct comprises a plurality of said linearized polypeptides and one or more of a sequencing adaptor; a linker; and a membrane anchor. In some embodiments the linker is a linker as described herein. In some embodiments the sequencing adaptor is a sequencing adaptor as described herein. In some embodiments the membrane anchor is a membrane anchor as described herein.
[0382] Also provided is a system, comprising a library of said constructs, which may be the same or different, and a nanopore. In some embodiments the nanopore is as described herein. In some embodiments the system further comprises a motor protein capable of controlling the movement of the constructs in the library with respect to the nanopore. In some embodiments the system comprises computing means configured to detect information characteristic of the constructs in the library and to selectively process the signal obtained as said constructs move with respect to the nanopore. In some embodiments the system comprises receiving means for receiving data from detection of the oligopeptides, processing means for processing the signal obtained as the constructs move with respect to the nanopore, and output means for outputting the characterisation information thus obtained.
[0383] Also provided is the use of a charge-modifying reagent to facilitate translocation of a target polypeptide through a nanopore,
[0384] wherein said use comprises modifying the side chains of one or more amino acids of said target polypeptide with said charge-modifying reagent,
[0385] thereby disrupting the charge of said one or more amino acids in the target polypeptide and destabilizing the secondary and / or tertiary structure of the target polypeptide,
[0386] thereby facilitating the translocation of the target polypeptide through the nanopore. Also provided is the use of a charge-modifying reagent to facilitate threading of a target polynucleotide through a nanopore.
[0387] wherein said use comprises modifying the side chains of one or more amino acids of said target polypeptide with said charge-modifying reagent,thereby disrupting the charge of said one or more amino acids in the target polypeptide and destabilizing the secondary and / or tertiary structure of the target polypeptide,
[0388] thereby facilitating the threading of the target polypeptide through the nanopore. Also provided is the use of a charge-modifying reagent to improve discrimination of amino acids in a target polynucleotide as it moves through a nanopore,
[0389] wherein said use comprises modifying the side chains of one or more amino acids of said target polypeptide with said charge-modifying reagent,
[0390] thereby disrupting the charge of said one or more amino acids in the target polypeptide and destabilizing the secondary and / or tertiary structure of the target polypeptide,
[0391] thereby improving discrimination of amino acids in the target polynucleotide as it moves through the nanopore.
[0392] In some embodiments the charge-modifying reagent is a reagent as described herein. In some embodiments the target polypeptide is a target polypeptide as described herein. In some embodiments the charge of the side chains of the one or more amino acids is disrupted as described herein. In some embodiments the nanopore is a nanopore as described herein.
[0393] In some embodiments, said use comprises the use of a motor protein to control the movement of the target polypeptide through the nanopore. In some such embodiments, the motor protein is a motor protein as described herein.
[0394] Exemplary workflows
[0395] The following non-limiting workflow is provided to illustrate the methods described herein. A protein or fragment thereof is provided in a sample. The protein constitutes the target polypeptide. The target polypeptide comprises a plurality of amino acids. A target amino acid is chosen by the user of the method. For example, in some embodiments the target amino acid may be chosen by the user of the method based on the sequence of the protein. For example, in some embodiments the target amino acid is chosen from Lys, Arg, Glu and Asp. In some embodiments the target amino acid is Lys. The method may comprise attaching a linker and / or one or more sequencing adapters to the to target polypeptide thereby forming a construct as defined herein. The construct may be contacted with a chemical and / or enzymatic reagent which reacts with the target amino acids in the target polypeptide. The reaction of the reagent with the target amino acidsalters the charge of the target amino acids. The alteration of the charge of the target amino acids alters the structure of the native target polypeptide. The destabilized polypeptide can be translocated through a nanopore, for example under the control of a motor protein, and one or more measurements (e.g. ionic current measurements) can be taken during the translocation. The measurements are characteristic of the destabilized polypeptide and allow the properties of the target polypeptide to be determined
[0396] It is to be understood that although particular embodiments, specific configurations as well as materials and / or molecules, have been discussed herein for methods according to the present invention, various changes or modifications in form and detail may be made without departing from the scope and spirit of this invention. The preceding embodiments and subsequent examples are provided for illustration only, and should not be considered limiting the application. The application is limited only by the claims.
[0397] EXAMPLES
[0398] Example 1
[0399] This example demonstrates a method of moving a target polypeptide with respect to a nanopore comprising modifying the side chains of said one or more amino acids with one or more charge-modifying moieties; and contacting the destabilized polypeptide with a motor protein under conditions such that the motor protein controls the movement of the polypeptide with respect to the nanopore. In this example the motor protein is an unfoldase.
[0400] Preparation of analytes for electrophysiology
[0401] Polypeptide analytes referred to as KSI (SEQ ID NO: 1) and DHFR (SEQ ID NO: 2) were prepared. Each analyte was labelled at the ybbR tag to introduce a linking moiety to facilitate leader attachment as described herein; labelling was carried out using the method of Yin et al., Nat. Protoc. 1, 280-285 (2006). A peptide corresponding to SEQ ID NO: 3 was attached to analyte KSI, and an oligonucleotide corresponding to SEQ ID NO: 4 was attached to analyte DHFR.
[0402] Leaders were prepared by annealing oligonucleotides corresponding to SEQ ID NO: 5 and SEQ ID NO: 6. The leader was attached to a peptide containing an unfoldaserecognition sequence (SEQ ID NO: 7) using click chemistry, and subsequently purified by SPRI using AMPure XP beads.
[0403] The leader with attached unfoldase recognition sequence was then modified by DNA ligation using T4 DNA ligase to attach an oligonucleotide corresponding to SEQ ID 8. This was then attached to the linker-modified KSI using click chemistry. The conjugate was subsequently purified by SPRI using AMPure XP beads. The leader with attached unfoldase recognition sequence was attached to DHFR using click chemistry and subsequently purified by SPRI using AMPure XP beads.
[0404] Modification of the charge of analyte DHFR
[0405] DHFR analyte protein to be modified was diluted into a buffer containing 6M Guanidine hydrochloride and DMSO. The sample was reduced through addition of TCEP and incubated for 15 mins at room temperature. The sample was then subjected to heating at 95°C for 30 mins. The protein was then exchanged into pH 9.0 borate buffer and labelled with 2500 equivalents of 4-sulfophenyl isothiocyanate for 1 hour at 37°C.
[0406] Labelling in this manner modifies the sidechain of each lysine residue, removing the positive charge and introducing a negative charge. The protein was then exchanged into 25 mM HEPES-KOH pH 7.5, 50 mM NaCl.
[0407] Electrophysiology
[0408] Electrical data were collected using FLO-MIN004RA flow cells (Oxford Nanopore Technologies pic). Sequencing libraries consisting of 200 nM analyte and 40 nME coli ClpX unfoldase enzyme were prepared in 75 pL Sequencing Buffer (Oxford Nanopore Technologies pic). Flow cells were first flushed with 1 mL of Sequencing Buffer. Sample (75 pL) was then introduced into the flow cell via the SpotON port. Electrical data were acquired with a sample rate of 4 kHz at 30°C and applied potential of 140 mV.
[0409] Results and discussion
[0410] KSI and DHFR polypeptides with attached oligonucleotide leader and unfoldase recognition sequence are depicted schematically in Figure 2.
[0411] Electrophysiology measurements were first collected with KSI; a depiction of the analyte delivery is shown in Figure 3. An example of an obtained trace is shown in Figure 4, demonstrating successful controlled translocation of the analyte through the nanopore.Data collection was then attempted with unmodified DHFR. No signals could be collected, suggesting that the unmodified DHFR analyte could not be translocated through the nanopore.
[0412] A charge-modified version of DHFR was subsequently tested. In this version, the DHFR polypeptide had its negative charge increased - after ybbR tagging and prior to attachment of the recognition sequence-appended leader - by modification with 4-sulfophenyl isothiocyanate, as described above. Electrophysiology measurements were conducted and signals for translocation were collected. An example of an obtained trace is shown in Figure 5, demonstrating that modification of the analyte to increase its net negative charge led to successful controlled translocation through the nanopore.
[0413] The charge distribution for unmodified and modified DHFR was calculated. Shown in Figures 6A and 6B are plots representing charge distribution for unmodified and modified DHFR domains. A subsequence from SEQ ID 3 (85-242 inclusive) was taken, corresponding to the folded DHFR domain. A moving average of charge was calculated and plotted taking into account a 20 residue window for both unmodified and modified sequences. Y axis on the plot is the moving average charge, and the X axis is the position in the sequence. A horizontal dotted line is shown indicating neutral net charge. The process of charge modification increased the net negative charge, and reduced the positive charge; this is shown as a reduction in the total area under the curve between 1 and 0.
[0414] Example 2
[0415] This example demonstrates the modification of amino acid side chains in a target peptide with charge-modifying moieties and the characterisation of the modified target peptide using a nanopore. The modified peptide is conjugated to DNA and movement of the DNA -peptide conjugate is controlled using a DNA motor protein, using methods such as described in WO 2021 / 111125.
[0416] Modification of target peptides to introduce additional negative charge
[0417] Target peptides to be modified were diluted in DIPEA (N,N -diisopropylethylamine) solution containing 2000 equivalents of succinic anhydride. The sample was then heated at 37°C for 1 hour. The peptides were exchanged into a HEPES-NaCl buffer using GT- 100 spin columns (G-Bioscience). Succinic anhydride is capable of modifying serine, threonine, cysteine, tyrosine, histidine and lysine residues to introduce a succinylation modification, resulting in these residues having a charge of-1.The lysine residue in SEQ ID NO. 11 below is not modified as it is pre-modified with a methyltetrazine group.
[0418] Preparation of analytes for electrophysiology
[0419] Peptides corresponding to SEQ ID NO: 9, SEQ ID NO: 10, and SEQ ID NO: 11 were prepared. Each analyte was constructed using ‘click’ chemistry to attach the target peptide to DNA handles comprising a duplex of SEQ ID NO: 12 and SEQ ID NO: 13, and a duplex of SEQ ID NO: 14 and SEQ ID NO: 15. Analytes were then ligated to a nanopore sequencing adapter comprising a DNA motor protein (SQK-NBD114 kit; Oxford Nanopore Technologies pic).
[0420] A schematic of a peptide analyte with DNA handles and sequencing adapter attached is shown in Figure 7.
[0421] Electrophysiology
[0422] Electrical data were collected on a MinlON nanopore sequencing device (Oxford Nanopore Technologies pic) using custom flow cells. Sequencing Libraries from peptide analytes were prepared in Sequencing Buffer (Oxford Nanopore Technologies pic). Flow cells were first flushed with 1 mb of Flush Buffer. Sample (75 pL) was then introduced into the flow cell via the SpotON port. Electrical data were acquired with a sample rate of 1 kHz at 21 °C and applied potential of 200 mV.
[0423] Results and discussion
[0424] Electrophysiology measurements were first collected for control peptides lacking charge modification. Example traces are shown in Figure 8, demonstrating successful controlled translocation of the analyte through the nanopore.
[0425] Data collection was then performed on the modified peptides, in which additional negative charge had been introduced by modifying residues that were previously neutral or positively charged. Electrophysiology measurements were conducted and signals for translocations were collected. Example ionic current traces of these peptides are shown in Figure 9.
[0426] Figure 9 shows that modification of the target peptide to increase its net negative charge results in more discrete ionic current steps compared to the corresponding control, suggesting that the modified peptide is in a more extended or linear conformation compared to the control peptide lacking charge modification.SEQUENCE LISTING
[0427]
[0428]
Claims
CLAIMS1. A method of moving a target polypeptide with respect to a nanopore;the method comprising destabilizing the secondary and / or tertiary structure of the target polypeptide by disrupting the charge of one or more amino acids in the target polypeptide; wherein disrupting the charge of one or more amino acids in the target polypeptide comprises modifying the side chains of said one or more amino acids with one or more charge-modifying moieties; andcontacting the destabilized polypeptide with a motor protein under conditions such that the motor protein controls the movement of the polypeptide with respect to the nanopore.
2. A method according to claim 1, wherein disrupting the charge of one or more amino acids in the target polypeptide comprises modifying the side chains of a plurality of amino acids in the target polypeptide.
3. A method according to claim 2, wherein the amino acids in said plurality of amino acids may be the same or different.
4. A method according to any one of the preceding claims, wherein the one or more amino acids that are modified with one or more charge-modifying moieties comprise one or more charged amino acids.
5. A method according to any one of the preceding claims, wherein the one or more amino acids that are modified with one or more charge-modifying moieties comprise one or more positively charged amino acids and wherein disrupting the charge of said one or more amino acids comprises reducing the positive charge of said amino acids.
6. A method according to any one of the preceding claims, wherein the one or more amino acids that are modified with one or more charge-modifying moieties comprise one or more negatively charged amino acids and wherein disrupting the charge of said one or more amino acids comprises reducing the negative charge of said amino acids.
7. A method according to any one of the preceding claims, wherein the one or more amino acids that are modified with one or more charge-modifying moieties are selected from lysine, arginine, glutamate and aspartate.
8. A method according to any one of the preceding claims, wherein modifying the side chains of said one or more amino acids with one or more charge-modifying moieties comprises covalently attaching said one or more charge-modifying moieties to said side chains.
9. A method according to any one of the preceding claims, wherein modifying the side chains of said one or more amino acids with one or more charge-modifying moieties comprises one or more of oxidation, esterification, O-glycosylation, alkylation, nitration, nitrosylation, succinylation, hydroxylation, amidation, deamidation, acylation, sulfhydration, carb amyl ati on, glycation, metal coordination, Schiff-base formation, and ADP-ribosylation of said side chains.
10. A method according to any one of the preceding claims, wherein modifying the side chains of said one or more amino acids with one or more charge-modifying moieties comprises contacting said one or more amino acids with one or more amino-acid modifying enzymes.
11. A method according to claim 10, wherein said one or more amino-acid modifying enzymes are selected from kinases, sulfotransferases, O-GlcNAc transferases (OGTs), methyltransferases, peroxidases, palmitoyltransferases (PATs), nitrosylases, acetyltransferases, hydroxylases, deiminases, succinyltransferases, glutamylases, oligosaccharyltransferases (OSTs), glycosyltransferases, and glutaminase.
12. A method according to any one of the preceding claims, wherein modifying the side chains of said one or more amino acids with one or more charge-modifying moieties comprises contacting said one or more amino acids with one or more chemical reagents.
13. A method according to claim 12, wherein the one or more chemical reagents are selected from oxidising agents, halogenating agents, nitrating agents, alkylating agents, phosphoric acid derivatives, glycosyl donors, acids, bases, NO donors, sulphide donors,glycation agents, nucleotides, acetyl anhydride, glyoxal, formaldehyde, cyanate, metal ions, succinic anhydride, hydroxyl radicals, and fatty acids.
14. A method according to any one of the preceding claims, wherein the target polypeptide comprises a protein or fragment thereof.
15. A method according to any one of the preceding claims, wherein the target polypeptide has a length of at least 50 amino acids.
16. A method according to any one of the preceding claims, wherein the target polypeptide is comprised in a construct comprising said target polypeptide and one or more of a sequencing adaptor; a linker; and a membrane anchor.
17. A method according to any one of claims 1 to 15, comprising attaching the destabilized polypeptide to one or more of a sequencing adaptor; a linker; and a membrane anchor, thereby forming a construct.
18. A method according to any one of claims 16 to 17, wherein the construct comprises a plurality of polypeptides attached together via one or more linkers.
19. A method according to any one of claims 16 to 18, wherein the or each linker independently comprises a polynucleotide, a polypeptide and / or a polysaccharide.
20. A method according to any one of the preceding claims, comprising loading the motor protein onto the target polypeptide or onto a sequencing adapter or linker attached to the target polypeptide.
21. A method according to any one of the preceding claims, wherein the motor protein is a helicase.
22. A method according to any one of the preceding claims, wherein the motor protein is a NTP driven unfoldase.
23. A method according to any one of the preceding claims, wherein disrupting the charge of one or more amino acids in the target polypeptide wholly or partially linearizes the target polypeptide.
24. A method of characterising a target polypeptide, the method comprising destabilizing the secondary and / or tertiary structure of the target polypeptide by disrupting the charge of one or more amino acids in the target polypeptide; wherein disrupting the charge of one or more amino acids in the target polypeptide comprises modifying the side chains of said one or more amino acids with one or more chargemodifying moieties;contacting the destabilized polypeptide with a motor protein;contacting the destabilized polypeptide with a nanopore; andtaking one or more measurements characteristic of the destabilized polypeptide as the motor protein controls the movement of the destabilized polypeptide with respect to the nanopore; thereby characterising the target polypeptide.
25. A method according to claim 24, wherein:- the target polypeptide is as defined in any one of claims 14 to 19; and / or destabilizing the secondary and / or tertiary structure of the target polypeptide is carried out as defined in any one of claims 2 to 13 or 23; and / or- the motor protein is as defined in any one of claims 20 to 22.
26. A kit for characterising a target polypeptide, comprisinga chemical or enzymatic reagent for disrupting the charge of one or more amino acids in a target polypeptide by modifying the side chains of said one or more amino acids;anda nanopore capable of detecting one or more characteristics of a target polypeptide as the target polypeptide moves with respect to the nanopore; and / ora motor protein capable of controlling the movement of a target polypeptide with respect to a nanopore;and optionally comprising a sequencing adapter capable of selectively reacting with the target polypeptide.
27. A construct comprising a linearized polypeptide having a length of at least 50 amino acids and comprising at least 5 charge-modified amino acids, and a motor protein;optionally wherein the construct comprises a plurality of said linearized polypeptides and one or more of a sequencing adaptor; a linker; and a membrane anchor.
28. Use of a charge-modifying reagent to facilitate translocation of a target polypeptide through a nanopore,wherein said use comprises modifying the side chains of one or more amino acids of said target polypeptide with said charge-modifying reagent,thereby disrupting the charge of said one or more amino acids in the target polypeptide and destabilizing the secondary and / or tertiary structure of the target polypeptide,thereby facilitating the translocation of the target polypeptide through the nanopore.