Methods of using nanopores to characterize target polypeptides

By combining target peptides with polynucleotide conjugates and controlling the movement of the conjugates within nanopores using polynucleotide-treated proteins, the problem of peptide characterization in existing technologies has been solved, enabling rapid and accurate peptide characterization at the single-molecule level.

CN118011004BActive Publication Date: 2025-12-30OXFORD NANOPORE TECH LTD
View PDF 30 Cites 0 Cited by

Patent Information

Application Number
CN202410066918.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2020-09-30
Filing Date
2020-12-01
Publication Date
2025-12-30
Estimated Expiration
2040-12-01

AI Technical Summary

Technical Problem

Existing peptide characterization techniques are difficult to perform efficient and accurate characterization at the single-molecule level. In particular, mass spectrometry and Edman degradation methods are affected by contaminants, are costly, and are not suitable for characterizing differences within peptide sample groups.

Method used

By conjugating target peptides with polynucleotides to form peptide-polynucleotide conjugates, and by using polynucleotide-treated proteins to control the movement of the conjugates relative to nanopores, peptide-specific measurements are performed to characterize peptides.

Benefits of technology

It enables rapid and accurate characterization of peptides at the single-molecule level, avoiding amplification bias and providing direct measurement of peptide length, identity, sequence, and secondary structure.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure SMS_1
    Figure SMS_1
  • Figure SMS_2
    Figure SMS_2
  • Figure SMS_12
    Figure SMS_12
Patent Text Reader

Abstract

Provided herein are methods of characterizing a target polypeptide as it moves relative to a nanopore. Related kits, systems, and devices for performing such methods are also provided.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This application is a divisional application of Chinese Patent Application No. 2020800834172, filed on December 1, 2020, entitled "Method for Characterizing Target Peptides Using Nanopores". Technical Field

[0002] This disclosure relates to a method for characterizing a target peptide by forming a conjugate of a target peptide with a polynucleotide and using a polynucleotide-treated protein to control the movement of the conjugate relative to a nanopore. This disclosure also relates to kits, systems, and apparatus for performing such methods. Background Technology

[0003] The characterization of biomolecules is becoming increasingly important in biomedical and biotechnological applications. For example, nucleic acid sequencing allows the study of genomes and the proteins they encode, enabling the correlation between nucleic acid mutations and observable phenomena such as disease indications. Nucleic acid sequencing can be used in evolutionary biology to study relationships between organisms. Metagenomics involves identifying organisms present in a sample, such as microorganisms in the microbiome, where nucleic acid sequencing allows for the identification of such organisms. While techniques for characterizing polynucleotides (e.g., sequencing polynucleotides) have been extensively developed, techniques for characterizing peptides are less advanced, despite their significant biotechnological importance. For example, knowledge of protein sequences can allow for the establishment of structure-activity relationships and influence rational drug development strategies for developing ligands targeting specific receptors. The identification of post-translational modifications is also crucial for understanding the functional properties of many proteins. For instance, typically 30%–50% of protein species are phosphorylated in eukaryotes. Some proteins may have multiple phosphorylation sites for activating or inactivating proteins, promoting protein degradation, or regulating interactions with protein couplers. Therefore, methods for characterizing proteins and other peptides are urgently needed.

[0004] Known methods for characterizing peptides include mass spectrometry and Edman degradation.

[0005] Protein mass spectrometry involves characterizing whole proteins or fragments thereof in ionized form. Known protein mass spectrometry methods include electrospray ionization (ESI) and matrix-assisted laser desorption / ionization (MALDI). Mass spectrometry has some advantages, but the results obtained can be affected by the presence of contaminants, and it can be difficult to process brittle molecules without breaking them. Furthermore, mass spectrometry is not a single-molecule technique and provides only a limited amount of information about the sample being queried. Mass spectrometry is not suitable for characterizing differences within a group of peptide samples and is inefficient when attempting to distinguish adjacent residues.

[0006] Edman degradation, which allows for residue-by-residue sequencing of peptides, is an alternative to mass spectrometry. Edman degradation sequences peptides by sequentially cleaving N-terminal amino acids and then characterizing the individually cleaved residues using chromatography or electrophoresis. However, edman sequencing is slow, involves expensive reagents, and, like mass spectrometry, is not a single-molecule technique.

[0007] Therefore, there remains a pressing need for new technologies to characterize peptides, especially at the single-molecule level. Single-molecule techniques for characterizing biomolecules such as polynucleotides have proven particularly attractive due to their high fidelity and avoidance of amplification bias.

[0008] One attractive approach for single-molecule characterization of biomolecules, such as peptides, is nanopore sensing. Nanopore sensing is an analyte detection and characterization method that relies on the observation of individual binding or interaction events between analyte molecules and ion conduction channels. Nanopore sensors can be generated by placing nanoscale single pores in an electrically insulating membrane and measuring the voltage-driven ion current flowing through the pore in the presence of analyte molecules. The presence of an analyte inside or near the nanopore will alter the ion flow through the pore, resulting in changes in the ions or current measured on the channel. The identity of the analyte is revealed by its unique current characteristics, particularly the duration and extent of the current block and the changes in current level during interaction with the pore. Nanopore sensing has the potential to allow for rapid and inexpensive peptide characterization.

[0009] Nanopore sensing and characterization of peptides have been proposed in this field. For example, WO 2013 / 123379 discloses the use of NTP-driven protein processing unfolding enzymes for processing proteins to be transported through nanopores. However, alternative and / or improved methods for characterizing peptides are still needed. Summary of the Invention

[0010] This disclosure relates to a method for characterizing a target peptide. The method includes conjugating the target peptide with a polynucleotide to form a peptide-polynucleotide conjugate. The method includes contacting the conjugate with a polynucleotide-treated protein. The polynucleotide-treated protein is capable of controlling the movement of the polynucleotide relative to a nanopore. As the conjugate moves relative to the nanopore, one or more measurements specific to the peptide are performed. In this manner, the target peptide included in the conjugate is characterized.

[0011] Therefore, this paper provides a method for characterizing target peptides, the method comprising:

[0012] - The target polypeptide is conjugated with a polynucleotide to form a polynucleotide-peptide conjugate;

[0013] - Contact the conjugate with a polynucleotide-treated protein capable of controlling the movement of the polynucleotide relative to the nanopore; and

[0014] - As the conjugate moves relative to the nanopore, one or more measurements specific to the polypeptide are performed to characterize the polypeptide.

[0015] In some embodiments, the nanopore has a contraction region. In some embodiments, the nanopore is modified to extend the distance between the polynucleotide-treated protein and the contraction region of the nanopore. In some embodiments, a replacement unit is used to separate the polynucleotide-treated protein from the nanopore, thereby extending the distance between the active site of the polynucleotide-treated protein and the nanopore. In some embodiments, the replacement unit comprises one or more proteins. In some embodiments, the polynucleotide-treated protein is modified to extend the distance from the active site of the polynucleotide-treated protein to the nanopore.

[0016] In some embodiments, when the portion of the conjugate in contact with the active site of the polynucleotide-treated protein comprises a polypeptide, the polynucleotide-treated protein is able to maintain binding to the conjugate. In some embodiments, the polynucleotide-treated protein is modified to prevent dissociation from the conjugate when the polynucleotide-treated protein contacts a portion of the conjugate comprising a polypeptide. In some embodiments, the polynucleotide-treated protein is modified to completely or partially close openings present in at least one conformational state of an unmodified protein through which the polynucleotide chain can unbind. In some embodiments, the polynucleotide-treated protein is a helicase.

[0017] In some embodiments, the conjugate comprises a plurality of polypeptide moieties and / or a plurality of polynucleotide moieties.

[0018] In some embodiments, the polypeptide has a length of 2 to about 50 peptide units. In some embodiments, the polypeptide is maintained in a linearized form.

[0019] In some embodiments, the polynucleotide is about 10 to about 1000 nucleotides in length. In some embodiments, one or more adaptors and / or one or more ligands and / or one or more anchors are linked to the polynucleotide in the conjugate.

[0020] In some embodiments of the disclosed method,

[0021] i) The polynucleotide-processing protein is located on the cis side of the nanopore, and the polynucleotide-processing protein controls the movement of the conjugate from the cis side of the nanopore to the trans side of the nanopore; or

[0022] ii) The polynucleotide processing protein is located on the trans side of the nanopore, and the polynucleotide processing protein controls the movement of the conjugate from the trans side of the nanopore to the cis side of the nanopore.

[0023] In some embodiments, the polynucleotide processing protein is located on the cis side of the nanopore, and the polynucleotide processing protein controls the movement of the polynucleotide from the cis side of the nanopore to the trans side of the nanopore, thereby controlling the movement of the polypeptide through the nanopore. In some embodiments, the polynucleotide processing protein is located on the trans side of the nanopore, and the polynucleotide processing protein controls the movement of the polynucleotide from the trans side of the nanopore to the cis side of the nanopore, thereby controlling the movement of the polypeptide through the nanopore.

[0024] In some embodiments, the conjugate comprises the form L-{PN}-P m One or more structures, wherein:

[0025] -L is the leading sequence, where L is optionally part N;

[0026] -P is a polypeptide;

[0027] -N includes polynucleotides; and

[0028] -m is 0 or 1;

[0029] The method includes passing the leader sequence (L) through the nanopore, thereby contacting the polypeptide (P) with the nanopore; and

[0030] i) The polynucleotide processing protein is located on the cis side of the nanopore, and the method includes allowing the polynucleotide processing protein to control the movement of the polynucleotide moiety (N) from the cis side of the nanopore to the trans side of the nanopore, thereby controlling the movement of the polypeptide (P) through the nanopore; or

[0031] ii) The polynucleotide processing protein is located on the trans side of the nanopore, and the method includes allowing the polynucleotide processing protein to control the movement of the polynucleotide moiety (N) from the trans side of the nanopore to the cis side of the nanopore, thereby controlling the movement of the polypeptide (P) through the nanopore.

[0032] In some embodiments, the conjugate comprises the form L-P1-N-{PN} n -P m One or more structures, wherein:

[0033] -n is a positive integer;

[0034] -L is the leading sequence, where L is optionally part N;

[0035] - Each P, which may be the same or different, is a polypeptide;

[0036] - Each N, which may be the same or different, includes a polynucleotide; and

[0037] -m is 0 or 1;

[0038] The method includes passing the leader sequence (L) through the nanopore, thereby contacting the polypeptide (P1) with the nanopore; and

[0039] i) The polynucleotide processing protein is located on the cis side of the nanopore, and the method includes allowing the polynucleotide processing protein to control the sequential movement of each polynucleotide (N) from the cis side of the nanopore to the trans side of the nanopore, thereby controlling the sequential movement of each polypeptide (P) through the nanopore; or

[0040] ii) The polynucleotide processing protein is located on the trans side of the nanopore, and the method includes allowing the polynucleotide processing protein to control the sequential movement of each polynucleotide (N) from the trans side of the nanopore to the cis side of the nanopore, thereby controlling the sequential movement of each polypeptide (P) through the nanopore.

[0041] In some embodiments of the disclosed method,

[0042] i) The polynucleotide-processing protein is located on the cis side of the nanopore, and the polynucleotide-processing protein controls the movement of the conjugate from the trans side of the nanopore to the cis side of the nanopore; or

[0043] ii) The polynucleotide processing protein is located on the trans side of the nanopore, and the polynucleotide processing protein controls the movement of the conjugate from the cis side of the nanopore to the trans side of the nanopore.

[0044] In some embodiments, the polynucleotide-processing protein is located on the cis side of the nanopore, and the polynucleotide-processing protein controls the movement of the polynucleotide from the trans side of the nanopore to the cis side of the nanopore, thereby controlling the movement of the polypeptide through the nanopore. In some embodiments, the polynucleotide-processing protein is located on the trans side of the nanopore, and the polynucleotide-processing protein controls the movement of the polynucleotide from the cis side of the nanopore to the trans side of the nanopore, thereby controlling the movement of the polypeptide through the nanopore.

[0045] In some embodiments, the conjugate comprises the form L-{PN}-P m One or more structures, wherein:

[0046] -L is the leading sequence, where L is optionally part N;

[0047] -P is a polypeptide;

[0048] -N includes polynucleotides;

[0049] -m is 0 or 1;

[0050] The method includes passing the leader sequence (L) through the nanopore, thereby contacting the polypeptide (P) with the nanopore; and

[0051] i) The polynucleotide-processing protein is located on the cis side of the nanopore, and the method includes allowing the polynucleotide-processing protein to control the movement of the polynucleotide (N) from the trans side of the nanopore to the cis side of the nanopore, thereby controlling the movement of the polypeptide (P) through the nanopore; or

[0052] i) The polynucleotide processing protein is located on the trans side of the nanopore, and the method includes allowing the polynucleotide processing protein to control the movement of the polynucleotide (N) from the cis side of the nanopore to the trans side of the nanopore, thereby controlling the movement of the polypeptide (P) through the nanopore.

[0053] In some embodiments, the conjugate includes a capped portion linked to the polypeptide via an optional linker, and the method includes...

[0054] i) Contact the conjugate with the nanopore such that the capping portion is located on the side of the nanopore opposite to the polynucleotide-treated protein;

[0055] ii) Contact the polynucleotide of the conjugate with the polynucleotide-treated protein;

[0056] iii) Allowing the polynucleotide-treated protein to control the movement of the polynucleotide relative to the nanopore, thereby controlling the movement of the polypeptide through the nanopore;

[0057] iv) When the capping portion contacts the nanopore, thereby preventing the conjugate from moving further through the nanopore, the polynucleotide-treated protein is allowed to transiently unbind from the polynucleotide, causing the conjugate to move through the nanopore under applied force in a direction opposite to the direction of movement controlled by the polynucleotide-treated protein; and

[0058] v) Optionally repeat steps (ii) to (iv) to cause the polypeptide to oscillate through the nanopore.

[0059] In some embodiments, the one or more measurements are specific to the polypeptide selected from one or more of the following characteristics: (i) the length of the polypeptide; (ii) the identity of the polypeptide; (iii) the sequence of the polypeptide; (iv) the secondary structure of the polypeptide; and (v) whether the polypeptide is modified.

[0060] This article also provides a nanopore including a contraction region, wherein the nanopore is modified to increase the distance between the contraction region and the polynucleotide-processing protein in contact with the nanopore.

[0061] A system is also provided, the system comprising:

[0062] - Nanopores, wherein the nanopores include contraction regions;

[0063] - Conjugates, said conjugates comprising polypeptides conjugated to polynucleotides; and

[0064] - Polynucleotide-treated proteins;

[0065] in

[0066] i) The nanopore is modified to increase the distance between the contractile region and the active site of the polynucleotide-treated protein when the polynucleotide-treated enzyme contacts the nanopore; and / or

[0067] ii) The system further includes one or more replacement units disposed between the nanopore and the polynucleotide-treated protein, thereby extending the distance between the nanopore and the active site of the polynucleotide-treated protein.

[0068] In some embodiments, the nanopore, the conjugate and / or the polynucleotide-treated protein, and optionally one or more substitutional units, if present, as defined herein.

[0069] A kit is also provided, the kit comprising:

[0070] - Nanopores, wherein the nanopores include contraction regions;

[0071] - A polynucleotide, said polynucleotide comprising a reactive functional group for conjugation to a target polynucleotide; and

[0072] - Polynucleotide-treated proteins.

[0073] In some embodiments, (i) the nanopore is modified to increase the distance between the contractile region and the polynucleotide-treated protein when the polynucleotide-treated enzyme contacts the nanopore; and / or (ii) the kit further includes one or more replacement units for extending the distance between the nanopore and the active site of the polynucleotide-treated protein. In some embodiments, the nanopore, the polynucleotide and / or the polynucleotide-treated protein, and optionally the one or more replacement units, if present, are as defined herein. Attached Figure Description

[0074] Figure 1 A schematic diagram of a non-limiting example of an embodiment of the disclosed method is shown, wherein a polynucleotide processing protein located on the cis side of a nanopore controls the movement of a conjugate comprising a polynucleotide (DNA2) conjugated with a polypeptide from the cis side of the nanopore to the trans side of the nanopore, thus allowing characterization of the polypeptide as it moves relative to the nanopore. As shown, an optional leader sequence (DNA1) is linked to the conjugate to facilitate polypeptide passage through the nanopore. RED (discussed herein) is shown as the conceptual distance between the contraction in the nanopore and the active site of the polynucleotide processing protein. (A) The substrate can be captured in the nanopore, for example, from the cis side of the membrane, by applying, for example, a positive voltage to the trans side of the membrane. The polynucleotide processing protein moves along the polynucleotide portion in the direction indicated by the dashed arrow and feeds the substrate into the pore, proceeding to state (B). As the polynucleotide processing protein moves along the polynucleotide (e.g., in a nucleotide-fueled step), the polynucleotide processing protein feeds the conjugate into the nanopore, and the peptide portion passes through the nanopore.

[0075] Figure 2 It shows Figure 1 A schematic diagram of an embodiment of the general setup shown. In this non-limiting example, the polynucleotide-treated protein initially stops at a spacer (X) in the polynucleotide moiety of the conjugate (DNA). An adaptor is attached to the polynucleotide moiety of the conjugate and has a tether attached thereto to position the conjugate in a membrane within a nanopore region for characterization. Steps (A) and (B) are consistent with those for… Figure 1 As described above. When a polynucleotide-processing protein treats a polynucleotide, the polynucleotide-processing protein may substitute an adaptor.

[0076] Figure 3 It shows Figure 1The diagram illustrates another embodiment of the general setup shown. In this non-limiting example, the conjugate includes multiple polynucleotide and polypeptide moieties that sequentially move through a nanopore under the control of a polynucleotide-treated protein. As shown, the polynucleotide-treated protein is initially loaded onto a first polynucleotide moiety (DNA1) of the conjugate, and this moiety of the conjugate is moved through the nanopore. The polynucleotide-treated protein causes the first polypeptide moiety of the conjugate to pass through the nanopore without dissociating from the conjugate. The polynucleotide-treated protein then contacts a second polynucleotide moiety (DNA2) of the conjugate and controls the movement of the second polynucleotide moiety through the nanopore. Additional polypeptide and polynucleotide moieties (not shown) can similarly move sequentially relative to the nanopore.

[0077] Figure 4 It shows the use of Figure 3 A schematic diagram of a non-limiting example of the substrate described in the embodiments. A: The first polynucleotide portion (DNA1) of the conjugate comprises: a sequencing Y adaptor having a leader sequence (dashed line) to facilitate capture in the nanopore; a tethering strand that enables tethering to the membrane to position the conjugate in the nanopore region; and a polynucleotide-processing protein stabilized by a spacer (X). As shown, the polynucleotide portion of the conjugate comprises double-stranded DNA. B: A variation of the embodiment of the substrate shown in (A); the tethering strand (or additional tethering strand) may be located on the second polynucleotide portion (DNA2) of the conjugate. The symbols “top” and “bottom” are used purely for ease of understanding. C: A schematic diagram illustrating the method of using the substrate shown in 4(A) of the present invention.

[0078] Figure 5 A schematic diagram of another non-limiting example of a substrate for the disclosed method is shown. The conjugate may comprise multiple polynucleotide and polypeptide moieties (n>0), which may be sequentially processed by a polynucleotide-treated protein for characterization via the nanopores described herein.

[0079] Figure 6A schematic diagram of a non-limiting example of an embodiment of the disclosed method is shown, wherein a polynucleotide processing protein located on the cis side of the nanopore controls the movement of a conjugate comprising a polynucleotide (DNA2) conjugated with a polypeptide from the trans side of the nanopore to the cis side of the nanopore, thus allowing characterization of the polypeptide as it moves relative to the nanopore. As shown, an optional leader sequence is linked to the conjugate to facilitate the initial passage of the polypeptide through the nanopore. (A) The substrate can be captured in the nanopore, for example, from the cis side of the membrane, by applying, for example, a positive voltage to the trans side of the membrane. The polynucleotide processing protein moves along the polynucleotide portion in the direction indicated by the dashed arrow to remove the substrate out of the pore, proceeding to state (B). As the polynucleotide processing protein moves along the polynucleotide (e.g., in a nucleotide-fueled step), the polynucleotide processing protein drives the conjugate out of the nanopore. The polypeptide portion of the conjugate thus passes through the nanopore (state C) and is thus characterized.

[0080] Figure 7 A schematic diagram illustrating a non-limiting example of the use of a capping portion (black square) that prevents the conjugate from moving through the nanopore upon contact. In the non-limiting embodiment shown, the polynucleotide treatment protein is a polymerase that controls the movement of the conjugate by extending the polynucleotide portion of the conjugate. Chain extension can continue until the capping portion reaches the nanopore. Dissociation of the newly synthesized chain allows the conjugate to move through the nanopore from the cis side back to the trans side, and the polynucleotide can then recycle the conjugate to move through the nanopore from the trans side to the cis side. In this way, the conjugate can be "flossed" through the nanopore. Other polynucleotide treatment proteins can be used in similar methods.

[0081] Figure 8 Schematic diagrams illustrating non-limiting examples of strategies for increasing the distance between a nanopore (e.g., contraction within the nanopore) and the active site of a polynucleotide-treated protein for controlling the movement of conjugates relative to the nanopore are shown. A: Schematic diagram of an unmodified pore of an unmodified RED. B: Nanopores can be modified to extend the RED. C: Substitution units can be used to displace polynucleotide-treated proteins from the nanopore, thus extending the RED. D: Various polynucleotide-treated proteins can be used to displace the active polynucleotide-treated protein controlling the movement of conjugates relative to the nanopore. These examples are described in more detail herein.

[0082] Figure 9 For clarity, a representative current-time trace of Example 1 with a sketch of the corresponding construct is shown. States A)-D) correspond to... Figure 4The states described in C are: A) - nanopore trapping leader strand, B) - Y-adaptor translocation across the nanopore reader head (RED), C) - peptide translocation across the RED, and D) - translocation of the polynucleotide tail (DNA2). The first trace represents only partial translocation events of the Y-adaptor (states A and B only), while the second trace shows the translocation of the entire conjugated polynucleotide-peptide across the nanopore. Data obtained as described in Example 1 (this data is for polynucleotide-peptide conjugates containing the peptide with sequence SEQ ID NO:20).

[0083] Figure 10 The current-to-time trace demonstrates the high throughput of data collection; there were 5 capture events within a 3-second timeframe, 4 of which corresponded to complete polynucleotide-peptide conjugates (event 3 was a partial translocation of the Y-adaptor only). The data is from Example 1 (this data pertains to polynucleotide-peptide conjugates containing the peptide with sequence SEQ ID NO:20).

[0084] Figure 11 The peptide sequence corresponding to GGSGRRSGSG (SEQ ID NO:21) Figure 4 Current traces of translocation of polynucleotide-peptide conjugates described in B and Example 1. A: relative to Figure 4 and Figure 9 A: Eleven instances of state-aligned traces as described in [the diagram]. B: A superposition of the same eleven traces. C: A stacked diagram of the eleven example traces to demonstrate how the time axis is normalized using a dynamic time warp algorithm to facilitate optimal alignment of key trace features.

[0085] Figure 12 The peptide sequence corresponding to GGSGYYSGSG (SEQ ID NO:22) Figure 4 Current traces of translocation of polynucleotide-peptide conjugates described in B and Example 1. A: relative to Figure 4 and Figure 9 A: Twelve instances of state-aligned traces as described in the diagram. B: A superposition of the same 12 traces. C: A stacked diagram of the 12 example traces.

[0086] Figure 13 The peptide sequence corresponding to GGSGDDSGSG (SEQ ID NO:20) Figure 4 Current traces of translocation of polynucleotide-peptide conjugates described in B and Example 1. A: relative to Figure 4 and Figure 9 A: Eleven instances of state-aligned traces as described in the diagram. B: A superposition of the same eleven traces. C: A stacked diagram of the eleven example traces.

[0087] Figure 14Schematic structure of the construct obtained using the peptide of SEQ ID NO:22; Y-adaptor including polynucleotide chains of SEQ ID NO:11, 12 and 13; and polynucleotide tail including polynucleotide chains of SEQ ID NO:14 and 16 (described in Example 1).

[0088] Figure 15 Representative current-time traces of Example 2 compared to Example 1. States A)-D) correspond to... Figure 4 The states described in C are: A) - nanopore trapping leader strand, B) - Y-adaptor translocation across the nanopore reader head (RED), C) - polypeptide translocation across the RED, and D) - translocation of the polynucleotide tail (DNA2). The traces in the top figure were collected according to the protocol in Example 1 (using a peptide pre-modified during synthesis); the traces in the bottom figure show the polynucleotide-peptide translocation conjugated according to the protocol in Example 2 (using the unmodified peptide of SEQ ID NO: 23; i.e., the same sequence as in the corresponding trace in Example 1).

[0089] Figure 16 Representative current-to-time traces of the translocation of polynucleotide-peptide conjugates between a .10-amino acid peptide (top plot; SEQ ID NO:20) and a 21-amino acid peptide (bottom plot; SEQ ID NO:24). States A)-D) correspond to... Figure 4 The states described in C are: A) - nanopore trapping leader chain, B) - Y-adaptor translocation across the nanopore reader head (RED), C) - polypeptide translocation across the RED, and D) - translocation of the polynucleotide tail. The results are described in Example 3. Detailed Implementation

[0090] This invention will be described with reference to specific embodiments and certain accompanying drawings, but is not limited thereto by the claims. No reference numerals in the claims should be construed as limiting the scope. It should be understood, of course, that not all aspects or advantages can be achieved according to any particular embodiment of the invention. Therefore, for example, those skilled in the art will recognize that the invention may be embodied or practiced in a manner that achieves or optimizes one or more advantages as taught herein, without necessarily achieving other aspects or advantages as may be taught or suggested herein.

[0091] The invention (both in terms of organization and method of operation) and its features and advantages can be best understood by referring to the following detailed description when read in conjunction with the accompanying drawings. Aspects and advantages of the invention will become apparent from one or more embodiments described below, and will be set forth with reference to said embodiments. Throughout this specification, reference to “one embodiment” or “an embodiment” means that a particular feature, structure, or characteristic described in connection with the embodiment is included in at least one embodiment of the invention. Therefore, the phrases “in one embodiment” or “in an embodiment” appearing in various places throughout this specification do not necessarily refer to the same embodiment, but may refer to the same embodiment. Similarly, it should be understood that in the description of exemplary embodiments of the invention, various features of the invention are sometimes grouped together in a single embodiment, drawing, or description thereof for the purpose of simplifying the disclosure and aiding in the understanding of one or more of the various inventive aspects. However, the method of this disclosure should not be construed as reflecting an intention to reflect more features required by the claimed invention than expressly recited in each claim. Rather, as reflected in the following claims, the inventive aspect lies in fewer features than all of the individual foregoing disclosed embodiments.

[0092] It should be understood that, unless the context otherwise requires, the “embodiments” of this disclosure can be specifically combined together. Specific combinations of all disclosed embodiments (unless the context otherwise implies) are further disclosed embodiments of the claimed invention.

[0093] Furthermore, as used in this specification and the appended claims, unless otherwise expressly indicated, the singular forms "a / an" and "the" both encompass the plural objects. Thus, for example, a reference to "polynucleotide" includes two or more polynucleotides; a reference to "motor protein" includes two or more such proteins; a reference to "helicase" includes two or more helicases; a reference to "monomer" refers to two or more monomers; a reference to "pore" includes two or more pores, etc.

[0094] All publications, patents, and patent applications cited in this article, whether mentioned above or below, are incorporated herein by reference in their entirety.

[0095] definition

[0096] When referring to singular nouns (e.g., "a / an", "the"), the use of indefinite or definite articles includes the plural form of the noun unless specifically stated otherwise. The use of the term "comprising" in this specification and claims does not exclude other elements or steps. Furthermore, the terms first, second, third, etc., in the specification and claims are used to distinguish similar elements and are not necessarily used to describe order or chronological sequence. It should be understood that the terms thus used are interchangeable where appropriate, and embodiments of the invention described herein can operate in orders other than those described or illustrated herein. The following terms or definitions are provided only to aid in understanding the invention. Unless specifically defined herein, all terms used herein have the same meaning to those skilled in the art to which this invention pertains. For the definitions and terminology used in this field, practicing physicians have specifically referred to Sambrook et al., *Molecular Cloning: A Laboratory Manual*, 4th edition, Cold Spring Harbor Press, Plainsview, New York (2012); and Ausubel et al., *Current Protocols in Molecular Biology* (Supplement 114), John Wiley & Sons, New York (2016). The definitions provided herein should not be construed as having a scope less than that understood by one of ordinary skill in the art.

[0097] When referring to measurable values ​​such as quantity or duration, the term “about” as used herein means to encompass deviations from the specified value of ±20% or ±10%, more preferably ±5%, even more preferably ±1%, and still more preferably ±0.1%, as such deviations are suitable for performing the disclosed method.

[0098] As used herein, the terms “nucleotide sequence,” “DNA sequence,” or “one or more nucleic acid molecules” refer to a polymer of nucleotides of any length, whether ribonucleotides or deoxyribonucleotides. This term refers only to the primary structure of the molecule. Therefore, this term encompasses both double-stranded and single-stranded DNA, as well as RNA. As used herein, the term “nucleic acid” is a single-stranded or double-stranded covalently linked sequence of nucleotides, wherein the 3' and 5' ends of each nucleotide are linked by a phosphodiester bond. Polynucleotides may consist of deoxyribonucleotide bases or ribonucleotide bases. Nucleic acids can be synthesized in vitro or isolated from natural sources. Nucleic acids may further comprise modified DNA or RNA, such as methylated DNA or RNA, or RNA that has undergone post-translational modifications, such as 5' capping with 7-methylguanosine, 3' processing such as cleavage and polyadenylation, and splicing. Nucleic acids can also include synthetic nucleic acids (XNAs), such as hexitol nucleic acid (HNA), cyclohexene nucleic acid (CeNA), threonine nucleic acid (TNA), glycerol nucleic acid (GNA), locked nucleic acid (LNA), and peptide nucleic acid (PNA). The size of a nucleic acid (also referred to herein as a “polynucleotide”) is typically expressed as the number of base pairs (bp) in a double-stranded polynucleotide, or, in the case of a single-stranded polynucleotide, as the number of nucleotides (nt). One thousand bp or nt equals one thousand bases (kb). Polynucleotides shorter than approximately 40 nucleotides are often referred to as “oligonucleotides” and may include primers used for manipulating DNA, such as via polymerase chain reaction (PCR).

[0099] In the context of this disclosure, the term “amino acid” is used in its broadest sense and refers to an organic compound containing amine (NH2) and carboxyl (COOH) functional groups and a side chain (e.g., an R group) specific to each amino acid. In some embodiments, an amino acid refers to a naturally occurring Lα-amino acid or residue. One and three commonly used letter abbreviations for naturally occurring amino acids are used herein: A = Ala; C = Cys; D = Asp; E = Glu; F = Phe; G = Gly; H = His; I = Ile; K = Lys; L = Leu; M = Met; N = Asn; P = Pro; Q = Gln; R = Arg; S = Ser; T = Thr; V = Val; W = Trp; and Y = Tyr (Lehninger, AL, (1975) Biochemistry, 2nd ed., pp. 71-92, Worth Publishers, New York). The general term “amino acid” further includes D-amino acids, trans-amino acids, and chemically modified amino acids, such as amino acid analogs, naturally occurring amino acids that are not typically incorporated into proteins (e.g., ortholeucine), and chemically synthesized compounds that have properties known in the art as amino acid characteristics (e.g., β-amino acids). For example, analogs or mimics of phenylalanine or proline that allow conformational restrictions to the same peptide compounds as natural Phe or Pro are included within the definition of an amino acid. Such analogs and mimics are referred to herein as “functional equivalents” of the corresponding amino acids. Other examples of amino acids are listed by Roberts and Vellaccio, *The Peptides: Analysis, Synthesis, Biology*, edited by Gross and Meiehofer, Vol. 5, p. 341, Academic Press, Inc., NY, 1983, which is incorporated herein by reference.

[0100] The terms “polypeptide” and “peptide” are used interchangeably herein to refer to polymers containing amino acid residues, as well as their variants and synthetic analogs. Therefore, these terms apply to amino acid polymers where one or more amino acid residues are synthetic, non-naturally occurring amino acids, such as chemical analogs of corresponding naturally occurring amino acids, and to polymers containing naturally occurring amino acids. Polypeptides may also undergo maturation or post-translational modification processes, which may include, but are not limited to, glycosylation, proteolytic cleavage, lipolysis, signal peptide cleavage, propeptide cleavage, phosphorylation, etc. Peptides can be prepared using recombinant techniques, for example, by expressing recombinant or synthetic polynucleotides. Recombinant peptides are typically substantially free of culture medium, for example, the culture medium comprises less than about 20% of the volume of the protein formulation, more preferably less than about 10%, and most preferably less than about 5%.

[0101] The term "protein" is used to describe folded polypeptides that have secondary or tertiary structures. Proteins can consist of a single polypeptide or can comprise multiple polypeptides that assemble to form a multimer. The multimer can be a homooligomer or a heterooligomer. Proteins can be naturally occurring or wild-type proteins, or modified or non-natural proteins. Proteins can differ from wild-type proteins, for example, through the addition, substitution, or deletion of one or more amino acids.

[0102] Protein “variants” encompass peptides, oligopeptides, polypeptides, proteins, and enzymes that have amino acid substitutions, deletions, and / or insertions relative to the unmodified or wild-type protein in question, and possess biological and functional activities similar to those of the unmodified protein from which they are derived. As used herein, the term “amino acid identity” refers to the degree to which sequences are identical on an amino acid-to-amino acid basis within a comparison window. Thus, the “sequence identity percentage” is calculated by comparing two optimally aligned sequences within a comparison window, determining the number of positions in which identical amino acid residues (e.g., Ala, Pro, Ser, Thr, Gly, Val, Leu, Ile, Phe, Tyr, Trp, Lys, Arg, His, Asp, Glu, Asn, Gln, Cys, and Met) appear in both sequences to produce the number of matching positions, dividing the number of matching positions by the total number of positions in the comparison window (i.e., the window size), and multiplying the result by 100 to produce the sequence identity percentage.

[0103] For all aspects and embodiments of the invention, the "variant" has at least 50%, 60%, 70%, 80%, 90%, 95%, or 99% complete sequence identity with the corresponding wild-type protein's amino acid sequence. Sequence identity can also be a fragment or portion of a full-length polynucleotide or polypeptide. Thus, a sequence may have only 50% overall sequence identity with a full-length reference sequence, but sequences of specific regions, domains, or subunits may share 80%, 90%, or up to 99% sequence identity with the reference sequence.

[0104] The term "wild-type" refers to a gene or gene product isolated from a naturally occurring source. Wild-type genes are the most frequently observed genes in a population and are therefore arbitrarily engineered to be in a "normal" or "wild-type" form. Conversely, the terms "modified," "mutant," or "variant" refer to a gene or gene product that exhibits sequence modifications (e.g., substitution, truncation, or insertion), post-translational modifications, and / or functional characteristics (e.g., altered properties) compared to a wild-type gene or gene product. Note that naturally occurring mutants can be isolated; these mutants are identified by the fact that they possess altered properties compared to a wild-type gene or gene product. Methods for introducing or substituting naturally occurring amino acids are well known in the art. For example, the codon for methionine (ATG) can be substituted with the codon for arginine (CGT) at the relevant position in the polynucleotide encoding the mutant monomer, while methionine (M) is replaced with arginine (R). Methods for introducing or substituting non-naturally occurring amino acids are also well known in the art. For example, non-naturally occurring amino acids can be introduced by including synthetic aminoacyl-tRNA in the IVTT system used to express mutant monomers. Alternatively, this can be introduced by expressing mutant monomers in *E. coli* that are auxotrophic for the specific amino acid in the presence of synthetic (i.e., non-naturally occurring) analogs of those specific amino acids. If the mutant monomer is produced using partial peptide synthesis, it can also be produced via naked linking. Conservative substitution replaces an amino acid with another amino acid having a similar chemical structure, similar chemical properties, or similar side chain volume. The introduced amino acid can have similar polarity, hydrophilicity, hydrophobicity, basicity, acidity, neutrality, or charge as the substituted amino acid. Alternatively, conservative substitution can introduce another aromatic or aliphatic amino acid to replace a pre-existing aromatic or aliphatic amino acid. Conservative amino acid changes are well known in the art and can be selected based on the properties of the 20 major amino acids defined in Table 1 below. In the case of amino acids having similar polarity, this can also be determined with reference to the hydrophilicity scale of the amino acid side chains in Table 2.

[0105] Table 1 - Chemical properties of amino acids

[0106]

[0107] Table 2 - Hydrophilicity Scale

[0108]

[0109] Mutants or modified proteins, monomers, or peptides can also be chemically modified in any manner and at any site. Preferably, mutants or modified monomers are chemically modified by attaching the molecule to one or more cysteine ​​residues (cysteine ​​linkage), attaching the molecule to one or more lysine residues, attaching the molecule to one or more non-natural amino acids, enzymatic modification of epitopes, or terminal modification. Suitable methods for performing such modifications are well known in the art. Mutants of modified proteins, monomers, or peptides can be chemically modified by attaching any molecule. For example, mutants of modified proteins, monomers, or peptides can be chemically modified by linking dyes or fluorophores.

[0110] The disclosed method

[0111] This disclosure relates to a method for characterizing peptides by forming conjugates with polynucleotides and using polynucleotide-treated proteins to control the movement of the conjugates relative to nanopores.

[0112] Compared to methods that seek to control the movement of peptides relative to nanopores using peptide-processing enzymes, the method disclosed herein enables the use of polynucleotide-processing enzymes to control the movement of peptides relative to nanopores.

[0113] The methods disclosed herein utilize the ability of polynucleotide-processing proteins to control the movement of conjugates, including not only polynucleotides. Specifically, polynucleotide-processing proteins as described herein can be used to move conjugates, including peptides, in a controlled manner. Polynucleotide-processing proteins suitable for the disclosed methods are described in more detail herein.

[0114] Therefore, this paper provides a method for characterizing target peptides, the method comprising:

[0115] - The target polypeptide is conjugated with a polynucleotide to form a polynucleotide-peptide conjugate;

[0116] - Contact the conjugate with a polynucleotide-treated protein capable of controlling the movement of the polynucleotide relative to the nanopore; and

[0117] - As the conjugate moves relative to the nanopore, one or more measurements specific to the peptide are performed.

[0118] This characterizes the polypeptide.

[0119] Any suitable polypeptide can be characterized using the methods disclosed herein. In some embodiments, the target polypeptide is a protein or a naturally occurring polypeptide. In some embodiments, the polypeptide is a synthetic polypeptide. Polypeptides that can be characterized according to the disclosed methods are described in more detail herein.

[0120] Any suitable polynucleotide can be used to form conjugates for the methods disclosed herein. In some embodiments, the length of the polynucleotide is at least as long as a portion of the target polypeptide to be characterized. In some embodiments, the length of the polynucleotide is greater than a portion of the target polypeptide to be characterized. This is discussed in more detail below. Polynucleotides suitable for the disclosed methods are disclosed in more detail herein.

[0121] In the disclosed methods, any suitable manner can be used to conjugate the target peptide to the polynucleotide. Some exemplary methods are described in more detail herein.

[0122] The conjugate formed in the disclosed method is contacted with a polynucleotide-treated protein capable of controlling the movement of polynucleotides relative to nanopores. An exemplary polynucleotide-treated protein is described in more detail herein.

[0123] Polynucleotide-treated proteins control the movement of polynucleotides relative to nanopores. Therefore, polynucleotide-treated proteins control the movement of conjugates relative to nanopores. Any suitable nanopore can be used in the disclosed method. Nanopores suitable for the disclosed method are described in more detail herein.

[0124] The disclosed method includes performing one or more measurements specific to the peptide as the conjugate moves relative to the nanopore. These measurements can be any suitable measurement. Typically, they are electrical measurements, such as current measurements, and / or one or more optical measurements. The apparatus for recording suitable measurements and the information such measurements can provide are described in more detail herein.

[0125] Characterizing target peptides

[0126] As disclosed herein, polynucleotides can be used to control the movement of peptides relative to nanopores. Polynucleotide movement is controlled by polynucleotide-treated proteins. Because the polynucleotide is conjugated to the peptide in the conjugate, the polynucleotide movement drives the peptide movement.

[0127] Compared to methods known in the art for characterizing peptides, using polynucleotide-treated proteins to control polynucleotide movement and thus peptide movement may have advantages. For example, polynucleotide-treated proteins are capable of processing polynucleotides at a higher turnover rate compared to peptide-treated enzymes. This means that characterization data can be obtained more rapidly for peptides characterized according to the disclosed methods compared to previously known methods.

[0128] These and other advantages will become apparent throughout this disclosure.

[0129] In developing the methods of this disclosure, the inventors have discovered that the length of the characterizable peptide is generally improved when using nanopores with longer tubes or channels compared to those with shorter tubes or channels. Without being bound by theory, the inventors believe this may be because pores with longer tubes or channels, when used in conjunction with polynucleotide-treated proteins in the disclosed methods, result in a longer distance between the active site of the polynucleotide-treated protein and the contraction within the nanopore. This distance may also be referred to as the RED (reader-enzyme distance). Those skilled in the art will understand that the form of the nanopore (discussed below) is not limiting. The nanopore can be a protein nanopore or a solid nanopore. If the nanopore does not have narrowing within the channel of the nanopore, then the contraction, as used herein, for example, in one embodiment, can be identified by the opening of the nanopore.

[0130] Unbound by theory, it is hypothesized that the length of the polypeptide portion within the conjugate, which can be characterized by nanopores, can correspond to or be determined by RED. In other words, the nanopore can include a readhead, and one or more measurements are specific to the "read portion" of the polypeptide, wherein the length of the read portion corresponds to or is determined by the distance between the readhead and the active site of the polynucleotide-treated protein.

[0131] Therefore, in some embodiments of the disclosed method, the length of the polynucleotide is at least as long as a portion of the target polypeptide to be characterized. In some embodiments, the length of the polynucleotide is longer than a portion of the target polypeptide to be characterized. This ensures that the length of the polypeptide portion that can be characterized is not limited by the amount of polynucleotide, since the polynucleotide-processing protein controls the movement of the polynucleotide.

[0132] The method can be referenced. Figure 1 To understand, Figure 1 A non-limiting example of the disclosed method is shown. The conjugate may include polynucleotides and peptides, and is contacted with a polynucleotide-treated protein, allowing the peptide to pass through a nanopore. In the illustrated embodiment, additional polynucleotides are used to facilitate peptide passage through the nanopore. Such use is within the scope of the disclosed method; however, it is not required.

[0133] Polynucleotide processing involves the processing of proteins with polynucleotides conjugated to peptides. When polynucleotides process proteins with polynucleotides, the conjugates pass through nanopores, and consequently, the peptides pass through the nanopores. The peptides are characterized as they pass through the nanopores.

[0134] exist Figure 1In the examples shown, from the "viewpoint" of the polynucleotide-treated protein, the polynucleotide-treated protein causes the conjugate to "move into the pore." For example, as shown in the figure, the polynucleotide-treated protein is located on the cis side of the nanopore and moves the conjugate into the pore, i.e., from the cis side to the trans side. The reverse setup can also be used.

[0135] In other words, in some embodiments, the polynucleotide-processing protein is located on the cis side of the nanopore, and the polynucleotide-processing protein controls the movement of the conjugate from the cis side of the nanopore to the trans side of the nanopore. Therefore, in some embodiments, the polynucleotide-processing protein is located on the cis side of the nanopore, and the polynucleotide-processing protein controls the movement of the polynucleotide from the cis side of the nanopore to the trans side of the nanopore, thereby controlling the movement of the polypeptide through the nanopore.

[0136] In other embodiments, the polynucleotide-processing protein is located on the trans side of the nanopore, and the polynucleotide-processing protein controls the movement of the conjugate from the trans side of the nanopore to the cis side of the nanopore. Therefore, in some embodiments, the polynucleotide-processing protein is located on the trans side of the nanopore, and the polynucleotide-processing protein controls the movement of the polynucleotide from the trans side of the nanopore to the cis side of the nanopore, thereby controlling the movement of the polypeptide through the nanopore.

[0137] As explained herein, conjugates may include a leader sequence. Any suitable leader sequence may be used, as explained herein. Optionally, the leader sequence may be a polynucleotide. The leader sequence may be the same as or different from the polynucleotide in the conjugate. As explained above, the leader sequence can facilitate the passage of the conjugate through nanopores.

[0138] In other words, in some embodiments, the conjugate comprises the form L-{PN}-P m One or more structures, wherein:

[0139] -L is the leading sequence, where L is optionally part N;

[0140] -P is a polypeptide;

[0141] -N includes polynucleotides; and

[0142] -m is 0 or 1;

[0143] The method may include passing the leader sequence (L) through the nanopore, thereby bringing the polypeptide (P) into contact with the nanopore.

[0144] In some such embodiments, the polynucleotide processing protein is located on the cis side of the nanopore, and the method includes allowing the polynucleotide processing protein to control the movement of the polynucleotide moiety (N) from the cis side of the nanopore to the trans side of the nanopore, thereby controlling the movement of the polypeptide (P) through the nanopore. In other embodiments, the polynucleotide processing protein is located on the trans side of the nanopore, and the method includes allowing the polynucleotide processing protein to control the movement of the polynucleotide moiety (N) from the trans side of the nanopore to the cis side of the nanopore, thereby controlling the movement of the polypeptide (P) through the nanopore.

[0145] As explained in more detail herein, a conjugate may include one or more connectors and / or anchors. Figure 2 A non-limiting example of this setup is shown.

[0146] As explained in more detail herein, in some embodiments, the conjugate comprises multiple polynucleotides and peptides. In such embodiments, the polynucleotide-treated protein sequentially controls the movement of the polynucleotides relative to the nanopore, and thus the sequential movement of the peptides relative to the nanopore. In this way, each peptide within the conjugate can be sequentially characterized in the disclosed methods.

[0147] For example, the conjugate may include the form L-P1-N-{PN}nP m One or more structures, wherein:

[0148] -n is a positive integer;

[0149] -L is the leading sequence, where L is optionally part N;

[0150] - Each P, which may be the same or different, is a polypeptide;

[0151] - Each N, which may be the same or different, includes a polynucleotide; and

[0152] -m is 0 or 1;

[0153] The method may include passing the leader sequence (L) through the nanopore, thereby bringing the polypeptide (P1) into contact with the nanopore.

[0154] Typically, in such embodiments, n is 1 to about 1000, for example 2 to about 100, for example about 3 to about 10, for example 1, 2, 3, 4, 5, 6, 7, 8, 9 or 10.

[0155] In some such embodiments, the polynucleotide processing protein is located on the cis side of the nanopore, and the method includes allowing the polynucleotide processing protein to control the sequential movement of each polynucleotide (N) from the cis side of the nanopore to the trans side of the nanopore, thereby controlling the sequential movement of each polypeptide (P) through the nanopore. In other such embodiments, the polynucleotide processing protein is located on the trans side of the nanopore, and the method includes allowing the polynucleotide processing protein to control the sequential movement of each polynucleotide (N) from the trans side of the nanopore to the cis side of the nanopore, thereby controlling the sequential movement of each polypeptide (P) through the nanopore.

[0156] Those skilled in the art will understand that when the conjugate comprises more than one polypeptide, it is advantageous (as described in more detail herein) for the polynucleotide-treated protein to remain bound to the conjugate without dissociation upon contact with the polypeptide. For example, as Figure 3 As shown, this allows polynucleotide-treated proteins to bypass the polypeptide moieties in the conjugates when they come into contact with them, so as to move onto the continuous moieties of the polynucleotides, thereby controlling the movement of the conjugates relative to the nanopores.

[0157] Figure 4 Depicting according to Figure 3 The embodiments shown are non-limiting examples of more complex settings in which various adaptors and linkers are used to facilitate peptide characterization. Figure 4 Only one polypeptide moiety is shown, although those skilled in the art will understand that multiple such moiety can be combined, such as Figure 5 It is shown schematically in the middle.

[0158] Figure 6 Another non-limiting embodiment of the disclosed method is illustrated schematically. The conjugate may include polynucleotides and peptides, and is contacted with a polynucleotide-treated protein, allowing the peptide to pass through a nanopore. In the illustrated embodiment, a leader sequence (which may optionally be another polynucleotide) is used to facilitate the peptide's passage through the nanopore. Such use is within the scope of the disclosed method; however, it is not required.

[0159] Polynucleotide processing involves the processing of proteins with polynucleotides conjugated to peptides. When polynucleotides process proteins with polynucleotides, the conjugates pass through nanopores, and consequently, the peptides pass through the nanopores. The peptides are characterized as they pass through the nanopores.

[0160] exist Figure 6 In the examples shown, from the "viewpoint" of the polynucleotide-treated protein, the protein moves the conjugate "out" of the pore. For example, as shown in the figure, the polynucleotide-treated protein is located on the cis side of the nanopore and moves the conjugate into the pore, i.e., from the trans side to the cis side. The reverse setup can also be used.

[0161] In other words, in some embodiments, the polynucleotide-processing protein is located on the cis side of the nanopore, and the polynucleotide-processing protein controls the movement of the conjugate from the trans side of the nanopore to the cis side of the nanopore. Therefore, in some embodiments, the polynucleotide-processing protein is located on the cis side of the nanopore, and the polynucleotide-processing protein controls the movement of the polynucleotide from the trans side of the nanopore to the cis side of the nanopore, thereby controlling the movement of the polypeptide through the nanopore.

[0162] In other embodiments, the polynucleotide-processing protein is located on the trans side of the nanopore, and the polynucleotide-processing protein controls the movement of the conjugate from the cis side to the trans side of the nanopore. Therefore, in some embodiments, the polynucleotide-processing protein is located on the trans side of the nanopore, and the polynucleotide-processing protein controls the movement of the polynucleotide from the cis side to the trans side of the nanopore, thereby controlling the movement of the polypeptide through the nanopore.

[0163] Using similar notation as described above, in some embodiments, the conjugate includes the form L-{PN}-P m One or more structures, wherein:

[0164] -L is the leading sequence, where L is optionally part N;

[0165] -P is a polypeptide;

[0166] -N includes polynucleotides;

[0167] -m is 0 or 1;

[0168] The method may include passing the leader sequence (L) through the nanopore, thereby bringing the polypeptide (P) into contact with the nanopore.

[0169] In some such embodiments, the polynucleotide processing protein is located on the cis side of the nanopore, and the method includes allowing the polynucleotide processing protein to control the movement of the polynucleotide (N) from the trans side of the nanopore to the cis side of the nanopore, thereby controlling the movement of the polypeptide (P) through the nanopore. In other such embodiments, the polynucleotide processing protein is located on the trans side of the nanopore, and the method includes allowing the polynucleotide processing protein to control the movement of the polynucleotide (N) from the cis side of the nanopore to the trans side of the nanopore, thereby controlling the movement of the polypeptide (P) through the nanopore.

[0170] In some embodiments, particularly those where the polynucleotide-treated protein controls the "exit" of the conjugate from the nanopore as discussed above, the conjugate may include a capped portion connected to a polypeptide via an optional linker. The capped portion is typically too large to pass through the nanopore and therefore prevents further movement of the conjugate through the nanopore when the conjugate moves relative to the nanopore and the capped portion comes into contact with the nanopore. This allows the polynucleotide-treated protein to transiently unbind from the conjugate. In embodiments of the disclosed method, where the conjugate moves relative to the nanopore under an applied force (e.g., a voltage potential or chemical potential), the conjugate may then move "backward" through the pore in the opposite direction to the movement controlled by the polynucleotide-treated protein. This backward movement of the conjugate through the pore allows the polypeptide portion of the conjugate to be recharacterized.

[0171] The process can be repeated multiple times by sequentially allowing the polynucleotide-treated protein to bind and recombine with the conjugate. In this way, the conjugate can oscillate through the pore (i.e., the conjugate can be "combed" through the nanopore). This "combing" allows the polypeptide portion of the conjugate to be repeatedly characterized through the nanopore. In some embodiments, this allows for increased accuracy of the characterization information.

[0172] In such embodiments, any suitable end-capping portion can be used. For example, the conjugate can be modified with biotin and the end-capping portion can be, for example, streptavidin, avidin, or neutral avidin. The end-capping portion can be a large chemical group, such as a dendritic. The end-capping portion can be nanoparticles or beads. Other suitable end-capping portions will be apparent to those skilled in the art.

[0173] Figure 7 A non-limiting example of such a method is shown.

[0174] Therefore, in some embodiments, the method includes

[0175] i) Contact the conjugate with the nanopore such that the capping portion is located on the side of the nanopore opposite to the polynucleotide-treated protein;

[0176] ii) Contact the polynucleotide of the conjugate with the polynucleotide-treated protein;

[0177] iii) Allowing the polynucleotide-treated protein to control the movement of the polynucleotide relative to the nanopore, thereby controlling the movement of the polypeptide through the nanopore;

[0178] iv) When the capping portion contacts the nanopore, thereby preventing the conjugate from moving further through the nanopore, the polynucleotide-treated protein is allowed to transiently unbind from the polynucleotide, causing the conjugate to move through the nanopore under applied force in a direction opposite to the direction of movement controlled by the polynucleotide-treated protein; and

[0179] v) Optionally repeat steps (ii) to (iv) to cause the polypeptide to oscillate through the nanopore.

[0180] Replacement unit

[0181] As described above and not bound by theory, the inventors have discovered that the length of characterizable peptides is generally improved when using nanopores with longer tubes or channels compared to those with shorter tubes or channels. As explained above, this is believed to be related to or determined by the “RED” distance, such as... Figure 8 As shown in A(A).

[0182] Therefore, in some embodiments, the nanopore is modified to extend the distance between the polynucleotide-treated protein and the contraction region of the nanopore. The nanopore is typically modified to extend the distance between the polynucleotide-treated protein and the contraction region of the nanopore, as determined, when the polynucleotide-treated protein is used to control the movement of the conjugate relative to the nanopore. In some embodiments, the polynucleotide-treated protein is in a “seated position” in contact with the nanopore when used to control the movement of the conjugate relative to the nanopore, for example, in contact with the cis or trans opening of the nanopore. This is described in more detail herein and Figure 8 A(B) is shown schematically.

[0183] In some embodiments, the distance between the active site of the polynucleotide-treated protein and the nanopore can be extended by using a replacement unit. Figure 8 A(C) schematically illustrates the use of the replacement unit. Therefore, in some embodiments, the methods provided herein include providing a replacement unit. In some embodiments, the replacement unit is used to separate a peptide-treated protein from a nanopore, thereby extending the distance between the polynucleotide-treated protein and the nanopore.

[0184] In such embodiments, any suitable substitution unit can be used. For example, the substitution unit can be provided as a protein.

[0185] Any suitable protein can be used as a replacement unit. Exemplary proteins include proteins that adopt a cyclic conformation, such as as polymers, and which can therefore be readily located at the entrance of a nanopore. Many suitable cyclic proteins are known in the art, including nanopores, helicases (e.g., T7 helicase) and variants thereof as described herein. The replacement unit itself is not required to have any activity. In some embodiments, the replacement unit does not provide any significant distinction between polynucleotides or peptides in the conjugate.

[0186] In some embodiments, the substitution unit may comprise one or more polynucleotide-treated proteins or their inactive variants. This is schematically illustrated in... Figure 8 In section B, as shown in the figure, the polynucleotide processing protein (E1) controls the movement of the conjugate relative to the nanopore. The polynucleotide processing proteins E2...En initially contact the polypeptide moiety of the conjugate and therefore do not control the movement of the conjugate relative to the nanopore; however, they displace the polynucleotide processing protein E1 from the nanopore, thus increasing RED.

[0187] In such embodiments, the polynucleotide processing protein used as a replacement unit may be the same as or different from the polynucleotide processing protein used to control conjugate movement. In some embodiments, the polynucleotide processing protein used as a replacement unit is formed from an inactive variant of the same polynucleotide processing protein used to control conjugate movement.

[0188] One or more substitutional units, such as those described herein, may be connected to a nanopore (e.g., covalently or non-covalently coupled to the nanopore). Alternatively, one or more substitutional units may be associated with the nanopore, for example, by using polynucleotides to control their position relative to the nanopore.

[0189] In other embodiments, the polynucleotide-treated protein is modified to extend the distance from the active site of the polynucleotide-treated protein to the nanopore. Polynucleotide-treated proteins are typically modified to extend the distance between the active site of the polynucleotide-treated protein and the nanopore, as measured when the polynucleotide-treated protein is used to control the movement of the conjugate relative to the nanopore. This is described in more detail herein.

[0190] polypeptide

[0191] As explained above, the disclosed method includes characterizing the target peptide within the conjugate as the conjugate moves relative to the nanopore.

[0192] Any suitable polypeptide can be characterized using the disclosed methods.

[0193] In some embodiments, the target polypeptide is an unmodified protein or a portion thereof, or a naturally occurring polypeptide or a portion thereof.

[0194] In some embodiments, the target peptide is secreted by the cell. Alternatively, the target peptide may be produced intracellularly, necessitating extraction from the cell for characterization by the disclosed methods. The peptide may include cellular expression products of plasmids, such as plasmids used to clone proteins according to the methods described in Sambrook et al., Molecular Cloning: A Laboratory Manual, 4th ed., Cold Spring Harbor, Plainview, New York (2012); and Ausubel et al., Laboratory Guide to Contemporary Molecular Biology (Supplement 114), John Wiley & Son, New York (2016).

[0195] Polypeptides can be obtained or extracted from any organism or microorganism. Polypeptides can be obtained from humans or animals, such as from urine, lymph, saliva, mucus, semen, or amniotic fluid, or from whole blood, plasma, or serum. Polypeptides can also be obtained from plants, such as grains, legumes, fruits, or vegetables.

[0196] The target peptide can be provided as an impure mixture of one or more peptides and one or more impurities. Impurities may include a truncated form of the target peptide, distinct from the “target peptide” used to characterize it in the disclosed methods. For example, the target peptide may be a full-length protein and the impurity may include a portion of the protein. Impurities may also include proteins other than the target protein, such as proteins that can be co-purified from cell cultures or obtained from a sample.

[0197] Peptides can include any combination of any amino acid, amino acid analogs, and modified amino acids (i.e., amino acid derivatives). The amino acids (and derivatives, analogs, etc.) in a peptide can be distinguished by their physical size and charge.

[0198] Amino acids / derivatives / analytes can be naturally occurring or artificial.

[0199] In some embodiments, the polypeptide may comprise any naturally occurring amino acid. Twenty amino acids are encoded by a universal genetic code. These codes are: alanine (A), arginine (R), asparagine (N), aspartic acid (D), cysteine ​​(C), glutamic acid (E), glutamine (Q), glycine (G), histidine (H), isoleucine (I), leucine (L), lysine (K), methionine (M), phenylalanine (F), proline (P), serine (S), threonine (T), tryptophan (W), tyrosine (Y), and valine (V). Other naturally occurring amino acids include selenocysteine ​​and pyrrolidone.

[0200] In some embodiments, the peptide is modified. In some embodiments, the peptide is modified for detection using the disclosed method. In some embodiments, the disclosed method is used to characterize modifications in a target peptide.

[0201] In some embodiments, one or more amino acids / derivatives / analogs in the polypeptide are modified. In some embodiments, one or more amino acids / derivatives / analogs in the polypeptide are post-translational modifications. Therefore, the methods disclosed herein can be used to detect the presence, absence, location, and quantity of post-translational modifications in a polypeptide. The disclosed methods can be used to characterize the degree to which a polypeptide has been post-translationally modified.

[0202] Peptides can contain one or more post-translational modifications. Typical post-translational modifications include modification with hydrophobic groups, modification with cofactors, addition of chemical groups, glycosylation (non-enzymatic linking of sugars), biotinylation, and PEGylation. Post-translational modifications can also be non-natural, meaning they are chemical modifications performed in the laboratory for biotechnological or biomedical purposes. This allows for monitoring the levels of laboratory-made peptides, polypeptides, or proteins compared to their natural counterparts.

[0203] Examples of post-translational modifications using hydrophobic groups include: myristylation, myristate linkage, C14 saturated acid; palmitoylation, palmitate linkage, C16 saturated acid; isopreneation or isopentenylation, linkage of isoprene-like groups; farnesylation, linkage of farnesol groups; geraniol geraniolation, linkage of geraniol groups; and glypiation, as well as the formation of glypiation (GPI) anchors via amide bonds.

[0204] Examples of post-translational modifications using cofactors include esterification, linking of thioclate (C8) functional groups; flavinization, linking of flavin moieties (e.g., flavin mononucleotide (FMN) or flavin adenine dinucleotide (FAD)); linking of heme C, for example via a thioether bond with cysteine; phosphopantetheinylation, linking of 4'-phosphopantetheinyl thioethylamine; and retinyl Schiff base formation.

[0205] Examples of post-translational modifications by adding chemical groups include acylation, such as O-acylation (ester), N-acylation (amide), or S-acylation (thioester); acetylation, such as linking an acetyl group to an N-terminus or lysine; formylation; alkylation, adding an alkyl group, such as methyl or ethyl; methylation, such as adding a methyl group to lysine or arginine; amidation; butyrylation; γ-carboxylation; glycosylation, enzymatic linking a glycosyl group to, for example, arginine, asparagine, cysteine, hydroxylysine, serine, threonine, tyrosine, or tryptophan; polysialylation, linking polysialic acid; malonylation; hydroxylation; iodination; and bromination. Citrullination; nucleotide addition, linking of any nucleotide, such as any nucleotide discussed above; ADP ribosylation; oxidation; phosphorylation, linking of a phosphate group to, for example, serine, threonine, or tyrosine (O-linked) or histidine (N-linked); adenosineylation, linking of the adenosine moiety to, for example, tyrosine (O-linked) or histidine or lysine (N-linked); propionylation; pyroglutamic acid formation; S-glutathioneylation; threonylation; S-nitrosylation; succinylation, linking of a succinyl group to, for example, lysine; selenoylation, incorporation of selenium; and ubiquitination, addition of a ubiquitin subunit (N-linked).

[0206] Labeling peptides with molecular markers is within the scope of the methods provided herein. Molecular markers can be peptide modifications that facilitate the detection of peptides in the methods provided herein. For example, a marker can be a modification of the peptide that alters the signal obtained when characterizing the conjugate. For instance, a marker might interfere with the flux of ions through a nanopore. In this way, labeling can improve the sensitivity of the method.

[0207] In some embodiments, the polypeptide contains one or more cross-linking moieties, such as CC bridges. In some embodiments, the polypeptide is not cross-linked prior to being characterized using the disclosed methods.

[0208] In some embodiments, the polypeptide comprises sulfur-containing amino acids and therefore has the potential to form disulfide bonds. Typically, in such embodiments, the polypeptide is reduced using reagents such as DTT (dithiothreitol) or TCEP (tris(2-carboxyethyl)phosphine) before characterizing multiple peptides using the disclosed methods.

[0209] In some embodiments, the polypeptide is a full-length protein or a naturally occurring polypeptide. In some embodiments, the protein or naturally occurring polypeptide is fragmented prior to conjugation with a polynucleotide. In some embodiments, the protein or polypeptide is chemically or enzymatically fragmented. In some embodiments, the polypeptide or polypeptide fragments may be conjugated to form a longer target polypeptide.

[0210] The polypeptide can be of any suitable length. In some embodiments, the polypeptide is about 2 to about 300 peptide units in length. In some embodiments, the length of the polypeptide is about 2 to about 100 peptide units, for example, about 2 to about 50 peptide units, for example, about 2 to about 40 peptide units, such as about 2 to about 30 peptide units, for example, about 2 to about 25 peptide units, for example, about 2 to about 20 peptide units; or about 3 to about 50 peptide units, for example, about 3 to about 40 peptide units, such as about 3 to about 30 peptide units, for example, about 3 to about 25 peptide units, for example, about 3 to about 20 peptide units; or about 5 to about 50 peptide units, for example, about 5 to about 40 peptide units, such as about 5 to about 30 peptide units, such as about 5 to about 25 peptide units, for example, about 5 to about 20 peptide units; for example, about 7 to about 16 peptide units, such as about 9 to about 12 peptide units; or about 16 to about 25 peptide units, such as about 18 to about 22 peptide units.

[0211] The disclosed methods can characterize any number of peptides. For example, the methods may include characterizing 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 50, 100 or more peptides. If two or more peptides are used, they may be different peptides or two or more instances of the same peptide.

[0212] Therefore, it is apparent that the measurements performed in the disclosed methods are typically specific to the polypeptide selected from one or more of the following characteristics: (i) the length of the polypeptide; (ii) the identity of the polypeptide; (iii) the sequence of the polypeptide; (iv) the secondary structure of the polypeptide; and (v) whether the polypeptide is modified. In typical embodiments, the measurements are specific to the sequence of the polypeptide, or to whether the polypeptide is modified, for example, by one or more post-translational modifications. In some embodiments, the measurements are characteristics of the polypeptide's sequence.

[0213] In some embodiments, the polypeptide is in a relaxed form. In some embodiments, the polypeptide is held in a linearized form. Holding the polypeptide in a linearized form facilitates characterization of the polypeptide on a residue-by-residue basis because it prevents the polypeptide from “aggregating” within the nanopores.

[0214] Any suitable method can be used to maintain the polypeptide in a linearized form.

[0215] For example, if a polypeptide is charged, it can be kept in a linearized form by applying a voltage.

[0216] If a polypeptide is uncharged or only weakly charged, the charge can be altered or controlled by adjusting the pH. For example, by increasing the relative negative charge of the polypeptide using a high pH, ​​the polypeptide can be maintained in a linearized form. Increasing the negative charge of the polypeptide allows it to be maintained in a linearized form, for example, under a positive voltage. Alternatively, by increasing the relative positive charge of the polypeptide using a low pH, the polypeptide can be maintained in a linearized form. Increasing the positive charge of the polypeptide allows it to be maintained in a linearized form, for example, under a negative voltage. In the disclosed method, polynucleotide treatment of the protein is used to control the movement of the polynucleotide relative to the nanopore. Since polynucleotides are generally negatively charged, it is generally best to increase the linearization of the polypeptide by increasing the pH, thereby making the polypeptide more negatively charged, as is the case with polynucleotides. In this way, the conjugate retains the total negative charge and can therefore move readily, for example, under an applied voltage.

[0217] Peptides can be maintained in a linearized form by using suitable denaturing conditions. Suitable denaturing conditions include, for example, the presence of an appropriate concentration of a denaturing agent, such as guanidine hydrochloride and / or urea. The concentration of such denaturing agents used in the disclosed methods depends on the target peptide to be characterized in the methods and can be readily selected by those skilled in the art.

[0218] The peptide can be maintained in a linearized form by using a suitable detergent. Detergents suitable for the disclosed method include SDS (sodium dodecyl sulfate).

[0219] By performing the disclosed method at elevated temperatures, peptides can be maintained in a linearized form. Elevated temperatures overcome intrachain bonds and allow the peptides to assume a linearized form.

[0220] By performing the disclosed method under strong electroosmotic forces, peptides can be maintained in a linearized form. Such forces can be provided by using asymmetric salt conditions and / or by providing a suitable charge within the channels of the nanopore. The charge within the protein nanopore channels can be altered, for example, by mutagenesis. Changing the charge of the nanopore is entirely within the capabilities of those skilled in the art. When a voltage is applied across the nanopore, altering the charge of the nanopore generates a strong electroosmotic force due to the unbalanced flow of cations and anions through the nanopore.

[0221] By allowing peptides to pass through structures such as arrays of nanopillars, through nanoslits, or across nanogaps, peptides can be maintained in a linearized form. In some embodiments, the physical constraints of such structures can force peptides into a linearized form.

[0222] Formation of conjugates

[0223] As explained in more detail in this article, conjugates include polynucleotides conjugated to target peptides.

[0224] The target peptide can be conjugated to a polynucleotide at any suitable position. For example, the peptide can be conjugated to a polynucleotide at the N-terminus or C-terminus of the peptide. The peptide can also be conjugated to a polynucleotide via side chain groups of residues (e.g., amino acid residues) in the peptide.

[0225] In some embodiments, the target polypeptide has naturally occurring reactive functional groups that can be used to facilitate conjugation with polynucleotides. For example, cysteine ​​residues can be used to form disulfide bonds with polynucleotides or modified groups thereon.

[0226] In some embodiments, the target peptide is modified to facilitate its conjugation to a polynucleotide. For example, in some embodiments, the peptide is modified by linking a portion comprising a reactive functional group for linking to the polynucleotide. For example, in some embodiments, the peptide may extend one or more residues (e.g., amino acid residues) at the N-terminus or C-terminus, said one or more residues comprising one or more reactive functional groups for reacting with the corresponding reactive functional group on the polynucleotide. For example, in some embodiments, the peptide may extend one or more cysteine ​​residues at the N-terminus and / or C-terminus. Such residues can be used for linking to the polynucleotide portion of the conjugate, for example, through maleimide chemistry (e.g., by reacting cysteine ​​with an azide-maleimide compound (such as azide-[Pol]-maleimide, where [Pol] is typically a short-chain polymer, such as PEG, e.g., PEG2, PEG3, or PEG4); then coupled with a suitably functionalized polynucleotide, such as a polynucleotide with a BCN group, for reaction with the azide). Such chemistry is described in Example 2. To avoid ambiguity, when a polypeptide includes appropriate naturally occurring residues at its N and / or C ends (e.g., naturally occurring cysteine ​​residues at the N and / or C ends), such residues can be used for linking with polynucleotides.

[0227] In some embodiments, residues in the target peptide are modified to facilitate the linkage of the target peptide to a polynucleotide. In some embodiments, residues in the peptide (e.g., amino acid residues) are chemically modified for linkage to a polynucleotide. In some embodiments, residues in the peptide (e.g., amino acid residues) are enzymatically modified for linkage to a polynucleotide.

[0228] The conjugation chemistry between polynucleotides and polypeptides in conjugates is not specifically limited. Any suitable combination of reactive functional groups can be used. Many suitable reactive groups and their chemical targets are known in the art. Some exemplary reactive groups and their corresponding targets include aryl azides that can react with amines, carbodiimides that can react with amines and carboxyl groups, acyl hydrazides that can react with carbohydrates, hydroxymethylphosphine that can react with amines, imine esters that can react with amines, isocyanates that can react with hydroxyl groups, carbonyl groups that can react with hydrazine groups, maleimides that can react with thiol groups, NHS-esters that can react with amines, PFP-esters that can react with amines, psoralen that can react with thymine, pyridyl disulfides that can react with thiol groups, vinyl sulfones that can react with thiol amines and hydroxyl groups, vinyl sulfonamides, etc.

[0229] Other suitable chemistry for conjugating peptides to polynucleotides includes click chemistry. Many suitable click chemical reagents are known in the art. Suitable examples of click chemistry include, but are not limited to, the following:

[0230] (a) Copper (I)-catalyzed azide-alkyne cycloaddition (azide-alkyne Huisgen cycloaddition);

[0231] (b) Strain-promoted azide-alkyne cycloaddition; including olefin and azide [3+2] cycloaddition; olefin and tetrazine reverse demand Diels-Alder reaction; and olefin and tetrazolium photoclick reaction;

[0232] (c) Copper-free variants of 1,3-dipolar cycloaddition reactions in which azides react with alkynes under strain, for example in cyclooctane rings, such as in bicyclic [6.1.0]nonyne (BCN);

[0233] (d) The reaction between the oxygen nucleophile on one linker and the reactive moiety of the epoxide or aziridine on the other linker; and

[0234] (e) Staudinger ligation, in which the alkyne moiety can be replaced by arylphosphine, resulting in a specific reaction with the azide to give an amide bond.

[0235] Any reactive group can be used to form conjugates. Some suitable reactive groups include [1,4-bis[3-(2-pyridyldithio)propamido]butane; 1,1,1-bis-maleimide triethylene glycol; 3,3'-dithiodipropionate di(N-hydroxysuccinimide); ethylene glycol-bis(succinate N-hydroxysuccinimide); 4,4'-diisothiocyanate stilbene-2,2'-disulfonic acid disodium salt; bis[2-(4-azidosalicylic acid amino)ethyl] disulfide; 3-(2-pyridyldithio)propionate N-hydroxysuccinimide; 4-maleimidebutyrate N-hydroxysuccinimide; iodoacetic acid N-hydroxysuccinimide; S-acetylthioacetic acid N-hydroxysuccinimide; azide-PEG-maleimide; and alkyne-PEG-maleimide. The reactive group can be any of those groups disclosed in WO 2010 / 086602, and particularly in Table 3 of this application.

[0236] In some embodiments, prior to the conjugation step, the reactive functional group is included in the polynucleotide and the target functional group is included in the polypeptide. In other embodiments, prior to the conjugation step, the reactive functional group is included in the polypeptide and the target functional group is included in the polynucleotide. In some embodiments, the reactive functional group is directly linked to the polypeptide. In some embodiments, the reactive functional group is linked to the polypeptide via a spacer. Any suitable spacer can be used. Suitable spacers include, for example, alkyl diamines, such as ethyldiamine, etc.

[0237] As will be apparent from the above discussion, in some embodiments, the conjugate comprises multiple polypeptide moieties and / or multiple polynucleotide moieties. For example, the conjugate may comprise a structure of the form ...-PNPNPN..., where P is a polypeptide and N is a polynucleotide. In such embodiments, the polynucleotide-treated protein sequentially controls the movement of the N moieties of the conjugate relative to the nanopore and thus sequentially controls the movement of the P moieties relative to the nanopore, thereby allowing for sequential characterization of the P moieties. In such embodiments, multiple polynucleotides and polypeptides may be conjugated together using the same or different chemical conjugates.

[0238] As explained herein, conjugates may include a leader sequence. As explained herein, any suitable leader sequence may be used. In some embodiments, the leader sequence is a polynucleotide. In embodiments where the leader sequence is a polynucleotide, the leader sequence may be the same type of polynucleotide as the polynucleotide used in the conjugate, or the leader sequence may be a different type of polynucleotide. For example, the polynucleotide in the conjugate may be DNA, and the leader sequence may be RNA, or vice versa.

[0239] In some embodiments, the leader sequence is a charged polymer, such as a negatively charged polymer. In some embodiments, the leader sequence includes a polymer, such as PEG or a polysaccharide. In such embodiments, the length of the leader sequence can be from 10 to 150 monomer units (e.g., ethylene glycol or sugar units), such as 20 to 120, such as 30 to 100, such as 40 to 80, such as 50 to 70 monomer units (e.g., ethylene glycol or sugar units).

[0240] Polynucleotides

[0241] As explained in more detail herein, the methods presented herein include conjugating peptides to polynucleotides and using polynucleotide-treated proteins to control the movement of the conjugates relative to nanopores.

[0242] In the disclosed method, any suitable polynucleotide can be used.

[0243] In some embodiments, the polynucleotide is secreted by the cell. Alternatively, the polynucleotide can be produced intracellularly, making it necessary to extract the polynucleotide from the cell for use in the disclosed method.

[0244] Polynucleotides can be provided as an impure mixture of one or more polynucleotides and one or more impurities. Impurities may include truncated forms of polynucleotides that differ from the polynucleotides used to form the conjugate. For example, the polynucleotide used to form the conjugate may be genomic DNA, and impurities may include portions of genomic DNA, plasmids, etc. Target polynucleotides may be coding regions of genomic DNA, and undesirable polynucleotides may include non-coding regions of DNA.

[0245] Examples of polynucleotides include DNA and RNA. The bases in DNA and RNA can be distinguished by their physical size.

[0246] Polynucleotides, or nucleic acids, can include any combination of any nucleotides. Nucleotides can be naturally occurring or artificial. One or more nucleotides in a polynucleotide can be oxidized or methylated. One or more nucleotides in a polynucleotide can be damaged. For example, polynucleotides can include pyrimidine dimers. Such dimers are commonly associated with UV damage and are a major cause of melanoma.

[0247] One or more nucleotides in a polynucleotide may be modified, for example, by a tag or label, suitable examples of which are known to those skilled in the art. A polynucleotide may include one or more spacers. Adaptors, such as sequencing adaptors, may be included in a polynucleotide. Adaptors, tags, and spacers are described in more detail herein.

[0248] Examples of modified bases are disclosed herein and can be incorporated into polynucleotides by means known in the art, such as by polymerase incorporation of modified nucleotide triphosphates during chain replication (e.g., in PCR) or by polymerase-filled methods. In some embodiments, one or more bases may be chemically modified using reagents known in the art.

[0249] Nucleotides typically contain a nucleobase, a sugar, and at least one phosphate group. The nucleobase and sugar form a nucleoside. The nucleobase is typically heterocyclic. Nucleobases include, but are not limited to, purines and pyrimidines, and more specifically, adenine (A), guanine (G), thymine (T), uracil (U), and cytosine (C). The sugar is typically a pentose sugar. Nucleotide sugars include, but are not limited to, ribose and deoxyribose. The sugar is preferably deoxyribose. Polynucleotides preferably include the following nucleosides: deoxyadenosine (dA), deoxyuridine (dU), and / or thymidine (dT), deoxyguanosine (dG), and deoxycytidine (dC). Nucleotides are typically ribonucleotides or deoxyribonucleotides. Nucleotides typically contain monophosphate, diphosphate, or triphosphate. Nucleotides may include more than three phosphates, such as four or five phosphates. Phosphates may be attached to the 5' or 3' side of the nucleotide. Nucleotides in polynucleotides may be linked to each other in any manner. Nucleotides are typically linked by their sugar and phosphate groups, as in nucleic acids. Nucleotides can also be linked by their nucleobases, as in pyrimidine dimers.

[0250] Polynucleotides can be double-stranded or single-stranded.

[0251] In some embodiments, the polynucleotide is single-stranded DNA. In some embodiments, the polynucleotide is single-stranded RNA. In some embodiments, the polynucleotide is a single-stranded DNA-RNA hybrid. DNA-RNA hybrids can be prepared by linking single-stranded DNA to RNA, or vice versa. Polynucleotides are most typically single-stranded deoxyribonucleic acid (DNA) or single-stranded ribonucleic acid (RNA).

[0252] In some embodiments, the polynucleotide is double-stranded DNA. In some embodiments, the polynucleotide is double-stranded RNA. In some embodiments, the polynucleotide is a double-stranded DNA-RNA hybrid. Double-stranded DNA-RNA hybrids can be prepared from single-stranded RNA by reverse transcription of cDNA complement.

[0253] Polynucleotides can be of any length. For example, the length of a polynucleotide can be at least 10, at least 50, at least 100, at least 150, at least 200, at least 250, at least 300, at least 400, or at least 500 nucleotides or nucleotide pairs. The length of a polynucleotide can be 1000 or more nucleotides or nucleotide pairs, 5000 or more nucleotides or nucleotide pairs, or 100,000 or more nucleotides or nucleotide pairs.

[0254] More typically, polynucleotides are about 1 to about 10,000 nucleotides or nucleotide pairs in length, such as about 1 to about 1,000 nucleotides or nucleotide pairs (e.g., about 10 to about 1,000 nucleotides or nucleotide pairs), such as about 5 to about 500 nucleotides or nucleotide pairs, such as about 10 to about 100 nucleotides or nucleotide pairs, such as about 20 to about 80 nucleotides or nucleotide pairs, such as about 30 to about 50 nucleotides or nucleotide pairs.

[0255] Any number of polynucleotides may be used in the disclosed methods. For example, the methods may include the use of 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 50, 100 or more polynucleotides. If two or more polynucleotides are used, they may be different polynucleotides or two instances of the same polynucleotide. Polynucleotides may be naturally occurring or artificial.

[0256] Nucleotides can have any identity and include, but are not limited to, adenosine monophosphate (AMP), guanosine monophosphate (GMP), thymidine monophosphate (TMP), uridine monophosphate (UMP), 5-methylcytidine monophosphate, 5-hydroxymethylcytidine monophosphate, cytidine monophosphate (CMP), cyclic adenosine monophosphate (cAMP), cyclic guanosine monophosphate (cGMP), deoxyadenosine monophosphate (dAMP), deoxyguanosine monophosphate (dGMP), deoxythymidine monophosphate (dTMP), deoxyuridine monophosphate (dUMP), deoxycytidine monophosphate (dCMP), and deoxymethylcytidine monophosphate. The nucleotide is preferably selected from AMP, TMP, GMP, CMP, UMP, dAMP, dTMP, dGMP, dCMP, and dUMP. Nucleotides can be baseless (i.e., lacking a nucleobase). Nucleotides can also lack a nucleobase and a sugar (i.e., are C3 spacers).

[0257] Polynucleotides can include products of PCR reactions, genomic DNA, products of endonuclease digestion, and / or DNA libraries. Polynucleotides can be obtained or extracted from any organism or microorganism. Polynucleotides can be obtained from humans or animals, for example from urine, lymph, saliva, mucus, semen, or amniotic fluid, or from whole blood, plasma, or serum. Polynucleotides can be obtained from plants, such as cereals, legumes, fruits, or vegetables. Polynucleotides can include genomic DNA. Genomic DNA can be fragmented. DNA can be fragmented by any suitable method. For example, methods for fragmenting DNA are known in the art, and such methods can use transposases, such as MuA transposase. Typically, genomic DNA is not fragmented.

[0258] Labeling polynucleotides with molecular markers is within the scope of the methods provided herein. Molecular markers can be polynucleotide modifications that facilitate the detection of polynucleotides or conjugates in the methods provided herein. For example, a marker can be a modification of the polynucleotide that alters the signal obtained when characterizing the conjugate. For instance, a marker might interfere with the flux of ions through a nanopore. In this way, labeling can improve the sensitivity of the method.

[0259] connector

[0260] In some embodiments of the methods provided herein, the polynucleotide has a polynucleotide adaptor to which it is attached. The adaptor typically comprises a polynucleotide chain capable of attaching to the end of the polynucleotide.

[0261] In some embodiments, the adaptor is linked to a polynucleotide prior to forming a conjugate with the polypeptide. In some embodiments, the adaptor is linked to a conjugate of the polynucleotide and the polypeptide.

[0262] Therefore, in some embodiments, the method includes linking an adaptor (e.g., an adaptor as described herein) to a polynucleotide and forming a conjugate by conjugating the polynucleotide / adaptor construct to a target peptide. In some embodiments, the conjugate is formed by linking an adaptor (e.g., an adaptor as described herein) to a polynucleotide and by linking the adaptor to a target peptide.

[0263] In some embodiments, the adaptor may be selected or modified to provide a specific site for conjugation with a polynucleotide.

[0264] An adaptor may be attached to only one end of a polynucleotide or conjugate. A polynucleotide adaptor may be added to both ends of a polynucleotide or conjugate. Alternatively, different adaptors may be added to both ends of a polynucleotide or conjugate.

[0265] An adaptor can be added to both strands of a double-stranded polynucleotide. An adaptor can also be added to a single-stranded polynucleotide. Methods for adding an adaptor to a polynucleotide are known in the art. The adaptor can be linked to the polynucleotide, for example, by ligation, by click chemistry, by labeling, by topoisomerization, or by any other suitable method.

[0266] In one embodiment, the adaptor or each adaptor is synthetic or artificial. Typically, the adaptor or each adaptor comprises a polymer as described herein. In some embodiments, the adaptor or each adaptor comprises a spacer as described herein. In some embodiments, the adaptor or each adaptor comprises a polynucleotide. The polynucleotide adaptor or each polynucleotide adaptor may comprise DNA, RNA, modified DNA (e.g., base-free DNA), RNA, PNA, LNA, BNA, and / or PEG. Typically, the adaptor or each adaptor comprises single-stranded and / or double-stranded DNA or RNA. The adaptor may comprise a polynucleotide of the same type as the polynucleotide chain it is linked to. The adaptor may comprise a polynucleotide of a different type than the polynucleotide chain it is linked to. In some embodiments, the polynucleotide chain used in the disclosed methods is a single-stranded DNA chain and the adaptor comprises DNA or RNA, typically single-stranded DNA. In some embodiments, the polynucleotide is a double-stranded DNA chain and the adaptor comprises DNA or RNA, for example, double-stranded or single-stranded DNA.

[0267] In some embodiments, the adapter may be a bridging portion. The bridging portion may be used to connect the two strands of a double-stranded polynucleotide. For example, in some embodiments, the bridging portion is used to connect the template strand of the double-stranded polynucleotide to its complementary strand.

[0268] The bridging portion typically covalently links the two strands of a double-stranded polynucleotide. The bridging portion can be anything capable of linking the two strands of a double-stranded polynucleotide, provided that it does not interfere with the movement of the polynucleotide relative to the nanopore. Suitable bridging portions include, but are not limited to, polymeric linkers, chemical linkers, polynucleotides, or polypeptides. Preferably, the bridging portion includes DNA, RNA, modified DNA (such as base-free DNA), RNA, PNA, LNA, or PEG. More preferably, the bridging portion is DNA or RNA.

[0269] In some embodiments, the bridging portion is a hairpin receptacle. A hairpin receptacle is an receptacle comprising a single polynucleotide chain, wherein the ends of the polynucleotide chains are capable of hybridizing to or being hybridized to each other, and wherein the middle segment of the polynucleotide forms a loop. Suitable hairpin receptacles can be designed using methods known in the art. In some embodiments, the length of the hairpin loop is typically 4 to 100 nucleotides, for example, 4 to 50 nucleotides, such as 4 to 20 nucleotides, for example, 4 to 8 nucleotides. In some embodiments, the bridging portion (e.g., the hairpin receptacle) is attached to one end of the double-stranded polynucleotide. The bridging portion (e.g., the hairpin receptacle) is typically not attached to both ends of the double-stranded polynucleotide.

[0270] In some embodiments, the adaptor is a linear adaptor. A linear adaptor can bind to either or both ends of a single-stranded polynucleotide. When the polynucleotide is a double-stranded polynucleotide, the linear adaptor can bind to either or both ends of either strand or both strands of the double-stranded polynucleotide. The linear adaptor may include a leader sequence as described herein. The linear adaptor may include a portion for hybridization with a tag (such as a pore tag) as described herein. The length of the linear adaptor can be from 10 to 150 nucleotides, such as 20 to 120, 30 to 100, 40 to 80, or 50 to 70 nucleotides. The linear adaptor can be single-stranded. The linear adaptor can be double-stranded.

[0271] In some embodiments, the adaptor may be a Y-adaptor. A Y-adaptor is typically a polynucleotide adaptor. A Y-adaptor is typically double-stranded and includes (a) a region at one end where the two strands hybridize, and (b) a region at the other end where the two strands are not complementary. The non-complementary portions of the strands typically form overhangs. The presence of non-complementary regions in the Y-adaptor gives it a Y-shape because the two strands typically do not hybridize with each other as they do in a double-stranded portion. The two single-stranded portions of the Y-adaptor may be of the same length or different lengths. For example, one single-stranded portion of the Y-adaptor may be 10 to 150 nucleotides long, such as 20 to 120, 30 to 100, 40 to 80, or 50 to 70 nucleotides long, and the other single-stranded portion of the Y-adaptor may independently be 10 to 150 nucleotides long, such as 20 to 120, 30 to 100, 40 to 80, or 50 to 70 nucleotides long. The length of the double-stranded "stem" portion of the Y-connector can be, for example, 10 to 150 nucleotides, such as 20 to 120, 30 to 100, 40 to 80, or 50 to 70 nucleotides.

[0272] The adaptor can be linked to the target polynucleotide by any suitable method known in the art. The adaptor can be synthesized separately and linked to the target polynucleotide by chemical linkage or enzymatic linkage. Alternatively, the adaptor can be generated during the processing of the target polynucleotide. In some embodiments, the adaptor is linked to the target polynucleotide at or near one end of the target polynucleotide. In some embodiments, the adaptor is linked to the target polynucleotide within 50 nucleotides, such as within 20 nucleotides, or even within 10 nucleotides from the end of the target polynucleotide. In some embodiments, the adaptor is linked to the target polynucleotide at its end. When the adaptor is linked to the target polynucleotide, the adaptor may comprise a nucleotide of the same type as the target polynucleotide or may comprise a nucleotide different from the target polynucleotide.

[0273] Integrators particularly suitable for the disclosed methods may include linear homopolymer regions (e.g., about 5 to about 20 nucleotides, such as about 10 to about 30 nucleotides, e.g., thymine or cytidine) and / or hybridization sites (as described in more detail herein) for hybridization with one or more tandem chains or anchors. Such integrators may also include reactive functional groups for binding to the target peptide. Click chemistry groups are particularly suitable in this regard. For example, exemplary groups included in the integrator include groups that can form particles in copper-free click chemistry, such as groups based on BCN (bicyclo[6.1.0]nonyne) and its derivatives, dibenzocyclooctyne (DBCO) groups, etc. Reactivity of such groups is well known in the art. For example, BCN groups typically react with groups such as azides, tetrazines, and nitrones, which can be incorporated into peptides, for example. DBCO groups are highly reactive to azide groups. Other particularly suitable chemical groups include 2-pyridinecarboxylaldehyde (2-PCA) groups and their derivatives. For example, 6-(azidomethyl)-2-pyridinecarboxylaldehyde can react with the N-terminal amino group of a peptide.

[0274] spacer

[0275] In some embodiments of the methods provided herein, the polynucleotide, conjugate formed by the reaction of a polynucleotide with a polypeptide, or the adaptor described herein may include spacers. For example, one or more spacers may be present in the polynucleotide adaptor. For example, the polynucleotide adaptor may include one to about 20 spacers, for example, about 1 to about 10, for example, 1 to about 5 spacers, for example, 1, 2, 3, 4, or 5 spacers. Spacers may include any suitable number of spacer units. Spacers can provide an energy barrier that impedes the movement of the polynucleotide-treated protein. For example, spacers can impede the polynucleotide-treated protein by reducing the traction force of the polynucleotide-treated protein on the polynucleotide. This can be achieved, for example, by using a base-free spacer, i.e., a spacer in which a base has been removed from one or more nucleotides in the polynucleotide adaptor. Spacers can physically prevent the movement of the polynucleotide-treated protein, for example, by introducing a large chemical group to physically impede the movement of the polynucleotide-treated protein.

[0276] In some embodiments, one or more spacers are included in the polynucleotide or conjugate or adaptor used in the method claimed herein, in order to provide a unique signal as the polynucleotide or conjugate or adaptor passes through or across the nanopore, i.e., as the polynucleotide or conjugate or adaptor moves relative to the nanopore.

[0277] In some embodiments, spacers may comprise linear molecules, such as polymers. Typically, such spacers have a structure different from the polynucleotide used in the conjugate. For example, if the polynucleotide is DNA, then the spacer or each spacer does not contain DNA. In particular, if the polynucleotide is deoxyribonucleic acid (DNA) or ribonucleic acid (RNA), then the spacer or each spacer preferably comprises peptide nucleic acid (PNA), glycerol nucleic acid (GNA), threonine nucleic acid (TNA), locked nucleic acid (LNA), or a synthetic polymer having nucleotide side chains. In some embodiments, the spacer may include one or more nitroindoles, one or more inosines, one or more acridines, one or more 2-aminopurines, one or more 2-6-diaminopurines, one or more 5-bromo-deoxyuridines, one or more reverse thymidines (reverse dT), one or more reverse dideoxythymidines (ddT), one or more dideoxycytidines (ddC), one or more 5-methylcytidines, one or more 5-hydroxymethylcytidines, one or more 2'-O-methylRNA bases, one or more isodeoxycytidines (Iso-dC), one or more isodeoxyguanosines (Iso-dG), one or more C3 (OC3H6OPO3) groups, one or more optically cleavable (PC) [OC3H6-C(O)NHCH2-C6H3NO2-CH(CH3)OPO3] groups, one or more hexanediol groups, one or more spacer 9 (iSp9) [(OCH2CH2)3OPO3] groups or one or more spacer 18 (iSp18) [(OCH2CH2)6OPO3] groups; or one or more thiol linkages. Spacers can include any combination of these groups. Many of these groups can be derived from... (Integrated DNA Commercially available. For example, C3, iSp9, and iSp18 spacers can all be obtained from... Obtained. Spacers may include any number of the above-mentioned groups as spacer units.

[0278] In some embodiments, the spacer may include one or more chemical groups that cause polynucleotide processing protein stagnation. In some embodiments, suitable chemical groups are one or more chemical side groups. One or more chemical groups may be linked to one or more nucleobases in the polynucleotide, construct, or adaptor. One or more chemical groups may be linked to the backbone of the polynucleotide adaptor. Any number of suitable chemical groups may be present, such as 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, or more. Suitable groups include, but are not limited to, fluorophores, streptavidin and / or biotin, cholesterol, methylene blue, dinitrophenol (DNP), digoxigenin and / or anti-digoxigenin and diphenylcyclooctynyl groups. In some embodiments, the spacer may include a polymer. In some embodiments, the spacer may include a polymer, said polymer being a polypeptide or polyethylene glycol (PEG).

[0279] Spacers may comprise one or more abasic nucleotides (i.e., nucleotides lacking a nucleobase), such as 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, or more abasic nucleotides. In the abasic nucleotides, the nucleobase may be replaced by -H (idSp) or -OH. Abasic spacers can be inserted into a target polynucleotide by removing a base from one or more adjacent nucleotides. For example, the polynucleotide may be modified to contain 3-methyladenine, 7-methylguanine, 1,N6-vinylidene adenine inosine, or hypoxanthine, and the nucleobase may be removed from these nucleotides using human alkyladenine DNA glycosidase (hAAG). Alternatively, the polynucleotide may be modified to contain uracil, and the nucleobase may be removed using uracil-DNA glycosidase (UDG). In one embodiment, one or more spacers do not include any abasic nucleotides.

[0280] A method for using spacers to arrest polynucleotide-processed proteins such as helicases on polynucleotide adaptors is described in WO 2014 / 135838, which is incorporated herein by reference in its entirety.

[0281] anchor

[0282] In some embodiments, polynucleotides, their conjugates with polypeptides, or adaptors linked to them may include, for example, membrane anchors or transmembrane pore anchors linked to the adaptor. In one embodiment, the anchor facilitates the characterization of the conjugate according to the methods disclosed herein. For example, membrane anchors or transmembrane pore anchors can facilitate the localization of the conjugate around nanopores in a membrane.

[0283] The anchor can be a peptide anchor and / or a hydrophobic anchor that can be inserted into the membrane. In one embodiment, the hydrophobic anchor is a lipid, fatty acid, sterol, carbon nanotube, peptide, protein, or amino acid, such as cholesterol, palmitate, or tocopherol. The anchor may include thiols, biotin, or surfactants.

[0284] On the one hand, the anchor can be biotin (for binding to streptavidin), amylose (for binding to maltose-binding proteins or fusion proteins), Ni-NTA (for binding to polyhistidine or polyhistidine-labeled proteins), or peptides (such as antigens).

[0285] In one embodiment, an anchor may include a linker, or two, three, four, or more linkers. Preferred linkers comprise, but are not limited to, polymers such as polynucleotides, polyethylene glycol (PEG), polysaccharides, and polypeptides. These linkers may be linear, branched, or cyclic. For example, a linker may be a cyclic polynucleotide. The linker may hybridize to a complementary sequence on the cyclic polynucleotide linker. One or more anchors or one or more linkers may include components that can be cleaved or broken down, such as restriction sites or photostable groups. Linkers may be functionalized with maleimide groups to link to cysteine ​​residues in a protein. Suitable linkers are described in WO 2010 / 086602.

[0286] In one embodiment, the anchor is a cholesterol or fatty acyl chain. For example, any fatty acyl chain having a length of 6 to 30 carbon atoms, such as hexadecanoic acid, can be used. Examples of suitable anchors and methods for connecting anchors to connectors are disclosed in WO 2012 / 164270 and WO 2015 / 150786.

[0287] Controlling the movement of conjugates relative to nanopores

[0288] As explained in more detail above, the methods provided herein include: contacting the conjugate with a polynucleotide-treated protein capable of controlling the movement of the polynucleotide relative to a nanopore; and performing one or more measurements specific to the peptide while the conjugate moves relative to the nanopore.

[0289] The movement of the conjugate relative to the nanopore can be driven by any suitable means. In some embodiments, the movement of the conjugate is driven by physical or chemical forces (potentials). In some embodiments, the physical force is provided by an electric potential (e.g., voltage potential) or a temperature gradient, etc.

[0290] In some embodiments, when a potential is applied across the nanopore, the conjugate moves relative to the nanopore. Polynucleotides are negatively charged, and therefore applying a voltage potential across the nanopore will cause the polynucleotide to move relative to the nanopore under the influence of the applied voltage potential. For example, if a positive voltage potential is applied relative to the cis side of the nanopore to the trans side, this will induce the negatively charged analyte to move from the cis side to the trans side. Similarly, if a positive voltage potential is applied relative to the cis side of the nanopore to the trans side, this will prevent the negatively charged analyte from moving from the trans side to the cis side. The opposite occurs if a negative voltage potential is applied relative to the cis side of the nanopore to the trans side. Apparatus and methods for applying appropriate voltages are described in more detail herein.

[0291] In some embodiments, the chemical force is provided by a concentration (e.g., pH) gradient.

[0292] In some embodiments, the polynucleotide-treated protein controls the movement of the conjugate in the same direction as a physical or chemical force (potential). For example, in some embodiments, a positive voltage is applied relative to the cis side of the nanopore towards the trans side, and the polynucleotide-treated protein controls the movement of the conjugate from the cis side to the trans side of the nanopore. In some embodiments, a positive voltage is applied relative to the trans side of the nanopore towards the cis side, and the polynucleotide-treated protein controls the movement of the conjugate from the trans side to the cis side of the nanopore.

[0293] In some embodiments, the polynucleotide-treated protein controls the movement of the conjugate in a direction opposite to a physical or chemical force (electric potential). For example, in some embodiments, a positive voltage is applied relative to the cis side of the nanopore towards the trans side, and the polynucleotide-treated protein controls the movement of the conjugate from the trans side to the cis side of the nanopore. In some embodiments, a positive voltage is applied relative to the trans side of the nanopore towards the cis side, and the polynucleotide-treated protein controls the movement of the conjugate from the cis side to the trans side of the nanopore.

[0294] In some embodiments, the movement of the conjugate is driven by a polynucleotide-treated protein in the absence of an applied potential.

[0295] In the disclosed method, the polynucleotide-treated protein is capable of controlling the movement of the polynucleotide relative to the nanopore. In other words, the polynucleotide-treated protein is capable of controlling the movement of the conjugate. In some embodiments, the polynucleotide-treated protein is capable of controlling the movement of both polynucleotides and peptides.

[0296] Suitable polynucleotide processing proteins are also referred to as motor proteins or polynucleotide processing enzymes. Suitable polynucleotide processing proteins are known in the art, and some exemplary polynucleotide processing proteins are described in more detail below.

[0297] In one embodiment, the motor protein is or is derived from a polynucleotide processing enzyme. A polynucleotide processing enzyme is a polypeptide capable of interacting with and modifying at least one property of a polynucleotide. The enzyme can modify a polynucleotide by cleaving it to form individual nucleotides or shorter nucleotide chains such as dinucleotides or trinucleotides. The enzyme can also modify a polynucleotide by orienting or moving it to a specific location.

[0298] In some embodiments, the polynucleotide-treated protein may be present on the conjugate prior to contact with the nanopore. For example, the polynucleotide-treated protein may be present on the polynucleotides within the conjugate. In some embodiments, the polynucleotide-treated protein is present on an adaptor that comprises a portion of the conjugate, or may be additionally present on a portion of the conjugate.

[0299] In some embodiments, when the portion of the conjugate in contact with the active site of the polynucleotide-treated protein comprises a polypeptide, the polynucleotide-treated protein is able to remain bound to the conjugate. In other words, in some embodiments, the polynucleotide-treated protein does not dissociate from the conjugate when it contacts the polypeptide portion of the conjugate. In some embodiments, the polynucleotide-treated protein moves freely relative to the polypeptide portion until it contacts one or more subsequent polynucleotide portions of the conjugate.

[0300] In some embodiments, the polynucleotide-treated protein is modified to prevent detachment from the conjugate, polynucleotide, or adaptor (except by crossing the end of the removed conjugate, polynucleotide, or adaptor) when the polynucleotide-treated protein comes into contact with a portion of the conjugate comprising a polypeptide. Such modified polynucleotide-treated proteins are particularly suitable for the disclosed methods.

[0301] Polynucleotide-treated proteins can be modified in any suitable manner. For example, a polynucleotide-treated protein can be loaded onto a polynucleotide, conjugate, or adaptor and then modified to prevent its detachment. Alternatively, a polynucleotide-treated protein can be modified to prevent its detachment before being loaded onto a polynucleotide, conjugate, or adaptor. Modification of polynucleotide-treated proteins to prevent their detachment from polynucleotides, conjugates, or adaptors can be achieved using methods known in the art, such as those discussed in WO 2014 / 013260 (hereinforced in its entirety by reference), and with particular reference to the paragraph describing the modification of polynucleotide-treated proteins (polynucleotide-binding proteins) (such as helicases) to prevent their detachment from the polynucleotide chain.

[0302] For example, a polynucleotide-treated protein may have a polynucleotide unbinding opening; for example, a cavity, crack, or gap through which the polynucleotide chain can pass when the polynucleotide-treated protein dissociates from the chain. In some embodiments, the polynucleotide unbinding opening of a given motor protein (polynucleotide-treated protein) can be determined by referring to its structure, such as its X-ray crystal structure. The X-ray crystal structure can be obtained in the presence and / or absence of a polynucleotide substrate. In some embodiments, the location of the polynucleotide unbinding opening in a given polynucleotide-treated protein can be inferred or confirmed by molecular modeling using standard packages known in the art. In some embodiments, the polynucleotide unbinding opening can be transiently generated by the movement of one or more portions of the polynucleotide-treated protein, such as one or more domains.

[0303] Polynucleotide-treated proteins (motor proteins) can be modified by closing polynucleotide unbinding openings. Therefore, closing polynucleotide unbinding openings prevents the polynucleotide-treated protein from dissociating from the polypeptide portion of the conjugate and from dissociating from the polynucleotide or adaptor. For example, motor proteins can be modified by covalently closing polynucleotide unbinding openings. In some embodiments, the motor protein used for addressing in this manner is a helicase as described herein. Thus, in some embodiments of the disclosed methods, the polynucleotide-treated protein is modified to completely or partially close openings present in at least one conformational state of the unmodified protein through which the polynucleotide chain can unbind.

[0304] Polynucleotide processing proteins can be selected or chosen based on the polynucleotides used in the conjugates characterized in the methods disclosed herein. Alternatively, polynucleotides can be selected or chosen based on the polynucleotide processing proteins used to control the movement of the conjugates. For example, when the polynucleotide is DNA, DNA motor proteins can typically be used. When the polynucleotide is RNA, RNA motor proteins can be used. When the polynucleotide is a hybrid of DNA and RNA, motor proteins capable of processing both DNA and RNA can be used.

[0305] In one embodiment, the motor protein is derived from any member of the enzyme classification (EC) group: 3.1.11, 3.1.13, 3.1.14, 3.1.15, 3.1.16, 3.1.21, 3.1.22, 3.1.25, 3.1.26, 3.1.27, 3.1.30, and 3.1.31.

[0306] In some embodiments, the motor protein is a helicase, polymerase, exonuclease, topoisomerase, or a variant thereof.

[0307] In one embodiment, the motor protein is an exonuclease. Suitable enzymes include, but are not limited to, exonuclease I (SEQ ID NO:1) from *Escherichia coli*, exonuclease III (SEQ ID NO:2) from *Escherichia coli*, RecJ (SEQ ID NO:3) from *T. thermophilus*, bacteriophage λ exonuclease (SEQ ID NO:4), TatD exonuclease, and variants thereof. The three subunits, including the sequence shown in SEQ ID NO:3, or variants thereof, interact to form a trimer exonuclease.

[0308] In one embodiment, the motor protein is a polymerase. The polymerase can be... 3173 DNA polymerase (which is commercially available) (Company), SD polymerase (commercially available) The enzyme is Klenow from NEB or a variant thereof. In one embodiment, the enzyme is Phi29 DNA polymerase (SEQ ID NO:5) or a variant thereof. A modified version of the Phi29 polymerase that can be used in the disclosed methods is disclosed in U.S. Patent No. 5,576,204.

[0309] In the embodiments provided herein, which include methods for controlling conjugate movement by synthesizing chains complementary to polynucleotides, the polynucleotide-treated protein is typically a polymerase, such as the polymerase described herein.

[0310] In one embodiment, the polynucleotide processing protein is a topoisomerase. In one embodiment, the topoisomerase is a member of either group 5.99.1.2 or 5.99.1.3 of the partial classification (EC). The topoisomerase can be a reverse transcriptase, which is an enzyme capable of catalyzing the formation of cDNA from an RNA template. These can be derived from, for example, New England... and Acquired through commercial purchase.

[0311] In one embodiment, the polynucleotide processing protein is a helicase. Any suitable helicase can be used according to the methods provided herein. For example, the motor protein used according to this disclosure, or each motor protein, can be independently selected from Hel308 helicase, RecD helicase, TraI helicase, TrwC helicase, XPD helicase, and Dda helicase, or variants thereof. Monomeric helicases can include several domains linked together. For example, TraI helicase and TraI subgroup helicases can contain two RecD helicase domains, a release enzyme domain, and a C-terminal domain. These domains typically form a monomeric helicase capable of functioning without forming oligomers. Specific examples of suitable helicases include Hel308, NS3, Dda, UvrD, Rep, PcrA, Pif1, and TraI. These helicases typically act on single-stranded DNA. Examples of helicases that can move along both strands of double-stranded DNA include FtfK and hexamethylenetetramer complexes, or multi-subunit complexes such as RecBCD. NS3 helicases are particularly suitable for the disclosed methods because they are capable of processing DNA and RNA, and therefore can be used in embodiments of the disclosed methods in which the target double-stranded nucleic acid is a DNA-RNA hybrid.

[0312] Hel308 helicase is described in publications such as WO 2013 / 057495, the entire contents of which are incorporated herein by reference. RecD helicase is described in publications such as WO 2013 / 098562, the entire contents of which are incorporated herein by reference. XPD helicase is described in publications such as WO 2013 / 098561, the entire contents of which are incorporated herein by reference. Dda helicase is described in publications such as WO 2015 / 055981 and WO 2016 / 055777, the entire contents of which are incorporated herein by reference.

[0313] In one embodiment, the helicase comprises the sequence shown in SEQ ID NO:6 (Trwc Cba) or a variant thereof, the sequence shown in SEQ ID NO:7 (Hel308 Mbu) or a variant thereof, or the sequence shown in SEQ ID NO:8 (Dda) or a variant thereof. The variants may differ from the natural sequence in any of the ways discussed below. An example variant of SEQ ID NO:8 includes E94C / A360C. Another example variant of SEQ ID NO:8 includes E94C / A360C, followed by (ΔM1)G1G2 (i.e., the deletion of M1, followed by the addition of G1 and G2).

[0314] In some embodiments, a motor protein (e.g., a helicase) can control conjugate movement in at least two active operating modes (when the motor protein has all the necessary components to facilitate movement, such as fuels and cofactors discussed herein, such as ATP and Mg2+) and one inactive operating mode (when the motor protein does not provide the components required to facilitate movement).

[0315] When all the necessary components are provided to facilitate movement (i.e., in active mode), motor proteins (e.g., helicases) move along polynucleotides in a 5' to 3' or 3' to 5' direction (depending on the motor protein). Motor proteins can be used to move conjugates away from (e.g., out of) a pore (e.g., against an applied force) or to move conjugates toward (e.g., into) a pore (e.g., using an applied force). For example, when the end of a conjugate moved by a motor protein is trapped in a pore, the motor protein works against the direction of the force and pulls the conjugate through the pore out (e.g., into the cis chamber). However, when the far end of a conjugate moved by a motor protein is trapped in a pore, the motor protein works in the direction of the force and pushes the conjugate through the pore into the pore (e.g., into the trans chamber).

[0316] When motor proteins (such as helicases) do not provide the necessary components to promote movement (i.e., they are in an inactive mode), they can bind to conjugates and act as brakes, slowing the movement of the construct relative to the nanopore, for example, by being pulled into the pore by force. In the inactive mode, it is not important which end of the conjugate is captured; the applied force determines the movement of the conjugate relative to the pore, and the polynucleotide-binding protein acts as a brake. The control of conjugate movement by polynucleotide-binding proteins in the inactive mode can be described in several ways (including ratcheting, sliding, and braking).

[0317] Motor proteins typically require fuel to process polynucleotides. This fuel is usually a free nucleotide or a free nucleotide analogue. Free nucleotides can be, but are not limited to, adenosine monophosphate (AMP), adenosine diphosphate (ADP), adenosine triphosphate (ATP), guanosine monophosphate (GMP), guanosine diphosphate (GDP), guanosine triphosphate (GTP), thymidine monophosphate (TMP), thymidine diphosphate (TDP), thymidine triphosphate (TTP), uridine monophosphate (UMP), uridine diphosphate (UDP), uridine triphosphate (UTP), cytidine monophosphate (CMP), cytidine diphosphate (CDP), cytidine triphosphate (CTP), cyclic adenosine monophosphate (cAMP), cyclic guanosine monophosphate (cGMP), and deoxygenated adenosine monophosphate (dAMP). Deoxyadenosine monophosphate (dAMP), deoxyadenosine diphosphate (dADP), deoxyadenosine triphosphate (dATP), deoxyguanosine monophosphate (dGMP), deoxyguanosine diphosphate (dGDP), deoxyguanosine triphosphate (dGTP), deoxythymidine monophosphate (dTMP), deoxythymidine diphosphate (dTDP), deoxythymidine triphosphate (dTTP), deoxyuridine monophosphate (dUMP), deoxyuridine diphosphate (dUDP), deoxyuridine triphosphate (dUTP), deoxycytidine monophosphate (dCMP), deoxycytidine diphosphate (dCDP), and deoxycytidine triphosphate (dCTP). Free nucleotides are typically selected from AMP, TMP, GMP, CMP, UMP, dAMP, dTMP, dGMP, or dCMP. The most common free nucleotide is adenosine triphosphate (ATP).

[0318] Cofactors of motor proteins are factors that allow motor proteins to function. Cofactors are preferably divalent metal cations. The preferred divalent metal cation is Mg. 2+ Mn 2+ Ca 2+ or Co 2+ The most preferred cofactor is Mg. 2+ .

[0319] As explained herein, in some embodiments, the polynucleotide treatment protein is modified to extend the distance between the polynucleotide treatment protein and the nanopore when the conjugate is used to control the movement of the polynucleotide treatment protein relative to the nanopore.

[0320] Polynucleotide-treated proteins can be modified in any suitable manner. Modification of proteins (such as polynucleotide-treated proteins) is within the knowledge of those skilled in the art.

[0321] Polynucleotide-treated proteins can be modified by introducing additional amino acids into the protein structure. In some embodiments, polynucleotide-treated proteins are modified by introducing one or more loop regions that extend beyond the native range of the protein. In embodiments where the polynucleotide-treated protein comprises multiple subunits, one or more loop regions may be introduced into one or more subunits of the polynucleotide-treated protein.

[0322] The polynucleotide processing protein can be modified by fusing one or more additional domains to displace the nanopore when the polynucleotide processing protein is in a “sitting position” relative to the nanopore.

[0323] Nanopores

[0324] As explained above, the methods disclosed herein include using polynucleotide-treated proteins to control the movement of conjugates relative to nanopores.

[0325] In the disclosed method, any suitable nanopore can be used. In one embodiment, the nanopore is a transmembrane pore.

[0326] A transmembrane pore is a structure that spans the membrane to some extent. It allows hydrated ions to flow across or within the membrane, driven by an applied potential. A transmembrane pore typically extends across the entire membrane, allowing hydrated ions to flow from one side to the other. However, a transmembrane pore does not necessarily extend across the membrane. It may be closed at one end. For example, a pore can be a hole, gap, channel, groove, or slit in the membrane, allowing hydrated ions to flow into or into the membrane.

[0327] Any transmembrane pore can be used in the methods provided herein. The pore can be biological or artificial. Suitable pores include, but are not limited to, protein pores, polynucleotide pores, and solid-state pores. In one embodiment, a solid-state pore can comprise a nanochannel. The pore can be a DNA origami pore (Langecker et al., Science, 2012; 338:932-936). Suitable DNA origami pores are disclosed in WO2013 / 083983.

[0328] In one embodiment, the nanopore is a transmembrane protein pore. A transmembrane protein pore is a polypeptide or aggregate of polypeptides that allows hydrated ions (e.g., polynucleotides) to flow from one side of a membrane to the other. In the methods provided herein, transmembrane protein pores are capable of forming pores that allow hydrated ions, driven by an applied potential, to flow from one side of a membrane to the other. Transmembrane protein pores preferably allow polynucleotides to flow from one side of a membrane (e.g., a triblock copolymer membrane) to the other. Transmembrane protein pores allow polynucleotides to move through the pore.

[0329] In one embodiment, the nanopore is a transmembrane protein pore, which is a monomer or oligomer. The pore is preferably composed of a plurality of repeating subunits, such as at least 6, at least 7, at least 8, at least 9, at least 10, at least 11, at least 12, at least 13, at least 14, at least 15, or at least 16 subunits. The pore is preferably a hexamer, heptamer, octamer, or non-merchanmeric pore. The pore can be a homooligomer or a heterooligomer.

[0330] In one embodiment, a transmembrane protein pore comprises a barrel or channel through which ions can flow. The subunits of the pore typically surround a central axis and contribute chains to either the transmembrane β-barrel or channel, or the transmembrane α-helical bundle or channel.

[0331] Typically, the barrels or channels of transmembrane protein pores comprise amino acids that facilitate interaction with analytes, such as target polynucleotides (as described herein). These amino acids are preferably located near the constriction of the barrel or channel. Transmembrane protein pores typically contain one or more positively charged amino acids, such as arginine, lysine, or histidine, or aromatic amino acids, such as tyrosine or tryptophan. These amino acids typically facilitate interactions between the pore and nucleotides, polynucleotides, or nucleic acids.

[0332] In one embodiment, the nanopore is a transmembrane protein pore derived from a β-barrel pore or an α-helical bundle pore. A β-barrel pore comprises a barrel or channel formed by β-chains. Suitable β-barrel pores include, but are not limited to, β-toxins such as α-hemolysin, anthrax toxin, and leukocyte toxin, as well as bacterial outer membrane proteins / porins, such as Mycobacterium smegmatis porins (Msp), such as MspA, MspB, MspC or MspD, CsgG, outer membrane porin F (OmpF), outer membrane porin G (OmpG), outer membrane phospholipase A, and Neisseria autotransporter (NalP), and other pores such as cytolysins. An α-helical bundle pore comprises a barrel or channel formed by α-helices. Suitable α-helical bundle pores include, but are not limited to, inner membrane proteins and α-outer membrane proteins, such as WZA and ClyA toxins.

[0333] In one embodiment, the nanopore is a transmembrane pore derived from or based on Msp, α-hemolysin (α-HL), cytolysin, CsgG, ClyA, Sp1, and the hemolysin fragaceatoxin C (FraC).

[0334] In one embodiment, the nanopore is a transmembrane protein pore derived from CsgG, such as CsgG derived from the Escherichia coli strain K-12 substrain MC4100. Such pores are oligomeric and typically comprise 7, 8, 9, or 10 monomers derived from CsgG. The pore can be a homopolymeric oligomeric pore derived from CsgG comprising the same monomers. Alternatively, the pore can be a heteropolymeric pore derived from CsgG, comprising at least one monomer different from the other monomers. Examples of suitable pores derived from CsgG are disclosed in WO 2016 / 034591.

[0335] In one embodiment, the nanopore is a transmembrane pore derived from cytosin. Examples of suitable pores derived from cytosin are disclosed in WO 2013 / 153359.

[0336] In one embodiment, the nanopore is a transmembrane pore derived from or based on α-hemolysin (α-HL). The wild-type α-hemolysin pore is formed by seven identical monomers or subunits (i.e., it is heptameric). The α-hemolysin pore can be α-hemolysin-NN or a variant thereof. Variants preferably include N residues at positions E111 and K147.

[0337] In one embodiment, the nanopore is a transmembrane protein pore derived from Msp, such as MspA. Examples of suitable pores derived from MspA are disclosed in WO 2012 / 107778.

[0338] In one embodiment, the nanopores are transmembrane pores derived from or based on ClyA.

[0339] As explained above, in some embodiments, the nanopore includes contraction. Contraction is typically a narrowing of the channel through the nanopore, which can determine or control the signal obtained when the conjugate moves relative to the nanopore. As used herein, both proteins and solid nanopores generally include “contraction”.

[0340] In some embodiments, the nanopore is modified to extend the distance between the polynucleotide-treated protein and the contractile region of the nanopore. In some embodiments, the nanopore is modified to extend the distance between the polynucleotide-treated protein and the contractile region of the nanopore when the polynucleotide-treated protein is used to control the movement of the conjugate relative to the nanopore. In some embodiments, the nanopore is modified to extend the distance between the polynucleotide-treated protein and the contractile region of the nanopore when the polynucleotide-treated protein contacts the nanopore.

[0341] In some embodiments, the nanopore is modified to extend the distance between the active site of the polynucleotide-treated protein and the contraction region of the nanopore. In such embodiments, the distance may be the distance between the active site of the polynucleotide-treated protein and the contraction of the nanopore when the polynucleotide-treated protein is used to control the movement of the conjugate relative to the nanopore and / or when the polynucleotide-treated protein contacts the nanopore.

[0342] Nanopores can be modified in any suitable manner. Modification of nanopores, such as protein nanopores, is within the knowledge of those skilled in the art. Modification of solid nanopores is conventional and can be achieved by controlling the substrate (e.g., its thickness) or the components that form the nanopores.

[0343] For example, nanopores can be modified to extend the length of the channels that pass through the pores.

[0344] Protein nanopores can be modified by introducing additional amino acids into the pore structure. In some embodiments, protein nanopores are modified by introducing one or more loop regions that extend beyond the native extent of the nanopore. In embodiments where the nanopore comprises multiple subunits, one or more loop regions can be introduced into one or more subunits of the nanopore. The loop regions can, for example, extend beyond the cis-entrance of the nanopore.

[0345] Protein nanopores can be modified to extend the length of the barrels or channels that pass through the pores. For example, β-barrel pores can be modified by introducing additional amino acids into the protein sequence in the barrel-forming portion, thereby extending the barrel length. The rational design of the relevant locations for such modifications can be, for example, by referencing the structure (e.g., X-ray) of the protein and / or its monomeric subunits.

[0346] Protein nanopores can be modified by fusing one or more additional domains to improve the “sitting position” of polynucleotide-treated proteins relative to the nanopore.

[0347] In some embodiments, a protein nanopore can be modified by fusing it with another protein nanopore. In this way, chains of nanopores with single channels running through them can be created to extend the distance between the contraction within the channel and the polynucleotide-processing protein. In such cases, the multiple nanopores can be identical or different.

[0348] Label

[0349] In some embodiments of the methods provided herein, tags on nanopores may be used, for example, to facilitate the capture of conjugates by the nanopores.

[0350] The interaction between the tag on the nanopore and the binding site on the polynucleotide (e.g., a binding site present in the polynucleotide portion of the conjugate or in the adaptor connected to the conjugate, wherein the binding site may be provided by an anchor or leader sequence of the adaptor or by a capture sequence within the double-stranded stem of the adaptor) can be reversible. For example, the polynucleotide can bind to the tag on the nanopore, for example, through its adaptor, and be released at certain points, for example, during characterization of the polynucleotide through the nanopore and / or during motor protein processing. Strong non-covalent binding (e.g., biotin / avidin) remains reversible and can be used in some embodiments of the methods described herein. For example, a pair of pore tags and polynucleotide adaptors can be designed to provide sufficient interaction between the complement of the double-stranded polynucleotide (or the complement-connected portion of the adaptor) and the nanopore, such that the complement remains close to the nanopore (without dissociating and diffusing from the nanopore) but is able to be released from the nanopore upon processing.

[0351] The pore tag and polynucleotide adapter can be configured such that the binding strength or affinity of the binding site on the polynucleotide (e.g., a binding site provided by the anchor or leader sequence of the adapter or by the capture sequence within the double-stranded stem of the adapter) to the tag on the nanopore is sufficient to maintain the coupling between the nanopore and the polynucleotide until an applied force is placed thereon to release the bound polynucleotide from the nanopore.

[0352] In some embodiments, the tag or tether is uncharged. This ensures that the tag or tether will not be pulled into the nanopore under the influence of a potential difference (if present).

[0353] One or more molecules that attract or bind conjugates, polynucleotides, or adaptors can be linked to nanopores. Any molecule that hybridizes with conjugates, adaptors, and / or polynucleotides can be used. The molecules linked to the pore can be selected from PNA tags, PEG linkers, short oligonucleotides, positively charged amino acids, and aptamers. Pores having such molecules linked to them are known in the art. For example, pores having short oligonucleotides linked to them are disclosed in Howarka et al. (2001) Nature Biotech 19:636-639 and WO 2010 / 086620, and pores including PEG linked within the lumen of the pore are disclosed in Howarka et al. (2000) J. Am. Chem. Soc. 122(11):2411-2416.

[0354] Short oligonucleotides linked to nanopores, including sequences complementary to the sequence in the conjugate (e.g., a leader sequence in an adaptor or another single-stranded sequence), can be used to enhance the capture of the conjugate in the methods described herein.

[0355] In some embodiments, the tag or tie may include or may be an oligonucleotide (e.g., DNA, RNA, LNA, BNA, PNA, or morpholino). The oligonucleotide may be about 10-30 nucleotides long or about 10-20 nucleotides long. In some embodiments, the oligonucleotide may have at least one end (e.g., a 3' or 5' end) modified for conjugation to other modified or solid substrate surfaces (including, for example, beads). The end modifier may add reactive functional groups that can be used for conjugation. Examples of functional groups that may be added include, but are not limited to, amino, carboxyl, thiol, maleimide, aminooxy, and any combination thereof. The functional group may be combined with spacers of different lengths (e.g., C3, C9, C12, spacer 9, and 18) to increase the physical distance between the functional group and the end of the oligonucleotide sequence.

[0356] Examples of modifications on the 3' and / or 5' ends of oligonucleotides include, but are not limited to, 3' affinity tags and functional groups for chemical linking (including, for example, 3'-biotin, 3'-primary amine, 3'-disulfide amide, 3'-pyridyl disulfide, and any combination thereof); 5' end modifications (including, for example, 5'-primary amine and / or 5'-fluorescein); modifications for click chemistry (including, for example, 3'-azide, 3'-alkynyl, 5'-azide, 5'-alkynyl) and any combination thereof.

[0357] In some embodiments, the tag or linker may further include a polymer linker, for example, to facilitate coupling to the nanopore. Exemplary polymer linkers include, but are not limited to, polyethylene glycol (PEG). The molecular weight of the polymer linker may be from about 500 Da to about 10 kDa (including end values), or from about 1 kDa to about 5 kDa (including end values). The polymer linker (e.g., PEG) may be functionalized with different functional groups, including, for example, but not limited to, maleimide, NHS ester, dibenzocyclooctylene (DBCO), azide, biotin, amine, alkyne, aldehyde, and any combination thereof.

[0358] Other examples of tags or chains include, but are not limited to, His tags, biotin or streptavidin, antibodies that bind to analytes, aptamers that bind to analytes, analyte-binding domains such as DNA-binding domains (including, for example, peptide zippers, such as leucine zippers, single-stranded DNA-binding proteins (SSBs)) and any combination thereof.

[0359] Tags or tethers can be attached to the outer surface of a nanopore, for example, on the cis side of a membrane, using any method known in the art. For example, one or more tags or tethers can be attached to the nanopore via one or more cysteine ​​residues (cysteine ​​bonds), one or more primary amines (such as lysine), one or more non-natural amino acids, one or more histidine residues (His tags), one or more biotin or streptavidin residues, one or more antibody-based tags, one or more enzymatic modifications of epitopes (including, for example, acetyltransferases), and any combination thereof. Suitable methods for making such modifications are well known in the art. Suitable non-natural amino acids include, but are not limited to, 4-azido-L-phenylalanine (Faz), and Liu C.C. and Schultz PG, *AnnuRev Biochem*, 2010, 79, 413-444. Figure 1 Any one of the amino acids numbered 1-71 in the Chinese language.

[0360] In some embodiments, one or more tags or chains are linked to nanopores via cysteine ​​bonds, one or more cysteine ​​residues may be introduced into one or more monomers that form nanopores by substitution. In some embodiments, the nanopores may be chemically modified by linking to: (i) maleimides, including dibromomaleimides such as: 4-benzodiazepine, 1,N-(2-hydroxyethyl)maleimide, N-cyclohexylmaleimide, 1,3-maleimide propionic acid, 1,1-4-aminophenyl-1H-pyrrole,2,5,dione, 1,1-4-hydroxyphenyl-1H-pyrrole,2,5,dione, N-ethylmaleimide, N... -Methoxycarbonylmaleimide, N-tert-butylmaleimide, N-(2-aminoethyl)maleimide, 3-maleimide-propoxy, N-(4-chlorophenyl)maleimide, 1-[4-(dimethylamino)-3,5-dinitrophenyl]-1H-pyrrole-2,5-dione, N-[4-(2-benzimidazolyl)phenyl]maleimide, N-[4-(2-benzoxazolyl)phenyl]maleimide, N-(1-naphthyl) Maleimide, N-(2,4-dimethyl)maleimide, N-(2,4-difluorophenyl)maleimide, N-(3-chloro-p-tolyl)-maleimide, 1-(2-amino-ethyl)-pyrrole-2,5-dione hydrochloride, 1-cyclopentyl-3-methyl-2,5-dihydro-1H-pyrrole-2,5-dione, 1-(3-aminopropyl)-2,5-dihydro-1H-pyrrole-2,5-dione hydrochloride, 3-methyl -1-[2-oxo-2-(piperazin-1-yl)ethyl]-2,5-dihydro-1H-pyrrole-2,5-dione hydrochloride, 1-benzyl-2,5-dihydro-1H-pyrrole-2,5-dione, 3-methyl-1-(3,3,3-trifluoropropyl)-2,5-dihydro-1H-pyrrole-2,5-dione, 1-[4-(methylamino)cyclohexyl]-2,5-dihydro-1H-pyrrole-2,5-dione trifluoroacetic acid, SMILES O=C1C=CC(=O)N1CC=2C=CN=CC2, SMILES O=C1C=CC(=O)N1CN2CCNCC2、1-Benzyl-3-methyl-2,5-dihydro-1H-pyrrole-2,5-dione、1-(2-fluorophenyl)-3-methyl-2,5-dihydro-1H-pyrrole-2,5-dione、N-(4-phenoxyphenyl)maleimide、N-(4-nitrophenyl)maleimide、(ii)iodoacetamide, such as 3-(2-iodoacetamido)-propoxy、N-(cyclopropylmethyl)-2-iodoacetamide、2-iodo-N-(2-phenylethyl)acetamide、2-iodo-N-(2,2,2-trifluoroethyl)acetamide、N-(4-acetylphenyl)-2-iodoacetamide、N-(4-(aminosulfonyl)phenyl)-2-iodoacetamide、N-(1,3-Benzothiazol-2-yl)-2-iodoacetamide, N-(2,6-diethylphenyl)-2-iodoacetamide, N-(2-benzoyl-4-chlorophenyl)-2-iodoacetamide, (iii) bromoacetamides: such as N-(4-(acetamido)phenyl)-2-bromoacetamide, N-(2-acetylphenyl)-2-bromoacetamide, 2-bromo-n-(2-cyanophenyl)acetamide, 2-bromo-N-(3-(trifluoromethyl)phenyl)acetamide, N-(2-benzoylphenyl)-2-bromoacetamide, 2-bromo-N-(4-fluorophenyl)-3-methylbutyramide, N-benzyl-2-bromo-N-phenylpropionamide, N-(2-bromo-butyryl)-4-chloro-benzenesulfonamide, 2-bromo-N-methyl-N-phenylacetamide, 2-bromo- N-Phenylacetamide, 2-adamantane-1-yl-2-bromo-N-cyclohexylacetamide, 2-bromo-N-(2-methylphenyl)butyramide, acetyl-p-bromoaniline; (iv) disulfides, such as aldrithiol-2, aldrithiol-4, isopropyl disulfide, 1-(isobutyldithioalkyl)-2-methylpropane, dibenzyl disulfide, 4-aminophenyl disulfide, 3-(2-pyridyldithio)propionic acid, 3-(2-pyridyldithio)propionic acid hydrazide, 3-(2-pyridyldithio)propionic acid N-succinimide, am6amPDP1-βCD; and (v) thiols, such as 4-phenylthiazolyl-2-thiol, Pulpald, 5,6,7,8-tetrahydro-quinazolin-2-thiol.

[0361] In some embodiments, the tag or tether can be attached to the nanopore directly or via one or more linkers. The tag or tether can be attached to the nanopore using hybrid linkers as described in WO 2010 / 086602. Alternatively, peptide linkers can be used. Peptide linkers are amino acid sequences. The length, flexibility, and hydrophilicity of peptide linkers are generally designed so that they do not interfere with the function of the monomer and the pore. Preferred flexible peptide linkers are extensions of 2 to 20, such as 4, 6, 8, 10, or 16 serine and / or glycine. More preferred flexible linkers comprise (SG)1, (SG)2, (SG)3, (SG)4, (SG)5, and (SG)8, where S is serine and G is glycine. Preferred rigid linkers are extensions of 2 to 30, such as 4, 6, 8, 16, or 24 proline. More preferred rigid linkers comprise (P)12, where P is proline.

[0362] membrane

[0363] Typically, in the disclosed methods, nanopores are present within the membrane. Any suitable membrane can be used in the system.

[0364] The membrane is preferably an amphiphilic layer. An amphiphilic layer is a layer formed by amphiphilic molecules such as phospholipids, which possess both hydrophilic and lipophilic properties. The amphiphilic molecules can be synthetic or naturally occurring. Non-naturally occurring amphiphiles and amphiphiles forming monolayers are known in the art, and include, for example, block copolymers (Gonzalez-Perez et al., Langmuir, 2009, 25, 10447-10450). A block copolymer is a polymeric material in which two or more monomer subunits polymerized together form a single polymer chain. Block copolymers typically possess properties contributed by each monomer subunit. However, block copolymers can possess unique properties not found in polymers formed from individual subunits. Block copolymers can be engineered so that one of the monomer subunits is hydrophobic (i.e., lipophilic) in an aqueous medium, while the other subunits are hydrophilic. In this case, the block copolymer can possess amphiphilic properties and can form a structure that mimics a biological membrane. Block copolymers can be diblock (consisting of two monomer subunits), but can also be constructed from more than two monomer subunits, forming a more complex arrangement exhibiting amphiphilic behavior. The copolymer can be triblock, tetrablock, or pentablock copolymers. The membrane is preferably a triblock copolymer membrane.

[0365] Archaea bipolar tetraether lipids are naturally occurring lipids that are constructed to form monolayer membranes. These lipids are generally found in extremophiles, thermophiles, halophiles, and acidophiles that survive in harsh biological environments. Their stability is thought to stem from the fusion properties of the final bilayer. A straightforward approach is to construct block copolymer materials that mimic these biological entities by generating triblock polymers with a general motif of hydrophilic-hydrophobic-hydrophilic properties. These materials can form monomeric membranes that exhibit lipid bilayer-like behavior and encompass a range of stages from vesicles to lamellar membranes. Membranes formed from these triblock copolymers retain several advantages over biological lipid membranes. Because of the synthesis of triblock copolymers, precise construction can be carefully controlled to provide the correct chain lengths and properties required for membrane formation and interaction with pores and other proteins.

[0366] Block copolymers can also be constructed from subunits not classified as lipid submaterials; for example, hydrophobic polymers can be made from siloxanes or other non-hydrocarbon-based monomers. The hydrophilic subsegments of the block copolymers can also possess low protein-binding properties, allowing for the creation of highly resistant membranes when exposed to pristine biological samples. This head unit can also be derived from non-classical lipid head units.

[0367] Compared to bio-lipid membranes, triblock copolymer membranes also exhibit increased mechanical and environmental stability, such as a much wider operating temperature or pH range. The synthetic properties of block copolymers provide a platform for customizing polymer-based membranes for a wide range of applications.

[0368] In some embodiments, the membrane is one of the membranes disclosed in International Application No. WO2014 / 064443 or WO2014 / 064444.

[0369] Amphiphilic molecules can be chemically modified or functionalized to facilitate the coupling of polynucleotides. The amphiphilic layer can be monolayer or bilayer. The amphiphilic layer is typically planar. The amphiphilic layer can be curved. The amphiphilic layer can be supportive.

[0370] Amphiphilic membranes are typically naturally mobile, essentially functioning as two-dimensional liquids with a lipid diffusion rate of approximately 10⁻⁸ cm s⁻¹. This means that pores and coupled polynucleotides can generally move within the amphiphilic membrane.

[0371] The membrane can be a lipid bilayer. Lipid bilayers are models of cell membranes and serve as excellent platforms for a range of experimental studies. For example, lipid bilayers can be used for in vitro studies of membrane proteins via single-channel recording. Alternatively, lipid bilayers can be used as biosensors to detect the presence of a range of substances. A lipid bilayer can be any lipid bilayer. Suitable lipid bilayers include, but are not limited to, planar lipid bilayers, supported bilayers, or liposomes. Preferably, a planar lipid bilayer is a flat lipid bilayer. Suitable lipid bilayers are disclosed in WO 2008 / 102121, WO 2009 / 077734, and WO 2006 / 100484.

[0372] Methods for forming lipid bilayers are known in the art. Lipid bilayers are typically formed by the method of Montal and Mueller (Proceedings of the National Academy of Sciences, 1972; 69:3561-3566), in which a lipid monolayer is carried on an aqueous / air interface across an opening perpendicular to the interface. The lipid is typically added to the surface of an aqueous electrolyte solution by first dissolving the lipid in an organic solvent and then evaporating a drop of solvent from the surface of the aqueous solution on either side of the opening. Once the organic solvent has evaporated, the solution / air interface on either side of the opening physically moves back and forth through the opening until a bilayer is formed. Planar lipid bilayers can be formed across openings in a membrane or across openings in a groove.

[0373] The Montal and Mueller method is commonly used because it is cost-effective and a relatively straightforward method for forming a good-quality lipid bilayer suitable for protein pore insertion. Other common methods for bilayer formation include tip immersion, bilayer brushing, and patch clamping.

[0374] Tip-immersion bilayer formation requires contact between the open end surface (e.g., a pipette tip) and the surface of the test solution carrying the lipid monolayer. Similarly, a lipid monolayer is first generated at the solution / air interface by evaporating a drop of lipid dissolved in an organic solvent at the solution surface. The bilayer is then formed via a Langmuir-Schaefer process, requiring mechanical automation to move the open end relative to the solution surface.

[0375] For the brush-coated bilayer, a drop of lipid dissolved in an organic solvent is applied directly to an opening immersed in an aqueous test solution. Using a brush or equivalent, the lipid solution is thinly diffused within the opening. This solvent dilution allows the formation of a lipid bilayer. However, completely removing the solvent from the bilayer is very difficult, and therefore the bilayer formed by this method is less stable and more prone to noise during electrochemical measurements.

[0376] Patch clamping is commonly used in biological cell membrane research. The cell membrane is clamped to the end of a pipette by aspiration, and the membrane patch becomes attached within an opening. This method is suitable for generating lipid bilayers by clamping and then bursting liposomes to detach from the lipid bilayer sealed within an opening of the pipette. This method requires stable, large, monolayer liposomes and the fabrication of small openings in a material with a glass surface.

[0377] Liposomes can be formed by sonication, extrusion or the Mozafari method (Colas et al. (2007) Micron 38:841-847).

[0378] In some embodiments, a lipid bilayer is formed as described in International Application WO 2009 / 077734. It is advantageous in this method to form the lipid bilayer from dried lipids. In a most preferred embodiment, the lipid bilayer is formed across an opening, as described in WO2009 / 077734.

[0379] A lipid bilayer is formed by two opposing lipid layers. The two lipid layers are arranged such that their hydrophobic tail groups face each other, forming a hydrophobic interior. The hydrophilic head groups of the lipids face outwards towards the aqueous environment on each side of the bilayer. The bilayer can exist in various lipid stages, including but not limited to liquid disordered stages (liquid sheets), liquid ordered stages, solid ordered stages (sheet-gel stages, interleaved gel stages), and planar bilayer crystals (sheet-subgel stages, sheet-crystalline stages).

[0380] Any lipid composition that forms a lipid bilayer can be used. The lipid composition is selected such that the lipid bilayer possesses desired properties, such as surface charge, ability to support membrane proteins, filling density, or the mechanical properties formed. The lipid composition may include one or more different lipids. For example, the lipid composition may contain up to 100 lipids. The lipid composition preferably contains 1 to 10 lipids. The lipid composition may include naturally occurring lipids and / or artificial lipids.

[0381] Lipids typically consist of a head group, an interfacial portion, and two hydrophobic tail groups that may be the same or different. Suitable head groups include (but are not limited to): neutral head groups, such as diacylglycerol esters (DG) and ceramides (CM); zwitterionic head groups, such as phosphatidylcholine (PC), phosphatidylethanolamine (PE), and sphingomyelin (SM); negatively charged head groups, such as phosphatidylglycerol (PG); phosphatidylserine (PS), phosphatidylinositol (PI), phosphatidic acid (PA), and cardiolipin (CA); and positively charged head groups, such as trimethylammonium propane (TAP). Suitable interfacial portions include, but are not limited to, naturally occurring interfacial portions, such as glycerol-based or ceramide-based portions. Suitable hydrophobic tail groups include, but are not limited to: saturated hydrocarbon chains, such as lauric acid (n-dodecanoic acid), myristic acid (n-tetradecanoic acid), palmitic acid (n-hexadecanoic acid), stearic acid (n-octadecanoic acid), and arachidic acid (n-eicosanoic acid); unsaturated hydrocarbon chains, such as oleic acid (cis-9-octadecanoic acid); and branched hydrocarbon chains, such as phytanoyl groups. The chain length and the position and number of double bonds in the unsaturated hydrocarbon chain can vary. The chain length and the position and number of branches (such as methyl groups) in the branched hydrocarbon chain can also vary. The hydrophobic tail group can be attached to the interfacial portion as an ether or ester. Lipids can be mycolic acids.

[0382] Lipids can also be chemically modified. The head or tail groups of lipids can be chemically modified. Suitable lipids with chemically modified head groups include, but are not limited to: PEG-modified lipids, such as 1,2-diacyl-sn-glycerol-3-phosphate ethanolamine-N-[methoxy(polyethylene glycol)-2000]; functionalized PEG lipids, such as 1,2-distearate-sn-glycerol-3-phosphate ethanolamine-N-[biotinyl(polyethylene glycol)2000]; and lipids for conjugation modification, such as 1,2-dioleoyl-sn-glycerol-3-phosphate ethanolamine-N-(succinyl) and 1,2-dispalmitoyl-sn-glycerol-3-phosphate ethanolamine-N-(biotinyl). Suitable lipids with chemically modified tails include, but are not limited to: polymerizable lipids, such as 1,2-bis(10,12-tetracarbadiynyl)-sn-glycerol-3-phosphate choline; fluorinated lipids, such as 1-palmitoyl-2-(16-fluoropalmitoyl)-sn-glycerol-3-phosphate choline; deuterated lipids, such as 1,2-dipalmitoyl-D62-sn-glycerol-3-phosphate choline; and ether-linked lipids, such as 1,2-di-O-phytyl-sn-glycerol-3-phosphate choline. Lipids may be chemically modified or functionalized to facilitate coupling with polynucleotides.

[0383] Amphiphilic layers, such as lipid compositions, typically include one or more additives that will affect the properties of the layer. Suitable additives include, but are not limited to: fatty acids, such as palmitic acid, myristic acid, and oleic acid; fatty alcohols, such as palmitol, myristicol, and oleyl alcohol; sterols, such as cholesterol, ergosterol, lanosterol, sitosterol, and stigmasterol; lysophospholipids, such as 1-acyl-2-hydroxy-sn-glycerol-3-phosphocholine; and ceramides.

[0384] In another embodiment, the film includes a solid layer. The solid layer can be formed of both organic and inorganic materials, including, but not limited to: microelectronic materials, insulating materials (such as Si3N4, Al2O3, and SiO), organic and inorganic polymers (such as polyamides), and plastics (such as...). The membrane may consist of an amphiphilic membrane or layer, such as a two-component addition-cured silicone rubber, or glass. The solid layer may be formed from graphene. Suitable graphene layers are disclosed in WO 2009 / 035647. If the membrane includes a solid layer, pores are typically present in the amphiphilic membrane or layer contained within the solid layer, for example, in pores, holes, gaps, channels, trenches, or slots within the solid layer. Suitable solid / amphiphilic hybrid systems can be prepared by those skilled in the art. Suitable systems are disclosed in WO2009 / 020682 and WO 2012 / 005857. Any of the amphiphilic membranes or layers discussed above may be used.

[0385] The methods disclosed herein are typically carried out using: (i) an artificial amphiphilic layer comprising a pore, (ii) a separated, naturally occurring lipid bilayer comprising a pore, or (iii) a cell into which the pore is inserted. Artificial amphiphilic layers (such as artificial triblock copolymer layers) are typically used to perform the methods. The layers may include other transmembrane and / or intramembrane proteins and other molecules besides the pore. Suitable apparatus and conditions are discussed below. The disclosed methods are typically performed in vitro.

[0386] condition

[0387] As explained above, the disclosed method includes characterizing the polypeptide as the conjugate comprising the polypeptide moves relative to the nanopore.

[0388] Characterization methods can be performed using any apparatus suitable for studying membrane / pore systems in which pores are inserted into the membrane. Characterization methods can be performed using any apparatus suitable for transmembrane pore sensing. For example, the apparatus may include a chamber containing an aqueous solution and a barrier dividing the chamber into two sections. The barrier typically has openings in which a pore-containing membrane is formed. Transmembrane pores are described herein.

[0389] The characterization method can be performed using the apparatus described in WO 2008 / 102120, WO 2010 / 122293 or WO 00 / 28312.

[0390] Characterization methods may involve measuring the ion current flowing through the pore, typically by measuring the current. Alternatively, the ion current through the pore may be measured optically, as disclosed in Heron et al., ACS Journal 9, Vol. 131, No. 5, 2009. Therefore, the apparatus may also include circuitry capable of applying a potential and measuring the electrical signal across the membrane and the pore. Characterization methods may be performed using patch clamps or voltage clamps. Characterization methods preferably involve the use of voltage clamps.

[0391] Characterization methods can be performed on silicon-based aperture arrays, each array comprising 128, 256, 512, 1024, 2000, 3000, 4000, 6000, 10000, 12000, 15000 or more apertures.

[0392] Characterization methods may involve measuring the current flowing through the pore. These methods are typically performed with a voltage applied across the membrane and through the pore. The voltage used is typically +2V to -2V, and usually -400mV to +400mV. Preferably, the voltage used is within a range having a lower limit selected from -400mV, -300mV, -200mV, -150mV, -100mV, -50mV, -20mV, and 0mV, and the upper limit is independently selected from +10mV, +20mV, +50mV, +100mV, +150mV, +200mV, +300mV, and +400mV. More preferably, the voltage used is within the range of 100mV to 240mV, and most preferably within the range of 120mV to 220mV. By using an increased applied potential, the distinguishability between different nucleotides can be increased through the pore.

[0393] Characterization methods are typically performed in the presence of any charge carriers, such as metal salts, for example alkali metal salts; halide salts, such as chloride salts, like alkali metal chloride salts. The charge carriers may comprise ionic liquids or organic salts, such as tetramethylammonium chloride, trimethylphenylammonium chloride, phenyltrimethylammonium chloride, or 1-ethyl-3-methylimidazolium chloride. In the exemplary apparatus discussed above, the salt is present in an aqueous solution within the chamber. Potassium chloride (KCl), sodium chloride (NaCl), or cesium chloride (CsCl) is typically used. KCl is preferred. The salt may be an alkaline earth metal salt, such as calcium chloride (CaCl2). The salt concentration may be saturated. The salt concentration may be 3 M or lower, and is typically 0.1 M to 2.5 M, 0.3 M to 1.9 M, 0.5 M to 1.8 M, 0.7 M to 1.7 M, 0.9 M to 1.6 M, or 1 M to 1.4 M. The salt concentration is preferably 150 mM to 1 M. The characterization method preferably uses a salt concentration of at least 0.3 M, such as at least 0.4 M, at least 0.5 M, at least 0.6 M, at least 0.8 M, at least 1.0 M, at least 1.5 M, at least 2.0 M, at least 2.5 M, or at least 3.0 M. High salt concentrations provide a high signal-to-noise ratio and allow identification of currents indicating binding / non-binding against a background of normal current fluctuations.

[0394] Characterization methods are typically performed in the presence of a buffer solution. In the exemplary apparatus discussed above, the buffer solution is present in an aqueous solution within the chamber. Any suitable buffer solution can be used. Typically, the buffer solution is HEPES. Another suitable buffer solution is Tris-HCl buffer. The methods are typically performed at the following pH values: 4.0 to 12.0, 4.5 to 10.0, 5.0 to 9.0, 5.5 to 8.8, 6.0 to 8.7, or 7.0 to 8.8, or 7.5 to 8.5. The pH value used is preferably about 7.5.

[0395] Characterization methods can be performed at the following temperatures: 0°C to 100°C, 15°C to 95°C, 16°C to 90°C, 17°C to 85°C, 18°C ​​to 80°C, 19°C to 70°C, or 20°C to 60°C. Characterization is typically performed at room temperature. Optionally, characterization can be performed at temperatures that support enzyme function, such as approximately 37°C.

[0396] Modified nanopores

[0397] A nanopore is also provided, comprising a contraction region, wherein the nanopore is modified to increase the distance between the contraction region and a polynucleotide-processing protein in contact with the nanopore. The nanopore can be as described herein. The nanopore can be modified as described herein.

[0398] system

[0399] A system is also provided, the system comprising:

[0400] - Nanopores, wherein the nanopores include contraction regions;

[0401] - Conjugates, said conjugates comprising polypeptides conjugated to polynucleotides; and

[0402] - Polynucleotide-treated proteins;

[0403] in

[0404] i) The nanopore is modified to increase the distance between the contractile region and the active site of the polynucleotide-treated protein when the polynucleotide-treated enzyme contacts the nanopore; and / or

[0405] ii) The system further includes one or more replacement units disposed between the nanopore and the polynucleotide-treated protein, thereby extending the distance between the nanopore and the active site of the polynucleotide-treated protein.

[0406] In some embodiments, the nanopore, the conjugate, and / or the polynucleotide-treated protein, and optionally one or more replacement units, if present, as described herein.

[0407] Reagent test kit

[0408] A kit is also provided, the kit comprising:

[0409] - Nanopores, wherein the nanopores include contraction regions;

[0410] - A polynucleotide, said polynucleotide comprising a reactive functional group for conjugation to a target polynucleotide; and

[0411] - Polynucleotide-treated proteins.

[0412] In some embodiments, the nanopore is modified to increase the distance between the contractile region and the polynucleotide-treated protein when the polynucleotide-treated enzyme comes into contact with the nanopore.

[0413] In some embodiments, the kit further includes one or more replacement units for extending the distance between the nanopore and the active site of the polynucleotide-treated protein.

[0414] In some embodiments, the nanopore, the polynucleotide and / or the polynucleotide-treated protein, and optionally one or more replacement units, if present, are as described herein.

[0415] The kit can be configured for use with the algorithm also provided herein, which is adapted to run on a computer system. The algorithm can be adapted to detect peptide-specific information (e.g., peptide sequence-specific information and / or whether the peptide is modified) and selectively process signals obtained when a conjugate comprising a polypeptide conjugated with a polynucleotide moves relative to a nanopore. A system including a computing device is also provided, configured to detect peptide-specific information (e.g., peptide sequence-specific information and / or whether the peptide is modified) and selectively process signals obtained when a conjugate comprising a polypeptide conjugated with a polynucleotide moves relative to a nanopore. In some embodiments, the system includes a receiving device for receiving data from the detection of the peptide, a processing device for processing signals obtained when the conjugate moves relative to a nanopore, and an output device for outputting the characterization information thus obtained.

[0416] It should be understood that although specific embodiments, constructions, and materials and / or molecules have been discussed herein with respect to the methods according to the invention, various changes or modifications in form and detail may be made without departing from the scope and spirit of the invention. The foregoing embodiments and the following examples are provided for illustrative purposes only and should not be considered as limiting the scope of this application. This application is limited only by the claims.

[0417] Example

[0418] Example 1

[0419] This example demonstrates the controlled translocation of a conjugate comprising a polypeptide with two polynucleotides side-attached: a dsDNA Y adaptor (DNA1) and a dsDNA tail (DNA2). The polynucleotide processing protein at the cis-side of the nanopore controls the conjugate movement by first unfolding DNA1 and transferring it 5'–3' onto the ssDNA, then sliding across the polypeptide moiety to ultimately unfold the DNA2 fragment. As this construct moves from the cis-side to the trans-side of the nanopore, thus traversing the RED, the polypeptide moiety can be visualized on a current-to-time plot, enabling characterization.

[0420] Y-adaptors were prepared by attaching DNA oligonucleotides (SEQ ID NO:11, SEQ ID NO:12, SEQ ID NO:13). A DNA motor (Dda helicase) was loaded onto and shut down on the adaptor, as described in WO 2014 / 013260. The material was subsequently purified by HPLC. The Y-adaptor contained 30 C3 leader sequences for capture by nanopores and side arms for tethering to membranes. DNA tails were prepared by attaching DNA oligonucleotides (SEQ ID NO:14, SEQ ID NO:16).

[0421] In this embodiment, the model peptide analytes (SEQ ID NO: 20, 21, 22) were obtained by pre-synthesizing an azide moiety at the N-terminus and directly after the C-terminus using an ethylenediamine spacer consistent with the peptide backbone. Each analyte was then conjugated to the Y-adaptor and DNA tail via a copper-free click chemistry between the azide and the BCN (bicyclic [6.1.0]nonyne) moiety. Figure 4 A schematic diagram of the resulting construct is shown. Samples were purified using Agencourt AMPure XP beads (Beckman Coulter) and washed twice with LNBs from an Oxford Nanopore Technologies sequencing kit (SQK-LSK109). The conjugated substrate was eluted in 10 mM Tris-Cl and 50 mM NaCl (pH 8.0).

[0422] Electrophysiological measurements were performed using the Oxford Nanopore Technologies MinION Mk1b and a custom MinION flow cell with MspA nanopores. The flow cell was flushed with a tethered mixture containing 50 nM DNA tethering and sequencing buffer lacking ATP. Initially, 800 μL of the tethered mixture was added for 5 minutes, followed by another 200 μL of the mixture flowing through the system with the SpotON port open. A 0.5 nM DNA peptide construct was prepared in LB of the Oxford Nanopore Technologies Sequencing Kit (SQK-LSK109) in ATP-free sequencing buffer, producing a “sequencing mixture.” 75 μL of the sequencing mixture was added to the MinION flow cell through the SpotON flow cell port. The mixture was incubated on the flow cell for 5–10 minutes to allow the constructs to tether and subsequently be captured by the nanopores. Without ATP, the DNA motor remained stationary in the spacer region of the Y-adjoint, and the conjugate was captured by the nanopores, but no translocation occurred. After incubation, 200 μL of sequencing buffer containing ATP was added; in the presence of ATP, the captured DNA-peptide conjugates moved through the nanopores via helicase, thereby generating a reproducible current footprint.

[0423] A standard sequencing script at 180 mV was run for 30 minutes to 1 hour, with static flipping performed every minute to remove expanded nanopore blocks. Raw data were collected in batch FAST5 files using MinKNOW software (Oxford Nanopore Technologies).

[0424] exist Figure 9 An exemplary current-to-time trace of one of the model peptides (SEQ ID NO:20) conjugated with the DNA Y-adaptor and tail can be seen. The Y-adaptor portion and the dsDNA tail can be separated from the peptide portion of the “curve” (trace), enabling peptide characterization.

[0425] High throughput was achieved by observing multiple translocation events per second. Figure 10 The text shows the relationship with... Figure 9 The same construct used (i.e., the peptide region containing SEQ ID NO:20) is shown as an example current-to-time trace of multiple capture and translocation events.

[0426] Characterize other conjugated polynucleotide-peptide constructs as described above. Figures 11 to 13 Reproducible current-to-time traces are shown, enabling characterization of constructs incorporating peptide regions containing positively charged amino acids (SEQ ID NO:21). Figure 11 Aromatic amino acids (SEQ ID NO:22, Figure 12 ) and negatively charged amino acids (SEQ ID NO:20, Figure 13 ).

[0427] For ease of reference, a schematic structure of the construct obtained using the peptide of SEQ ID NO:22 is shown below. Figure 14 As shown.

[0428] Example 2

[0429] This example demonstrates the utility of the disclosed method in characterizing polynucleotide-peptide constructs obtained from peptides that have never been pre-synthesized to include linker groups.

[0430] In this example, the Y-adaptor is the same as in Example 1, and dsDNA tails are prepared by attaching DNA oligonucleotides (SEQ ID NO:15, SEQ ID NO:16). Data collection was performed using the protocol established in Example 1 on a MinION Mk1b from Oxford Nanopore Technologies and a custom MinION flow cell with MspA nanopores.

[0431] The peptide analyte used in this example is almost identical to the model peptide used in Example 1 (GGSGDDSGSG, SEQ ID NO: 20 for Example 1; SEQ ID NO: 23 for Example 2), but lacks the pre-synthesized azide molecule for click chemical conjugation of the polynucleotide intransitive linker and tail. An additional C-terminal cysteine ​​is included to enable maleimide chemistry. The N-terminus of the peptide is functionalized with a tetrazine-NHS ester compound (BroadPharm, product code: BP-22946). Unconjugated tetrazines are removed using amino-functionalized magnetic particles (Sigma Aldrich, product code: 53572).

[0432] The peptide and DNA tail (SEQ ID NO:15, SEQ ID NO:16) were then incubated overnight at 4°C to promote the click reaction between the tetrazine and TCO (trans-cyclooctene). After incubation, the possible disulfide bonds between the C-terminal peptide cysteine ​​residues were reduced with 5 mM DTT at room temperature for 30 min, and the peptide-DNA conjugate was purified using Agencourt AMPure XP beads (Beckman Coulter) to remove unreacted peptide and DTT. The exposed cysteine ​​residues were then reacted with azide-PEG3-maleimide (BroadPharm, product code: BP-22468). Excess maleimide linkers were removed using Agencourt AMPure XP beads, and the construct was reacted with the Y-linker overnight at 4°C via click chemistry between BCN and the azide. The construct formed by the conjugation between the C-terminus and Y-linker, and the N-terminus and DNA tail, was purified using Agencourt AMPure XP beads to separate the complete construct from the peptide-DNA tail.

[0433] The final construct is evaluated as described in Example 1, and exemplary current traces are presented in... Figure 15 As can be seen, peptide characterization is possible without the need for the pre-synthesis of linkage sites.

[0434] Example 3

[0435] This example compares the disclosed methods for characterizing 21-amino acid peptides and 10-amino acid peptides.

[0436] The 21-amino acid peptide polynucleotide-peptide conjugates were prepared and analyzed according to the method described in Example 1. Current-time traces obtained using the 21-amino acid construct were compared with traces obtained using the 10-amino acid construct from Example 2. The peptide sequences used were GDDDGSASGDDDGSASGDDDG (21aa; SEQ ID NO: 24) and GGSGDDSGSG (10aa; SEQ ID NO: 20).

[0437] Data showing the current versus time traces of the translocation of polynucleotide-peptide conjugates of 10-amino acid peptides and 21-amino acid peptides are presented in Figure 16 As shown in the figure, two traces placed on the same timescale emphasize that the current portion of the 21-amino acid polypeptide is approximately twice as long as the 10-amino acid polypeptide.

[0438] Therefore, this example demonstrates that the disclosed method can be used to characterize peptides of different lengths and elongations.

[0439] Sequence List Description

[0440] SEQ ID NO:1 shows the amino acid sequence of (hexahistidine-labeled) exonuclease I (EcoExo I) from Escherichia coli.

[0441] SEQ ID NO:2 shows the amino acid sequence of exonuclease III from Escherichia coli.

[0442] SEQ ID NO:3 shows the amino acid sequence of the RecJ enzyme (TthRecJ-cd) from thermophilic bacteria.

[0443] SEQ ID NO:4 shows the amino acid sequence of the bacteriophage λ exonuclease. This sequence is one of the three identical subunits that assemble into the trimer. (http: / / www.neb.com / nebecomm / products / productM0262.asp).

[0444] SEQ ID NO:5 shows the amino acid sequence of Phi29 DNA polymerase from Bacillus subtilis.

[0445] SEQ ID NO:6 shows the amino acid sequence of the Trwc Cba (Citromicrobium bathyomarinum) helicase.

[0446] SEQ ID NO:7 shows the amino acid sequence of the helicase from Hel308 Mbu (Methanococcoides burtonii).

[0447] SEQ ID NO:8 shows the amino acid sequence of Dda helicase 1993 from Enterobacter T4 bacteriophage.

[0448] SEQ ID NO:11 shows the sequence of the first polynucleotide chain used to generate the Y adaptor as described in Example 1 (DNA1-top, with a C3[(OC3H6OPO3)] leader sequence, a 3'BCN click link, and an enzyme stagnation chemistry; 8=iSp18[(OCH2CH2)6OPO3]).

[0449] SEQ ID NO:12 shows the sequence of the second polynucleotide chain used to generate the Y adaptor as described in Example 1 (DNA1-with a side arm for connecting the strand).

[0450] SEQ ID NO:13 shows the sequence (DNA1-bottom) of the third polynucleotide chain used to generate the Y-adaptor as described in Example 1.

[0451] SEQ ID NO:14 shows the sequence (DNA 2-top strand, 5' BCN click chemistry) used to generate the first polynucleotide chain of the dsDNA tail as described in Example 1.

[0452] SEQ ID NO:15 shows the sequence (DNA2-top chain, 5'TCO (orthogonal click chemistry)) used to generate the second polynucleotide chain of the dsDNA tail as described in Example 1.

[0453] SEQ ID NO:16 shows the sequence (DNA2-bottom strand, no side arms) used to generate the polynucleotide chain for producing the dsDNA tail as described in Example 2.

[0454] SEQ ID NO:20 shows the amino acid sequence of the first peptide fragment used to generate the first polynucleotide-peptide construct as described in Example 1.

[0455] SEQ ID NO:21 shows the amino acid sequence of the second peptide fragment used to generate the second polynucleotide-peptide construct as described in Example 1.

[0456] SEQ ID NO:22 shows the amino acid sequence of the third peptide fragment used to generate the third polynucleotide-peptide construct as described in Example 1.

[0457] SEQ ID NO:23 shows the amino acid sequence of the peptide fragment used to generate the polynucleotide-peptide construct as described in Example 2.

[0458] SEQ ID NO:24 shows the amino acid sequence of the 21-amino acid peptide fragment used to generate the polynucleotide-peptide construct as described in Example 3.

[0459] sequence list

[0460] SEQ ID NO:1 - Exonuclease I from Escherichia coli

[0461]

[0462] SEQ ID NO:2 - Exonuclease III from Escherichia coli

[0463] SEQ ID NO:3 - RecJ enzyme from thermophilic bacteria

[0464] SEQ ID NO:4-phage λ exonuclease

[0465] SEQ ID NO:5-Phi29 DNA polymerase

[0466] SEQ ID NO:6-Trwc Cba helicase

[0467] SEQ ID NO:7-Hel308 Mbu helicase

[0468]

[0469] SEQ ID NO:8-Dda helicase

[0470]

[0471] SEQ ID NO:11

[0472]

[0473] SEQ ID NO:12

[0474] CGACGTACGCATTCGACGATTGCTTTGAGGCGAGCGGTCAA

[0475] SEQ ID NO:13

[0476] GCAATACGTAACTGAACGAAGT / iBNA - A / / iBNA - meC / / iBNA - A / / iBNAT / / 3BNA - T / SEQ IDNO:14

[0477] BCN - AATGTACTTCGTTCAGTTACGTATTGCTGCTTGGGTGTTTAACC

[0478] SEQ ID NO:15

[0479] TCO - AATGTACTTCGTTCAGTTACGTATTGCTGCTTGGGTGTTTAACC

[0480] SEQ ID NO:16

[0481] GGTTAAACACCCAAGCAGCAATACGTAACTGAACGAAGTACATT

[0482] SEQ ID NO:20

[0483] X - GGSGDDSGSG - ed - X

[0484] (X = azidoacetyl; ed = ethylenediamine)

[0485] SEQ ID NO:21

[0486] X - GGSGRRSGSG - ed - X

[0487] (X = azidoacetyl; ed = ethylenediamine)

[0488] SEQ ID NO:22

[0489] X - GGSGYYSGSG - ed - X

[0490] (X = azidoacetyl; ed = ethylenediamine)

[0491] SEQ ID NO:23

[0492] GGSGDDSGSC

[0493] SEQ ID NO:24

[0494] GDDDGSASGDDDGSASGDDDG

Claims

1. A method of characterizing a target polypeptide, comprising: - conjugating the target polypeptide to a polynucleotide to form a polynucleotide-polypeptide conjugate; - contacting the conjugate with a polynucleotide handling protein capable of controlling movement of the polynucleotide relative to a nanopore; and - while the conjugate is moving relative to the nanopore, making one or more measurements that characterize the polypeptide, thereby characterizing the polypeptide; wherein the conjugate comprises a leader sequence that facilitates passage of the conjugate through the nanopore; and wherein the nanopore is a transmembrane protein pore.

2. The method of claim 1, wherein the leader sequence comprises a polynucleotide or a charged polymer.

3. The method of claim 1, wherein the polynucleotide handling protein is capable of remaining bound to the conjugate when a portion of the conjugate in contact with an active site of the polynucleotide handling protein comprises a polypeptide.

4. The method of claim 1, wherein the polynucleotide handling protein is modified to prevent dissociation of the polynucleotide handling protein from the conjugate when the polynucleotide handling protein contacts a portion of the conjugate comprising a polypeptide.

5. The method of claim 1, wherein the polynucleotide handling protein is modified to completely or partially occlude an opening present in at least one conformational state of the unmodified protein through which a polynucleotide strand can unbind.

6. The method of claim 1, wherein the polynucleotide handling protein is a helicase.

7. The method of claim 1, wherein the conjugate comprises a plurality of polypeptide portions and / or a plurality of polynucleotide portions.

8. The method of claim 1, wherein the polypeptide is 2 to 50 peptide units in length, and / or wherein the polynucleotide is 10 to 1000 nucleotides in length.

9. The method of claim 1, wherein one or more adaptors are linked to the polynucleotide in the conjugate.

10. The method of claim 1, wherein the conjugate comprises one or more structures in the form of L-{P-N}-P m wherein:​ -L is a leader sequence, wherein L is optionally an N portion; -P is a polypeptide; -N comprises a polynucleotide; and -m is 0 or 1; and wherein the method comprises passing the leader sequence (L) through the nanopore, thereby contacting the polypeptide (P) with the nanopore.

11. The method of claim 1, wherein: i) the polynucleotide handling protein is located on the cis side of the nanopore, and the polynucleotide handling protein controls movement of the conjugate from the cis side of the nanopore to the trans side of the nanopore; or ii) the polynucleotide handling protein is located on the trans side of the nanopore, and the polynucleotide handling protein controls movement of the conjugate from the trans side of the nanopore to the cis side of the nanopore.

12. The method of claim 1, wherein the conjugate comprises one or more structures in the form of L-{P-N}-P m wherein:​ -L is a leader sequence, wherein L is optionally an N portion; -P is a polypeptide; -N comprises a polynucleotide; and -m is 0 or 1; and wherein the method comprises passing the leader sequence (L) through the nanopore, thereby contacting the polypeptide (P) with the nanopore; and i) the polynucleotide handling protein is located on the cis side of the nanopore, and the method comprises allowing the polynucleotide handling protein to control movement of the polynucleotide moiety (N) from the cis side of the nanopore to the trans side of the nanopore, thereby controlling movement of the polypeptide (P) through the nanopore; or ii) the polynucleotide handling protein is located on the trans side of the nanopore, and the method comprises allowing the polynucleotide handling protein to control movement of the polynucleotide moiety (N) from the trans side of the nanopore to the cis side of the nanopore, thereby controlling movement of the polypeptide (P) through the nanopore.

13. The method of claim 12, wherein the conjugate comprises one or more structures of the form L-P1-N-{P-N} n -P m wherein: - n is a positive integer; - L is a leader sequence, wherein L is optionally an N moiety; - each P, which can be the same or different, is a polypeptide; - each N, which can be the same or different, comprises a polynucleotide; and - m is 0 or 1 ; and wherein the method comprises passing the leader sequence (L) through the nanopore, thereby bringing the polypeptide (P1) into contact with the nanopore; and i) the polynucleotide handling protein is located on the cis side of the nanopore, and the method comprises allowing the polynucleotide handling protein to control movement of each polynucleotide (N) sequentially from the cis side of the nanopore to the trans side of the nanopore, thereby controlling movement of each polypeptide (P) sequentially through the nanopore; or ii) the polynucleotide handling protein is located on the trans side of the nanopore, and the method comprises allowing the polynucleotide handling protein to control movement of each polynucleotide (N) sequentially from the trans side of the nanopore to the cis side of the nanopore, thereby controlling movement of each polypeptide (P) sequentially through the nanopore.

14. The method of claim 1, wherein: i) the polynucleotide handling protein is located on the cis side of the nanopore, and the polynucleotide handling protein controls movement of the conjugate from the trans side of the nanopore to the cis side of the nanopore; or ii) the polynucleotide handling protein is located on the trans side of the nanopore, and the polynucleotide handling protein controls movement of the conjugate from the cis side of the nanopore to the trans side of the nanopore.

15. The method of claim 1, wherein the conjugate comprises one or more structures in the form of L-{P-N}-P m wherein:​ - L is a leader sequence, wherein L is optionally an N moiety; - P is a polypeptide; - N comprises a polynucleotide; - m is 0 or 1 ; and wherein the method comprises passing the leader sequence (L) through the nanopore, thereby bringing the polypeptide (P) into contact with the nanopore; and i) the polynucleotide handling protein is located on the cis side of the nanopore, and the method comprises allowing the polynucleotide handling protein to control movement of the polynucleotide (N) from the trans side of the nanopore to the cis side of the nanopore, thereby controlling movement of the polypeptide (P) through the nanopore; or i) the polynucleotide handling protein is located on the trans side of the nanopore, and the method comprises allowing the polynucleotide handling protein to control movement of the polynucleotide (N) from the cis side of the nanopore to the trans side of the nanopore, thereby controlling movement of the polypeptide (P) through the nanopore.

16. The method of claim 1, wherein the one or more measurements are characterizations of one or more properties of the polypeptide selected from the group consisting of: (i) a length of the polypeptide; (ii) an identity of the polypeptide; (iii) a sequence of the polypeptide; (iv) a secondary structure of the polypeptide; and (v) whether the polypeptide is modified.

17. The method of claim 1, wherein the step of taking one or more measurements that characterize the polypeptide comprises measuring ionic current flowing through the nanopore as the conjugate moves relative to the nanopore.

18. The method of claim 1, wherein the transmembrane protein pore is a transmembrane protein pore derived from or based on Msp, a-hemolysin, CsgG, ClyA, Sp1, or FraC.

19. The method of claim 18, (i) wherein the transmembrane protein pore is a transmembrane protein pore derived from MspA; or (ii) wherein the transmembrane protein pore is a transmembrane protein pore derived from CsgG.

20. The method of claim 1, wherein the nanopore comprises a constriction.

21. A kit comprising: - a nanopore comprising a constriction, wherein the nanopore is a transmembrane protein pore; - a polynucleotide comprising a reactive functional group for conjugation to a target polypeptide; and - a polynucleotide-handling protein capable of controlling movement of the polynucleotide relative to the nanopore.

Citation Information

Patent Citations

  • phi 29 DNA polymerase

    US5576204A

  • A miniature support for thin films containing single channels or nanopores and methods for using same

    WO2000028312A1

  • Deliver of molecules to a li id bila

    WO2006100484A2

  • Lipid bilayer sensor system

    WO2008102120A1

  • Formation of lipid bilayers

    WO2008102121A1