mutant pores
By modifying the amino acids of cytolysin monomers to form mutant cytolysin monomers, the problems of difficulty in distinguishing nucleotides and high current changes in nucleic acid sequencing are solved, and more efficient nucleotide recognition and sequencing performance is achieved.
Patent Information
- Application Number
- CN202410563199.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2016-05-11
- Filing Date
- 2017-04-06
- Publication Date
- 2026-01-06
- Estimated Expiration
- 2037-04-06
AI Technical Summary
Existing nanopore sensing technology suffers from problems such as difficulty in distinguishing nucleotides, high current changes, and low signal-to-noise ratio in nucleic acid sequencing, resulting in insufficient performance of sequencing systems.
By modifying the cytolysin monomer with a series of amino acids, mutant cytolysin monomers are formed, which enhance their interaction with polynucleotides, improve their ability to distinguish nucleotides, and reduce changes in current state.
The mutant cytolysin monomer significantly improved nucleotide discrimination ability, increased current range, reduced current state changes, and improved signal-to-noise ratio, making it easier to identify polynucleotide sequences.
Smart Images

Figure BDA0004828401700000101 
Figure BDA0004828401700000111 
Figure BDA0004828401700000131
Abstract
Description
[0001] This application is a divisional application filed again with respect to divisional application 202310422087.0. Divisional application 202310422087.0 is a divisional application filed on April 6, 2017, with application number 201780022553.9 and the invention title "Mutant Pore". Technical Field
[0002] This invention relates to mutant forms of lysenin. It also relates to analyte characterization using said mutant forms of lysenin. Background Technology
[0003] Nanopore sensing is a sensing method that relies on the observation of individual binding or interaction events between analyte molecules and acceptors. Nanopore sensors can be generated by placing nanoscale monopores in an insulating membrane and measuring voltage-driven ion transport through the pores in the presence of analyte molecules. The identity of the analyte is revealed by its unique current characteristics, particularly the duration and extent of the current block and changes in the current level. Such nanopore sensors are commercially available, for example, the MinION sold by Oxford Nanopore Technologies Ltd. TM The device includes an array of nanopores integrated within an electronic chip.
[0004] There is a current need for rapid and inexpensive nucleic acid (e.g., DNA or RNA) sequencing technologies across a wide range of applications. Existing technologies are slow and expensive, primarily because they rely on amplification techniques to generate large quantities of nucleic acids and require significant amounts of specialized fluorescent chemicals for signal detection. Nanopore sensing has the potential to provide rapid and inexpensive nucleic acid sequencing by reducing the amount of nucleotides and reagents required.
[0005] One of the fundamental elements of using nanopore sensing for nucleic acid sequencing is controlling the movement of nucleic acids through the pore. Another element is distinguishing nucleotides as nucleic acid polymers move through the pore. In the past, to achieve nucleotide differentiation, nucleic acids have been passed through mutants of hemolysin. This has provided current characteristics that have shown sequence dependence. It has also been shown that when using hemolysin pores, a large number of nucleotides contribute to the observed current, making the direct relationship between the observed current and polynucleotides challenging.
[0006] While the range of currents used for nucleotide differentiation has been improved through mutations in the hemolysin wells, further increases in the current differences between nucleotides would result in even higher performance for the sequencing system. Additionally, some current states have shown high changes as nucleic acids move through the wells. Some mutant hemolysin wells have also exhibited higher changes than other mutant hemolysins. While these state changes may contain sequence-specific information, it is desirable to produce wells with low changes to simplify the system. It is also desirable to reduce the number of nucleotides contributing to the observed currents.
[0007] Cytolysin (also known as efL1) is a pore-forming toxin purified from the coelomic fluid of the earthworm *Eisenia fetida*. It specifically binds to sphingomyelin, which inhibits cytolysin-induced hemolysis (Yamaji et al., *J. Biol. Chem.*, 1998, Vol. 273, No. 9, pp. 5300–5306). The crystal structure of the cytolysin monomer is disclosed in *Structure*, 2012, Vol. 20, pp. 1498–1507, by De Colbis et al. Summary of the Invention
[0008] The inventors have surprisingly discovered a novel mutant cytolysin monomer, in which one or more modifications have been made to enhance the monomer's ability to interact with polynucleotides. The inventors have also surprisingly demonstrated that pores incorporating the novel mutant monomer possess enhanced ability to interact with polynucleotides and thus exhibit improved properties for estimating polynucleotide characteristics such as their sequence. The mutant pores unexpectedly exhibit improved nucleotide differentiation. Specifically, the mutant pores unexpectedly exhibit an increased current range and reduced state changes; the increased current range makes it easier to distinguish different nucleotides, and the reduced state changes increase the signal-to-noise ratio. Furthermore, the number of nucleotides contributing to the current as the polynucleotide moves through the pore is reduced. This makes it easier to identify the direct relationship between the observed current and the polynucleotide sequence as the polynucleotide moves through the pore.
[0009] Unless otherwise stated, all amino acid substitutions, deletions and / or additions disclosed herein refer to mutant cytolysin monomers including variants of the sequence shown in SEQ ID NO: 2.
[0010] References to mutant cytolysin monomers including variants of the sequence shown in SEQ ID NO: 2 encompass mutant cytolysin monomers including variants of the sequences described in SEQ ID NO: 14 to 16. Cytolysin monomers including variants of the sequence shown in SEQ ID NO: 2 may be modified with amino acid substitutions, deletions, and / or additions equivalent to those disclosed herein with reference to SEQ ID NO: 2.
[0011] Mutant monomers can be considered as isolated monomers.
[0012] Therefore, the present invention provides a mutant cytosolic monomer comprising a variant of the sequence shown in SEQ ID NO: 2, wherein the monomer is capable of forming pores, and wherein the variant comprises modifications at one or more of the following positions: K37, G43, K45, V47, S49, T51, H83, V88, T91, T93, V95, Y96, S98, K99, V100, I101, P108, P109, T110, S111, K112, and T114.
[0013] The present invention also provides a mutant cytolysin monomer comprising a variant of the sequence shown in SEQ ID NO: 2, wherein the monomer is capable of forming pores, and wherein the variant comprises one or more of the following substitutions:
[0014] D35N / S;
[0015] S74K / R;
[0016] E76D / N;
[0017] S78R / K / N / Q;
[0018] S80K / R / N / Q;
[0019] S82K / R / N / Q;
[0020] E84R / K / N / A;
[0021] E85N;
[0022] S86K / Q;
[0023] S89K;
[0024] M90K / I / A;
[0025] E92D / S;
[0026] E94D / Q / G / A / K / R / S / N;
[0027] E102N / Q / D / S;
[0028] T104R / K / Q;
[0029] T106R / K / Q;
[0030] R115S;
[0031] Q117S; and
[0032] N119S.
[0033] The present invention also provides a mutant cytolysin monomer comprising a variant of the sequence shown in SEQ ID NO: 2, wherein the monomer is capable of forming pores, and wherein the variant comprises mutations at one or more of the following:
[0034] D35 / E94 / T106;
[0035] K37 / E94 / E102 / T106;
[0036] K37 / E94 / T104 / T106;
[0037] K37 / E94 / T106;
[0038] K37 / E94 / E102 / T106;
[0039] G43 / E94 / T106;
[0040] K45 / V47 / E92 / E94 / T106;
[0041] K45 / V47 / E94 / T106;
[0042] K45 / S49 / E92 / E94 / T106;
[0043] K45 / S49 / E94 / T106;
[0044] K45 / E94 / T106;
[0045] K45 / T106;
[0046] V47 / E94 / T106;
[0047] V47 / V88 / E94 / T106;
[0048] S49 / E94 / T106;
[0049] T51 / E94D / T106;
[0050] S74 / E94;
[0051] E76 / E94;
[0052] S78 / E94;
[0053] Y79 / E94;
[0054] S80 / E94;
[0055] S82 / E94;
[0056] S82 / E94 / T106;
[0057] H83 / E94;
[0058] H83 / E94 / T106;
[0059] E85 / E94 / T106;
[0060] S86 / E94;
[0061] V88 / M90 / E94 / T106;
[0062] S89 / E94;
[0063] M90 / E94 / T106;
[0064] T91 / E94 / T106;
[0065] E92 / E94 / T106;
[0066] T93 / E94 / T106;
[0067] E94 / Y96 / T106;
[0068] E94 / S98 / K99 / T106;
[0069] E94 / K99 / T106;
[0070] E94 / E102;
[0071] E94 / T104;
[0072] E94 / T106;
[0073] E94 / P108;
[0074] E94 / P109;
[0075] E94 / T110;
[0076] E94 / S111;
[0077] E94 / T114;
[0078] E94 / R115;
[0079] E94 / Q117; and
[0080] E94 / E119.
[0081] The present invention also provides a mutant cytolysin monomer comprising a variant of the sequence shown in SEQ ID NO: 2, wherein the monomer is capable of forming pores, and wherein the variant comprises one or more of the following substitutions:
[0082] E84R / E94D;
[0083] E84K / E94D;
[0084] E84N / E94D;
[0085] E84A / E94Q;
[0086] E84K / E94Q and
[0087] E94Q / D121S.
[0088] The present invention also provides a mutant cytolysin monomer comprising a variant of the sequence shown in SEQ ID NO: 2, wherein the variant comprises one of the following substitution combinations:
[0089] -E84Q / E85K / E92Q / E94D / E97S / D126G;
[0090] -E84Q / E85K / E92Q / E94Q / E97S / D126G; or
[0091] -E84Q / E85K / E92Q / E94D / E97S / T106K / D126G.
[0092] The present invention also provides a mutant cytolysin monomer comprising a variant of the sequence shown in SEQ ID NO: 2, wherein in the variant, (a) 2, 4, 6, 8, 10, 12, 14, 16, 18, or 20 amino acids at positions 34 to 70 of SEQ ID NO: 2 or corresponding to those positions are deleted, and (b) 2, 4, 6, 8, 10, 12, 14, 16, 18, or 20 amino acids at positions 71 to 107 of SEQ ID NO: 2 or corresponding to those positions are deleted.
[0093] The present invention also provides:
[0094] - A construct comprising two or more covalently linked monomers derived from cytolysin, wherein at least one of the monomers is a mutant cytolysin monomer of the present invention;
[0095] - A polynucleotide that encodes the mutant cytolysin monomer of the present invention or the gene fusion construct of the present invention;
[0096] - A homooligomeric pore derived from cytolysin, comprising a sufficient number of mutant cytolysin monomers of the present invention;
[0097] - A heterooligoporous pore derived from cytolysin, comprising at least one mutant cytolysin monomer of the present invention;
[0098] - A hole comprising at least one construct of the present invention;
[0099] - A method for characterizing a target analyte, comprising: (a) contacting the target analyte with an aperture of the present invention, such that the target analyte moves through the aperture; and (b) acquiring one or more measurements as the analyte moves relative to the aperture, wherein the measurements indicate one or more properties of the target analyte and thereby characterize the target analyte.
[0100] - A method for forming a sensor for characterizing a target polynucleotide, comprising forming a complex between a pore of the invention and a polynucleotide-binding protein and thereby forming a sensor for characterizing the target polynucleotide;
[0101] - A sensor for characterizing target polynucleotides, comprising a complex between the pore of the present invention and a polynucleotide-binding protein;
[0102] - The use of the pores in this invention for characterizing target analytes;
[0103] - A kit for characterizing target polynucleotides, comprising (a) the wells of the present invention and (b) a membrane;
[0104] - An apparatus for characterizing target polynucleotides in a sample, comprising (a) a plurality of pores of the present invention and (b) a plurality of polynucleotide-binding proteins;
[0105] - A method for improving the ability of a cytolysin monomer comprising the sequence shown in SEQ ID NO: 2 to characterize a polynucleotide, comprising: performing one or more modifications and / or substitutions according to the present invention;
[0106] - A method for generating a construct of the present invention, comprising: covalently linking at least one mutant cytolysin monomer of the present invention to one or more monomers derived from cytolysin; and
[0107] - A method of forming pores according to the present invention, comprising: allowing at least one mutant monomer of the present invention or at least one construct of the present invention to oligomerize with a sufficient number of monomers of the present invention, constructs of the present invention or monomers derived from cytosine to form pores. Attached Figure Description
[0108] Figure 1 The median plot of cytolysin mutant 1 is shown.
[0109] Figure 2 The median plot of cytolysin mutant 10 is shown.
[0110] Figure 3 The median plot of the cytolysin mutant - cytolysin-(E84Q / E85K / E92Q / E94D / E97S / T106K / D126G / C272A / C283A)9 (SEQ ID NO: 2 with the mutation E84Q / E85K / E92Q / E94D / E97S / T106K / D126G / C272A / C283A) is shown.
[0111] Figure 4 The median plot of cytolysin mutant-cytolysin-with 2-iodo-N-(2,2,2-trifluoroethyl)acetamide attached via E94C (SEQ ID NO: 2 with mutant E84Q / E85K / E92Q / E94C / E97S / T106K / D126G / C272A / C283A) is shown.
[0112] Figure 5 The connectives used in the examples are shown. A corresponds to 30 iSpC3. B corresponds to SEQ IN NO: 19. C corresponds to 4 iSp18. D corresponds to SEQ ID NO: 20. E corresponds to SEQ ID NO: 21, which has 5BNA-G / / iBNA-G / / iBNA-T / / iBNA-T / / i-BNA-A attached to its 5' end. F corresponds to SEQ ID NO: 22, which has a 5' phosphate ester. G corresponds to SEQ ID NO: 24. H corresponds to cholesterol.
[0113] Figure 6The 3D structure of the cytolysin monomer is shown. Upon interaction with a sphingomyelin-containing membrane, the cytolysin monomer assembles together through a central anterior pore to form a nonameric pore. During the assembly process, the polypeptide segment shown in black (corresponding to amino acids 65 to 74 of SEQ ID NO: 2) transforms into the bottom loop of the β-barrel shown in Figure 7. Two β-sheets on either side of the polypeptide segment shown in black, as well as polypeptide segments (corresponding to amino acids 34 to 64 and 75 to 107 of SEQ ID NO: 2) linking those β-sheets to the polypeptide segment shown in black, extend to form the β-barrel of the pore, as shown in Figure 7. This large structural change makes it difficult to predict the β-barrel region of the cytolysin pore by studying the monomer structure.
[0114] Figure 7 depicts the region of the cytolysin pore. Figure 7A The 3D structure of the nonameric pore of cytolysin is shown, and Figure 7B The structure of the monomers derived from the cytolysin pores is shown. Each monomer contributes two β-sheets to the barrel of the cytolysin pore. The β-sheets (containing amino acids corresponding to amino acids 34 to 64 and 75 to 107 of SEQ ID NO: 2) are linked by an unstructured loop at the bottom of the pore (amino acids corresponding to positions 65 to 74 of SEQ ID NO: 2).
[0115] Figure 8 This is an alignment of the amino acid sequence of cytolysin (SEQ ID NO: 2) with the amino acid sequences of three cytolysin-associated proteins (SEQ ID NO: 14 to 16). Three cytolysin homologues with sequences closely related to cytolysin were identified using a BLAST search performed using a database of non-redundant protein sequences. The protein sequences of cytolysin-associated protein 1 (LRP1), cytolysin-associated protein 2 (LRP2), and cytolysin-associated protein 3 (LRP3) were aligned with the sequence of cytolysin to show the similarity of the four proteins. Dark gray shading indicates the position of the consistent amino acid in all four sequences. LRP1 has approximately 75% similarity to cytolysin, LRP2 approximately 88%, and LRP3 approximately 79%.
[0116] Sequence List Description
[0117] SEQ ID NO: 1 shows a polynucleotide sequence encoding a cytosolic monomer.
[0118] SEQ ID NO: 2 shows the amino acid sequence of the cytolysin monomer.
[0119] SEQ ID NO: 3 shows the polynucleotide sequence encoding Phi29 DNA polymerase.
[0120] SEQ ID NO: 4 shows the amino acid sequence of Phi29 DNA polymerase.
[0121] SEQ ID NO: 5 shows a codon-optimized polynucleotide sequence derived from the sbcB gene of *E. coli*. It encodes an exonuclease I (EcoExo I) from *E. coli*.
[0122] SEQ ID NO: 6 shows the amino acid sequence of an exonuclease I (EcoExo I) from Escherichia coli.
[0123] SEQ ID NO: 7 shows a codon-optimized polynucleotide sequence derived from the xthA gene of *E. coli*. It encodes an exonuclease III enzyme derived from *E. coli*.
[0124] SEQ ID NO: 8 shows the amino acid sequence of an exonuclease III enzyme from Escherichia coli. This enzyme performs partitioning digestion of 5' monophosphate nucleotides from one strand of double-stranded DNA (dsDNA) in the 3' to 5' direction. Enzyme initiation on the strand requires approximately 4 nucleotides of 5' protrusion.
[0125] SEQ ID NO: 9 shows a codon-optimized polynucleotide sequence derived from the recJ gene of thermophilic bacteria. It encodes the RecJ enzyme (TthRecJ-cd) from thermophilic bacteria.
[0126] SEQ ID NO: 10 shows the amino acid sequence of a RecJ enzyme (TthRecJ-cd) from thermophilic bacteria. This enzyme performs progressive digestion of 5' monophosphate nucleotides from ssDNA in the 5' to 3' direction. Initiation of the enzyme on the strand requires at least 4 nucleotides.
[0127] SEQ ID NO: 11 shows a codon-optimized polynucleotide sequence derived from the phage λexo(redX) gene. It encodes the phage λ exonuclease.
[0128] SEQ ID NO: 12 shows the amino acid sequence of a bacteriophage λ exonuclease. The sequence is one of three identical subunits that assemble into a trimer. The enzyme performs highly progressive digestion of nucleotides from one strand of dsDNA in the 5' to 3' direction (http: / / www.neb.com / nebecomm / products / productM0262.asp). Enzyme initiation on the strand preferentially requires a 5' overhang of approximately four nucleotides with a 5' phosphate ester.
[0129] SEQ ID NO: 13 shows the amino acid sequence of Hel308 Mbu.
[0130] SEQ ID NO: 14 shows the amino acid sequence of cytolysin-associated protein (LRP) 1.
[0131] SEQ ID NO: 15 shows the amino acid sequence of cytolysin-associated protein (LRP)2.
[0132] SEQ ID NO: 16 shows the amino acid sequence of cytolysin-associated protein (LRP)3.
[0133] SEQ ID NO: 17 shows the amino acid sequence of the activated version of parasporin-2. The full-length protein is cleaved at its amino and carboxyl ends to form the activated version capable of forming pores.
[0134] SEQ ID NO: 18 shows the amino acid sequence of Dda 1993.
[0135] SEQ ID NO: 19 to 24 show the polynucleotide sequences used in the examples. Detailed Implementation
[0136] It should be understood that different applications of the disclosed products and methods can be tailored to the specific needs of the respective fields. It should also be understood that the terminology used herein is for the purpose of describing specific embodiments of the invention only and is not intended to be limiting.
[0137] Furthermore, unless otherwise expressly stated, the singular forms “a / an” and “the” as used in this specification and the appended claims include plural indicators. Thus, for example, reference to “a mutant monomer” includes “a plurality of mutant monomers,” reference to “a substitution” includes two or more such substitutions, reference to “a pore” includes two or more such pores, reference to “a polynucleotide” includes two or more such polynucleotides, and so on.
[0138] In this specification, when different amino acids are separated by the symbol " / " at a specific position, the " / " symbol means "or". For example, P108R / K means P108R or P108K. In this specification, when different positions or different substitutions are separated by the symbol " / ", the " / " symbol means "and". For example, E94 / P108 means E94 and P108, or E94D / P108K means E94D and P108K.
[0139] All publications, patents and patent applications cited in this document (both above and below) are hereby incorporated in their entirety by reference.
[0140] Mutant cytolysin monomer
[0141] In one aspect, the present invention provides mutant cytolysin monomers. Mutant cytolysin monomers can be used to form pores according to the present invention. A mutant cytolysin monomer is a monomer whose sequence differs from that of a wild-type cytolysin monomer (e.g., SEQ ID NO: 2, SEQ ID NO: 14, SEQ ID NO: 15, or SEQ ID NO: 16). Mutant cytolysin monomers generally retain the ability to form pores in the presence of other monomers of the present invention or other monomers derived from or containing cytolysin. Therefore, mutant monomers are generally capable of forming pores. Methods for confirming the ability of mutant monomers to form pores are well known in the art and are described in examples. For example, pore formation can be determined by electrophysiology. The pore is typically inserted into a membrane, which may be, for example, a lipid membrane or a block copolymer membrane. Electrical or optical measurements can be obtained from monocytolysin pores inserted into the membrane, such as pores comprising one or more monomers of the present invention. A potential difference can be applied across the membrane, and a current through the membrane can be detected. The current can be detected by any suitable method, such as by electrical or optical means. Pore transfer polynucleotides, excellent Selected single-stranded polynucleotides The ability can be determined by adding a premixture of polynucleotide-binding protein, DNA, and fuel (e.g., MgCl2, ATP); applying a potential difference (e.g., 180 mV); and monitoring the current through the pore to detect DNA movement controlled by the polynucleotide-binding protein.
[0142] When present in a pore, the mutant monomer has an altered ability to interact with polynucleotides. Therefore, pores containing one or more of the mutant monomers exhibit enhanced nucleotide readout properties, for example, showing (1) enhanced polynucleotide capture and (2) enhanced polynucleotide recognition or differentiation. Specifically, pores constructed from mutant monomers capture nucleotides and polynucleotides more readily than wild-type pores. Additionally, pores constructed from mutant monomers exhibit an increased current range and reduced state changes; the increased current range makes it easier to distinguish different nucleotides, and the reduced state changes increase the signal-to-noise ratio. Furthermore, the number of nucleotides contributing to the current as polynucleotides move through pores constructed from mutant monomers is reduced. This makes it easier to identify a direct relationship between the observed current and the polynucleotide sequence as polynucleotides move through the pore. The enhanced nucleotide readout properties of mutants are achieved through five main mechanisms, namely, through changes in the following:
[0143] • Steric hindrance (increasing or decreasing the size of amino acid residues);
[0144] • Charge (e.g., introducing or removing -ve charge and / or introducing or removing +ve charge);
[0145] • Hydrogen binding (e.g., introducing an amino acid that can bind hydrogen to a base pair);
[0146] • π stacking (e.g., introducing amino acids that interact through delocalized electron π systems); and / or
[0147] • Changes in pore structure (e.g., introduction of amino acids that increase the size of barrels or channels).
[0148] Any one or more of these five mechanisms may be the cause of the improved properties of the pores formed by the mutant monomers of the present invention. For example, pores comprising the mutant monomers of the present invention may exhibit improved nucleotide readout properties due to altered steric hindrance, altered hydrogen binding, and altered structure.
[0149] The mutant monomers of the present invention include variants of the sequence shown in SEQ ID NO: 2. SEQ ID NO: 2 is the wild-type sequence of a cytolysin monomer. Variant versions of SEQ ID NO: 2 are polypeptides with amino acid sequences different from those of SEQ ID NO: 2. Typically, the variants retain their ability to form pores.
[0150] Wells containing one or more mutant monomers with substitutions in S80, T106, and T104 exhibit enhanced polynucleotide capture. Specific examples of such substitutions include S80K / R, T104R / K, and T106R / K. Other substitutions that increase the positive charge of the amino acid side chains at any one or more of these positions, such as 2, 3, 4, or 5, can be used to enhance the properties of wells containing mutant monomers, i.e., to improve polynucleotide capture compared to wild-type wells or wells containing mutant monomers with other capture-enhancing mutations, such as E84Q / E85K / E92Q / E97S / D126G, for example, wells containing only mutant monomers with those mutations or wells containing mutant monomers with the following mutations: E84Q / E85K / E92Q / E94D / E97S / D126G. Typically, when an improvement is determined relative to wells containing other mutations, such as E84Q / E85K / E92Q / E97S / D126G or E84Q / E85K / E92Q / E94D / E97S / D126G, those mutations are also present in the mutant monomer being tested; that is, one or more effects of the mutation or a combination of mutations are determined relative to a baseline monomer / well consistent with the monomer / well being tested, rather than at one or more test sites. The nature of the wells containing the mutant monomer or control monomer can be determined using heterooligomeric wells or, more preferably, homooligomeric wells. Examples of preferred mutation combinations are described throughout the specification, for example, in Table 9.
[0151] Wells containing one or more of the mutant monomers with substitutions at D35, K37, K45, V47, S49, E76, S78, S82, V88, S89, M90, T91, E92, E94, Y96, S98, V100, and T104 exhibited enhanced polynucleotide recognition or differentiation. Specific examples of such substitutions include D35N, K37N / S, K45R / K / D / T / Y / N, V47K / R, S49K / R / L, T51KE76S / N, S78N, S82N, V88I, S89Q, M90I / A, T91S, E92D / E, E94D / Q / N, Y96D, S98Q, V100S, and T104K. As described in Table 9, these mutations can each reduce noise, increase current range, and / or reduce channel gating. The specified positions in SEQ ID NO: 2 or the corresponding positions in variants of SEQ ID NO: 2 were subjected to other mutations that increased or decreased the size of the amino acid side chains, increased or decreased the charge, resulted in the same hydrogen bond formation, and / or affected π stacking in the same manner as any one or more of these exemplary mutations. Mutations can be introduced individually or in combination. For example, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, or 18 mutations at these positions can be introduced to improve the properties of pores containing mutant monomers, i.e., improve the signal-to-noise ratio, increase the range, and / or reduce channel gating, making them more suitable for wild-type pores, including those containing the mutations E84Q / E85K / E92QE97S / D126G. Polynucleotide recognition and differentiation are improved in wells containing only mutant monomers, mutant monomers containing E84Q / E85K / E92Q / E94D / E97S / D126G, mutant monomers containing E84Q / E85K / E92Q / E94Q / E97S / D126G, and / or mutant monomers containing E84Q / E85K / E92Q / E94D / E97S / T106K / D126G. Typically, when an improvement is determined relative to wells containing other mutations, such as E84Q / E85K / E92Q / E97S / D126G, E84Q / E85K / E92Q / E94D / E97S / D126G, E84Q / E85K / E92Q / E94Q / E97S / D126G, or E84Q / E85K / E92Q / E94D / E97S / T106K / D126G, those mutations are also present in the tested mutant monomer; that is, one or more effects of the mutation, or a combination of mutations, are determined relative to a baseline monomer / well consistent with the monomer / well being tested, rather than at one or more test sites. The nature of the wells containing the mutant monomer or control monomer can be determined using heterooligomeric wells or, more preferably, homooligomeric wells.Examples of preferred mutation combinations are described throughout the specification, for example in Table 9.
[0152] Compared to wild-type pores or pores comprising mutant monomers containing the mutations E84Q / E85K / E92QE97S / D126G, pores comprising one or more mutant monomers containing substitutions at E94 and / or Y96 can reduce the number of nucleotides contributing to the current as the polynucleotide moves through the pore. For example, substitutions can be made for Y96D / E, preferably in combination with E94Q / D, to reduce the size of the read head. The reduction in the number of nucleotides contributing to the current as the polynucleotide moves through the pore, compared to wild-type pores or pores comprising mutant monomers containing the mutations E84Q / E85K / E92QE97S / D126G, can also be achieved by deleting an even number of amino acids (typically, amino acids present in the lumen of the pore and adjacent amino acids facing away from the lumen of the pore) from each of the two β chains of a portion of the pore-forming barrel of the monomer, i.e., positions corresponding to amino acids 34 to 65 and 74 to 107 of SEQ ID NO: 2, as described herein.
[0153] Modifications of the present invention
[0154] This invention provides a mutant cytolysin monomer wherein the amino acid sequence contributing to the structure of the barrel in the cytolysin pore is modified compared to wild-type cytolysin and compared to cytolysin mutants disclosed in the art, for example, in WO 2013 / / 153359. The modification of this invention is made in the region of the cytolysin monomer corresponding to amino acids 34 to 107 of SEQ ID NO: 2, specifically amino acids 34 to 65 and 74 to 107 of SEQ ID NO: 2. The corresponding regions of the LR1, LR2, and LR3 monomers are shown in... Figure 8 In the comparison.
[0155] Therefore, the present invention provides a mutant cytolysin monomer comprising a variant of the sequence shown in SEQ ID NO: 2, wherein the monomer is capable of forming pores, and wherein the variant comprises modifications at one or more of the following positions: K37, G43, K45, V47, S49, T51, H83, V88, T91, T93, V95, Y96, S98, K99, V100, I101, P108, P109, T110, S111, K112, and T114. The variant may comprise modifications at any number of the stated positions and any combination of the stated positions. In one aspect, the modification may be a substitution, deletion, or addition of amino acids, and preferably a substitution or deletion mutation. Preferred modifications are discussed below under the heading “Further Modifications.” The mutant cytolysin monomer may include modifications at other positions of SEQ ID NO: 2. For example, in addition to one or more modifications of the present invention, such as 2 to 20, 3 to 15, 4 to 10, or 6 to 8, the mutant cytolysin monomer may have one or more amino acid substitutions or deletions in the sequence of SEQ ID NO: 2 that are described in the relevant field, such as in WO 2013 / 153359, such as 2 to 20, 3 to 15, 4 to 10, or 6 to 8.
[0156] The variant preferably includes modifications at one or more of the following positions: T91, V95, Y96, S98, K99, V100, I101, and K112. The variant may have modifications at any number of these positions and any combination of these positions. The modifications are preferably substitutions using serine (S) or glutamine (Q). The variant preferably includes substitutions at one or more of T91S, V95S, Y96S, S98Q, K99S, V100S, I101S, and K112S. The variant may include any number of these substitutions and any combination of these substitutions.
[0157] The variant preferably includes modifications at one or more of the following positions: K37, G43, K45, V47, S49, T51, H83, V88, T91, T93, Y96, S98, K99, P108, P109, T110, S111, and T114. The variant may include modifications at any number of the stated positions and any combination of the stated positions. The modifications are preferably substitutions using asparagine (N), tryptophan (W), serine (S), glutamine (Q), lysine (K), aspartic acid (D), arginine (R), threonine (T), tyrosine (Y), leucine (L), or isoleucine (I). The variant preferably includes one or more of the following substitutions: K37N / W / S / Q, G43K, K45D / R / N / Q / T / Y, V47K / S / N, S49K / L, T51K, H83S / K, V88I / T, T91K, T93K, Y96D, S98K, K99Q / L, P108K / R, P109K, T110K / R, S111K, and T114K. The variant preferably includes modifications at one or more of the following positions:
[0158] The variant preferably includes one or more of the following substitutions:
[0159]
[0160] The present invention also provides a mutant cytolysin monomer comprising a variant of the sequence shown in SEQ ID NO: 2, wherein the monomer is capable of forming pores, and wherein the variant comprises one or more of the following substitutions:
[0161] D35N / S;
[0162] S74K / R;
[0163] E76D / N;
[0164] S78R / K / N / Q;
[0165] S80K / R / N / Q;
[0166] S82K / R / N / Q;
[0167] E84R / K / N / A;
[0168] E85N;
[0169] S86K / Q;
[0170] S89K;
[0171] M90K / I / A;
[0172] E92D / S;
[0173] E94D / Q / G / A / K / R / S / N;
[0174] E102N / Q / D / S;
[0175] T104R / K / Q;
[0176] T106R / K / Q;
[0177] R115S;
[0178] Q117S; and
[0179] N119S.
[0180] The variants may include any number of these substitutions and any combination of these substitutions. The variants preferably include one or more of the following substitutions: E94D / Q / G / A / K / R / S, S86Q, and E92S, such as E94D / Q / G / A / K / R / S; S86Q; E92S; E94D / Q / G / A / K / R / S and S86Q; E94D / Q / G / A / K / R / S and E92S; S86Q and E92S; or E94D / Q / G / A / K / R / S, S86Q, and E92S.
[0181] The variant preferably includes one or more of the following substitutions:
[0182] D35N / S;
[0183] S74K / R;
[0184] E76D / N;
[0185] S78R / K / N / Q;
[0186] S80K / R / N / Q;
[0187] S82K / R / N / Q;
[0188] E84R / K / N / A;
[0189] E85N;
[0190] S86K;
[0191] S89K;
[0192] M90K / I / A;
[0193] E92D;
[0194] E94D / Q / K / N;
[0195] E102N / Q / D / S;
[0196] T104R / K / Q;
[0197] T106R / K / Q;
[0198] R115S;
[0199] Q117S; and
[0200] N119S.
[0201] The variants may include any number of these substitutions and combinations thereof.
[0202] The variant preferably includes one or more of the following substitutions:
[0203]
[0204]
[0205] The variants may include any number of these substitutions and any combination of these substitutions.
[0206] The present invention also provides a mutant cytolysin monomer comprising a variant of the sequence shown in SEQ ID NO: 2, wherein the monomer is capable of forming pores, and wherein the variant comprises mutations at one or more of the following:
[0207] D35 / E94 / T106;
[0208] K37 / E94 / E102 / T106;
[0209] K37 / E94 / T104 / T106;
[0210] K37 / E94 / T106;
[0211] K37 / E94 / E102 / T106;
[0212] G43 / E94 / T106;
[0213] K45 / V47 / E92 / E94 / T106;
[0214] K45 / V47 / E94 / T106;
[0215] K45 / S49 / E92 / E94 / T106;
[0216] K45 / S49 / E94 / T106;
[0217] K45 / E94 / T106;
[0218] K45 / T106;
[0219] V47 / E94 / T106;
[0220] V47 / V88 / E94 / T106;
[0221] S49 / E94 / T106;
[0222] T51 / E94D / T106;
[0223] S74 / E94;
[0224] E76 / E94;
[0225] S78 / E94;
[0226] Y79 / E94;
[0227] S80 / E94;
[0228] S82 / E94;
[0229] S82 / E94 / T106;
[0230] H83 / E94;
[0231] H83 / E94 / T106;
[0232] E85 / E94 / T106;
[0233] S86 / E94;
[0234] V88 / M90 / E94 / T106;
[0235] S89 / E94;
[0236] M90 / E94 / T106;
[0237] T91 / E94 / T106;
[0238] E92 / E94 / T106;
[0239] T93 / E94 / T106;
[0240] E94 / Y96 / T106;
[0241] E94 / S98 / K99 / T106;
[0242] E94 / K99 / T106;
[0243] E94 / E102;
[0244] E94 / T104;
[0245] E94 / T106;
[0246] E94 / P108;
[0247] E94 / P109;
[0248] E94 / T110;
[0249] E94 / S111;
[0250] E94 / T114;
[0251] E94 / R115;
[0252] E94 / Q117; and
[0253] E94 / E119.
[0254] The variant preferably includes one or more of the following substitutions: D35N / E94D / T106K;
[0255] D35S / E94D / T106K;
[0256] K37Q / E94D / E102N / T106K;
[0257] K37S / E94D / E102S / T106K;
[0258] K37S / E94D / T104K / T106K;
[0259] K37N / E94D / T106K;
[0260] K37W / E94D / T106K;
[0261] K37S / E94D / T106K;
[0262] G43K / E94D / T106K;
[0263] K45N / V47K / E92D / E94N / T106K;
[0264] K45T / V47K / E94D / T106K;
[0265] K45N / S49K / E94N / E92D / T106K;
[0266] K45Y / S49K / E94D / T106K;
[0267] K45D / E94K / T106K;
[0268] K45R / E94D / T106K;
[0269] K45N / E94N / T106K;
[0270] K45Q / E94Q / T106K;
[0271] K45R / T106K;
[0272] V47S / E94D / T106K;
[0273] V47K / E94D / T106K;
[0274] V47N / V88T / E94D / T106K;
[0275] S49L / E94D / T106K;
[0276] T51K / E94D / T106K;
[0277] S74K / E94D;
[0278] S74R / E94D;
[0279] E76D / E94D;
[0280] E76N / E94D;
[0281] E76S / E94Q;
[0282] E76N / E94Q;
[0283] S78R / E94D;
[0284] S78K / E94D;
[0285] S78N / E94D;
[0286] S78Q / E94Q;
[0287] Y79S / E94Q;
[0288] S80K / E94D;
[0289] S80R / E94D;
[0290] S80N / E94D;
[0291] S80Q / E94Q;
[0292] S82K / E94D;
[0293] S82R / E94D;
[0294] S82N / E94D;
[0295] S82Q / E94Q;
[0296] S82K / E94D / T106K;
[0297] H83S / E94Q;
[0298] H83K / E94D / T106K;
[0299] E85N / E94D / T106K;
[0300] S86K / E94D;
[0301] V88I / M90A / E94D / T106K;
[0302] S89K / E94D;
[0303] M90K / E94D / T106K;
[0304] M90I / E94D / T106K;
[0305] T91K / E94D / T106K;
[0306] E92D / E94Q / T106K;
[0307] T93K / E94D / T106K;
[0308] E94Q / Y96D / T106K;
[0309] E94D / S98K / K99L / T106K;
[0310] E94D / K99Q / T106K;
[0311] E94D / E102N;
[0312] E94D / E102Q;
[0313] E94D / E102D;
[0314] E94D / T104R;
[0315] E94D / T104K;
[0316] E94Q / T104Q;
[0317] E94D / T106R;
[0318] E94D / T106K;
[0319] E94Q / T106Q;
[0320] E94Q / T106K;
[0321] E94D / P108K;
[0322] E94D / P108R;
[0323] E94D / P109K;
[0324] E94D / T110K;
[0325] E94D / T110R;
[0326] E94D / S111K;
[0327] E94D / T114K;
[0328] E94Q / R115S;
[0329] E94Q / Q117S; and
[0330] E94Q / N119S.
[0331] The variants may include any number of these substitutions and any combination of these substitutions.
[0332] The present invention also provides a mutant cytolysin monomer comprising a variant of the sequence shown in SEQ ID NO: 2, wherein the monomer is capable of forming pores, and wherein the variant comprises one or more of the following substitutions:
[0333] E84R / E94D;
[0334] E84K / E94D;
[0335] E84N / E94D;
[0336] E84A / E94Q;
[0337] E84K / E94Q and
[0338] E94Q / D121S.
[0339] The variants may include any number of these substitutions and any combination of these substitutions.
[0340] The mutant monomers of the present invention preferably comprise any combination of the modifications and / or substitutions described above. Exemplary combinations are disclosed in the examples.
[0341] Bucket missing
[0342] In another embodiment, the present invention also provides a mutant cytolysin monomer comprising a variant of the sequence shown in SEQ ID NO: 2, wherein in the variant, (a) 2, 4, 6, 8, 10, 12, 14, 16, 18, or 20 amino acids at positions 34 to 70 of SEQ ID NO: 2 are missing, or wherein the amino acids corresponding to positions 34 to 70 of SEQ ID NO: 2 are missing, and (b) 2, 4, 6, 8, 10, 12, 14, 16, 18, or 20 amino acids at positions 71 to 107 of SEQ ID NO: 2 are missing, or wherein the amino acid residues at positions 71 to 107 of SEQ ID NO: 2 are missing.
[0343] The number of amino acids missing from positions 34 to 70 may differ from the number of amino acids missing from positions 71 to 107. Preferably, the number of amino acids missing from positions 34 to 70 is the same as the number of amino acids missing from positions 71 to 107.
[0344] Amino acids from positions 34 to 70 and from positions 71 to 107 may be omitted. The positions of the omitted amino acids are preferably shown in one row of Table 1 or Table 2, or in more than one row of Table 1 and / or Table 2. For example, if D35 and V34 are omitted from positions 34 to 70, then T104 and I105 may be omitted from positions 71 to 107. Similarly, D35, V34, K37, and I38 may be omitted from positions 34 to 70, and E102, H103, T104, and I105 may be omitted from positions 71 to 107. This ensures the β-fold structure liner of the barrel maintaining the pores.
[0345] Table 1
[0346]
[0347]
[0348] Table 2
[0349]
[0350]
[0351]
[0352] Amino acids missing from positions 34 to 70 and positions 71 to 107 do not need to be in a single row of Table 1 or 2. For example, if D35 and V34 are missing from positions 34 to 70, then I72 and E71 can be missing from positions 71 to 107.
[0353] The amino acids deleted from positions 34 to 70 are preferably consecutive. The amino acids deleted from positions 71 to 107 are preferably consecutive. The amino acids deleted from positions 34 to 70 and from positions 71 to 107 are preferably consecutive.
[0354] The present invention preferably provides a mutant monomer, wherein the following are missing:
[0355] (i) N46 / V47 / T91 / T92; or
[0356] (ii)N48 / S49 / T91 / T92.
[0357] Those skilled in the art can identify other combinations of amino acids that may be omitted according to the present invention. The following discussion uses the residue numbers in SEQ ID NO: 2 (i.e., before any amino acid is omitted as described above).
[0358] The barrel deletion variant further preferably includes any modifications and / or substitutions discussed above or below, where appropriate. "Where appropriate" means whether the position remains in the mutant monomer after the barrel deletion.
[0359] Chemical modification
[0360] In another aspect, the present invention provides chemically modified mutant cytosolic monomers. The mutant monomer can be any of the mutant monomers discussed above or below. Therefore, the mutant monomers of the present invention, such as variants of SEQ ID NO: 2 including modifications at one or more of the following positions: K37, G43, K45, V47, S49, T51, H83, V88, T91, T93, V95, Y96, S98, K99, V100, I101, P108, P109, T110, S111, K112, and T114, or variants including barrel deletions of the above, can be chemically modified according to the present invention, as discussed below.
[0361] Any further modifications, including those discussed below, can be chemically modified in mutant monomers that alter the ability of the monomer or preferably the region described herein to interact with polynucleotides, within the region of SEQ ID NO: 2 from approximately position 44 to approximately position 126. These chemically modified monomers do not need to include the modifications of the present invention, i.e., they do not need to include modifications at one or more of the following positions: K37, G43, K45, V47, S49, T51, H83, V88, T91, T93, V95, Y96, S98, K99, V100, I101, P108, P109, T110, S111, K112, and T114. The chemically modified mutant monomer preferably comprises a variant of SEQ ID NO: 2, said variant comprising substitutions at one or more of the following positions of SEQ ID NO: 2: (a) E84, E85, E92, E97, and D126; (b) E85, E97, and D126; or (c) E84 and E92. Any number of substitutions discussed below or any combination thereof may be made.
[0362] The mutant monomer can be chemically modified in any way to reduce or shrink the diameter of the barrels or channels formed by the monomer. This is discussed in more detail below.
[0363] Chemical modification is performed to preferably covalently link chemical molecules to mutant monomers. Any method known in the art can be used to covalently link chemical molecules to mutant monomers. Chemical molecules are typically attached via chemical linkages.
[0364] Preferably, the mutant monomer is chemically modified by enzymatic modification, such as attaching the molecule to one or more cysteine residues (cysteine linkage), attaching the molecule to one or more lysine residues, attaching the molecule to one or more unnatural amino acids, or epitope modification. If the chemical modifier is attached via cysteine linkage, the one or more cysteine residues have preferably been introduced into the mutant monomer by substitution. Suitable methods for performing such modifications are well known in the art. Suitable unnatural amino acids include, but are not limited to, 4-azido-L-phenylalanine (Faz) and Liu CC and Schultz P.G., *Annu. Rev. Biochem.*, 2010, Vol. 79, pp. 413-444. Figure 1 Any amino acid numbered 1 to 71 in the amino acid spectrum.
[0365] Mutant monomers can be chemically modified by attaching any molecule that has the effect of reducing or shrinking the diameter of the barrel formed by the monomer at any position or site. Mutant monomers can be chemically modified by attaching the following: (i) maleimides, such as: 4-benzodiazepine, 1,N-(2-hydroxyethyl)maleimide, N-cyclohexylmaleimide, 1,3-maleimide propionic acid, 1,1-4-aminophenyl-1H-pyrrole,2,5,dione, 1,1-4-hydroxyphenyl-1H-pyrrole,2,5,dione, N-ethylmaleimide, N-methoxycarbonylmaleimide, N-tert-butylmaleimide, N-(2-aminoethyl)maleimide, 3-maleimide-PROXYL, N-(4-chlorophenyl)maleimide, 1-[4-(dimethylamino)-3,5-dinitrophenyl]-1H-pyrrole-2,5-dione, N-[4-(2-benzimidazolyl)phenyl]maleimide, N-[4-(2-benzoxazolyl)phenyl]maleimide, N-(1-naphthyl)maleimide N-(2,4-dimethyl)maleimide, N-(2,4-difluorophenyl)maleimide, N-(3-chloro-p-tolyl)maleimide, 1-(2-amino-ethyl)-pyrrole-2,5-dione hydrochloride, 1-cyclopentyl-3-methyl-2,5-dihydro-1H-pyrrole-2,5-dione, 1-(3-aminopropyl)-2,5-dihydro-1H-pyrrole-2,5-dione hydrochloride, 3-methyl-1- [2-Oxo-2-(piperazin-1-yl)ethyl]-2,5-dihydro-1H-pyrrole-2,5-dione hydrochloride, 1-benzyl-2,5-dihydro-1H-pyrrole-2,5-dione, 3-methyl-1-(3,3,3-trifluoropropyl)-2,5-dihydro-1H-pyrrole-2,5-dione, 1-[4-(methylamino)cyclohexyl]-2,5-dihydro-1H-pyrrole-2,5-dione trifluoroacetic acid, SMILES O=C1C=CC(=O)N1CC=2C=CN=CC2、SMILES O=C1C=CC(=O)N1CN2CCNCC2、1-Benzyl-3-methyl-2,5-dihydro-1H-pyrrole-2,5-dione、1-(2-fluorophenyl)-3-methyl-2,5-dihydro-1H-pyrrole-2,5-dione、N-(4-phenoxyphenyl)maleimide、N-(4-nitrophenyl)maleimide;(ii) Iodoacetamides, such as 3-(2-iodoacetamido)-PROXYL, N-(cyclopropylmethyl)-2-iodoacetamide, 2-iodo-N-(2-phenylethyl)acetamide, 2-iodo-N-(2,2,2-trifluoroethyl)acetamide, N-(4-acetylphenyl)-2-iodoacetamide, N-(4-(aminosulfonyl)phenyl)-2-iodoacetamide, N-(1,3-benzothiazolyl-2-yl)-2-iodoacetamide, N-(2,6-diethyl) (iii) Bromoacetamides: such as N-(4-(acetamido)phenyl)-2-bromoacetamide, N-(2-acetylphenyl)-2-bromoacetamide, 2-bromo-N-(2-cyanophenyl)acetamide, 2-bromo-N-(3-(trifluoromethyl)phenyl)acetamide, N-(2-benzoylphenyl)-2-bromoacetamide, 2-bromo-N-(4-fluorophenyl)acetamide (iv) 3-methylbutyramide, N-benzyl-2-bromo-N-phenylpropionamide, N-(2-bromo-butyryl)-4-chloro-benzenesulfonamide, 2-bromo-N-methyl-N-phenylacetamide, 2-bromo-N-phenylethyl-acetamide, 2-adamantane-1-yl-2-bromo-N-cyclohexyl-acetamide, 2-bromo-N-(2-methylphenyl)butyramide, acetyl-p-bromoaniline; (iv) disulfides, such as: ALDRITHIOL-2, ALDRITHIOL-4, iso- Propyl disulfide, 1-(isobutyldithioalkyl)-2-methylpropane, dibenzyl disulfide, 4-aminophenyl disulfide, 3-(2-pyridyldithio)propionic acid, 3-(2-pyridyldithio)propionic acid hydrazide, 3-(2-pyridyldithio)propionic acid N-succinimide ester, am6amPDP1-βCD; and (v) thiols, such as: 4-phenylthiazolyl-2-thiol, Pulpald, 5,6,7,8-tetrahydro-quinazoline-2-thiol.
[0366] Mutant monomers can be chemically modified by attaching polyethylene glycol (PEG), nucleic acids such as DNA, dyes, fluorophores, or chromophores. In some embodiments, the mutant monomers are chemically modified using molecular adaptors that promote interactions between the pore containing the monomer and the target analyte, target nucleotide, or target polynucleotide sequence. The presence of the adaptor improves the host-guest chemistry of the pore and the nucleotide or polynucleotide, thereby enhancing the sequencing capability of the pores formed from the mutant monomers.
[0367] The chemically modified mutant monomer preferably comprises a variant of the sequence shown in SEQ ID NO: 2. The variant is defined as follows. The variant typically includes one or more substitutions, wherein one or more residues are replaced by cysteine, lysine, or a non-natural amino acid.Non-natural amino acids include, but are not limited to: 4-azido-L-phenylalanine (Faz), 4-acetyl-L-phenylalanine, 3-acetyl-L-phenylalanine, 4-acetoacetyl-L-phenylalanine, O-allyl-L-tyrosine, 3-(phenylvinyl)-L-alanine, O-2-propynyl-1-yl-L-tyrosine, 4-(dihydroxyboryl)-L-phenylalanine, 4-[(ethylthioalkyl)carbonyl]-L-phenylalanine, (2S)-2-amino-3-4-[(prop-2-ylthioalkyl)carbonyl]phenyl;propionic acid, (2S)-2-amino-3-4-[(2-amino-3-thioalkylpropionyl)amino]phenyl;propionic acid, O-methyl-L-tyrosine, 4-amino-L-phenylalanine, etc. Amino acids, 4-cyano-L-phenylalanine, 3-cyano-L-phenylalanine, 4-fluoro-L-phenylalanine, 4-iodo-L-phenylalanine, 4-bromo-L-phenylalanine, O-(trifluoromethyl)tyrosine, 4-nitro-L-phenylalanine, 3-hydroxy-L-tyrosine, 3-amino-L-tyrosine, 3-iodo-L-tyrosine, 4-isopropyl-L-phenylalanine, 3-(2-naphthyl)-L-alanine, 4-phenyl-L-phenylalanine, (2S)-2-amino-3-(naphthyl-2-ylamino)propionic acid, 6-(methylthioalkyl)-leucine, 6-oxo-L-lysine, D-tyrosine, (2R)-2-hydroxy-3-(4-hydroxyphenyl)propionic acid, (2R)-2-aminooctanoate 3-( 2,2'-Bipyridin-5-yl)-D-alanine, 2-amino-3-(8-hydroxy-3-quinolinyl)propionic acid, 4-benzoyl-L-phenylalanine, S-(2-nitrobenzyl)cysteine, (2R)-2-amino-3-[(2-nitrobenzyl)thioalkyl]propionic acid, (2S)-2-amino-3-[(2-nitrobenzyl)oxy]propionic acid, O-(4,5-dimethoxy-2-nitrobenzyl)-L-serine, (2S)-2-amino-6-([(2-nitrobenzyl)oxy]carbonyl;amino)hexanoic acid, O-(2-nitrobenzyl)-L-tyrosine, 2-nitrophenylalanine, 4-[(E)-phenyldiazeninyl]-L-phenylalanine, 4-[3-(trifluoromethyl)-3H -diazacycloheptane-3-yl]-D-phenylalanine, 2-amino-3-[[5-(dimethylamino)-1-naphthyl]sulfonylamino]propionic acid, (2S)-2-amino-4-(7-hydroxy-2-oxo-2H-benzopyran-4-yl)butyric acid, (2S)-3-[(6-acetylnaphthylacetamide-2-yl)amino]-2-aminopropionic acid, 4-(carboxymethyl)phenylalanine, 3-nitro-L-tyrosine, O-sulfon-L-tyrosine, (2R)-6-acetamido-2-aminohexanoic acid, 1-methylhistidine, 2-aminononanoic acid, 2-aminodecanoic acid, -L-homocysteine, 5-thioalkyl-n-valine, 6-thioalkyl-L-n-leucine, 5-(methylthio)-L-n-valine, N.6 -[(2R,3R)-3-methyl-3,4-dihydro-2H-pyrrolo-2-yl]carbonyl; -L-lysine, N 6 -[(benzyloxy)carbonyl]lysine, (2S)-2-amino-6-[(cyclopentylcarbonyl)amino]hexanoic acid, N 6 -[(cyclopentyloxy)carbonyl]-L-lysine, (2S)-2-amino-6-[(2R)-tetrahydrofuran-2-ylcarbonyl]amino; hexanoic acid, (2S)-2-amino-8-[(2R,3S)-3-ethynyltetrahydrofuran-2-yl]-8-oxooctanoic acid, N 6 -(tert-Butoxycarbonyl)-L-lysine, (2S)-2-hydroxy-6-([(2-methyl-2-propyl)oxy]carbonyl;amino)hexanoic acid, N 6 -[(allyloxy)carbonyl]lysine, (2S)-2-amino-6-([(2-azidobenzyl)oxy]carbonyl;amino)hexanoic acid, N 6 -L-prolyl-L-lysine, (2S)-2-amino-6-[(prop-2-yn-1-yloxy)carbonyl]amino; hexanoic acid and N 6 -[(2-azidoethoxy)carbonyl]-L-lysine. The preferred non-natural amino acid is 4-azido-L-phenylalanine (Faz).
[0368] The mutant monomer can be chemically modified by attaching any molecule to any of the following positions in SEQ ID NO: 2: K37, V47, S49, T55, S86, E92, and E94. More preferably, the mutant monomer can be chemically modified by attaching any molecule to position E92 and / or E94. In one embodiment, the mutant monomer is chemically modified by attaching molecules to one or more cysteine residues (cysteine linkages), one or more lysine residues, or one or more non-natural amino acids at these positions. The mutant monomer preferably comprises a variant of the sequence shown in SEQ ID NO: 2, which includes one or more of K37C, V47C, S49C, T55C, S86C, E92C, and E94C, wherein one or more molecules are attached to the one or more introduced cysteine residues. The mutant monomer more preferably comprises a variant of the sequence shown in SEQ ID NO: 2, which includes E92C and / or E94C, wherein one or more molecules are attached to the one or more introduced cysteine residues. In each of these two preferred embodiments, the one or more cysteines (Cs) may be replaced by one or more lysines or one or more non-natural amino acids such as one or more Faz.
[0369] The reactivity of cysteine residues can be enhanced by modifying adjacent residues. For example, attaching a basic group to an arginine, histidine, or lysine residue can alter the pKa of the cysteine thiol group to make it more reactive. - The pKa of the cysteine residue can be protected by thiol protecting groups such as dTNB. These can react with one or more cysteine residues of the mutant monomer prior to attachment linker.
[0370] The molecule can be directly attached to the mutant monomer. Preferably, the molecule is attached to the mutant monomer using a connector such as a chemical crosslinking agent or a peptide linker. Suitable chemical crosslinking agents are well known in the art. Preferred crosslinking agents include 2,5-dioxopyrrolidine-1-yl ester of 3-(pyridin-2-yldisulfonyl)propionate, 2,5-dioxopyrrolidine-1-yl ester of 4-(pyridin-2-yldisulfonyl)butyrate, and 2,5-dioxopyrrolidine-1-yl ester of 8-(pyridin-2-yldisulfonyl)octanoate. The most preferred crosslinking agent is succinimide ester of 3-(2-pyridinedithio)propionate (SPDP). Typically, the molecule is covalently linked to the bifunctional crosslinking agent before the molecule / crosslinking agent complex is covalently linked to the mutant monomer, but it is also possible to covalently link the bifunctional crosslinking agent to the monomer before the bifunctional crosslinking agent / monomer complex is attached to the molecule.
[0371] Preferably, the joint is resistant to dithiothreitol (DTT). Suitable joints include, but are not limited to, iodoacetamide-based and maleimide-based joints.
[0372] The advantages of the pores in the chemically modified mutant monomers of the present invention are discussed in more detail below.
[0373] Further chemical modifications that can be performed according to the present invention are discussed below.
[0374] Further refinement
[0375] Where appropriate, any of the mutant monomers discussed above may include further modifications in the region approximately 44 to approximately 126 of SEQ ID NO: 2 (i.e., where the relevant amino position remains in the mutant monomer or is not modified / substituted by another amino acid). At least a portion of this region typically contributes to the transmembrane region of the cytolysin. At least a portion of this region typically contributes to the barrel or channel of the cytolysin. At least a portion of this region typically contributes to the inner wall or liner of the cytolysin.
[0376] The transmembrane region of cytosin has been identified as positions 44 to 67 of SEQ ID NO: 2 (De Colbis et al., Structure, 2012, Vol. 20, pp. 1498 to 1507).
[0377] The variant preferably comprises one or more modifications within the region approximately 44 to approximately 126 of SEQ ID NO: 2, said modifications altering the ability of the monomer or preferably said region to interact with the polynucleotide. The interaction between the monomer and the polynucleotide can be increased or decreased. Increased interaction between the monomer and the polynucleotide will, for example, facilitate the capture of the polynucleotide through a pore containing the mutant monomer. Decreased interaction between said region and the polynucleotide will, for example, improve the recognition or differentiation of the polynucleotide. The recognition or differentiation of the polynucleotide can be improved by reducing changes in the state of the pore containing the mutant monomer (which increases the signal-to-noise ratio) and / or by reducing the number of nucleotides in the polynucleotide that contribute to the current as the polynucleotide moves through the pore containing the mutant monomer.
[0378] The ability of a monomer to interact with a polynucleotide can be determined using methods well known in the art. The monomer can interact with the polynucleotide in any manner, such as through non-covalent interactions like hydrophobic interactions, hydrogen binding, van der Waals forces, π(π)-cation interactions, or electrostatic forces. For example, the ability of the region to bind to a polynucleotide can be measured using conventional binding assays. Suitable assays include, but are not limited to, fluorescence-based binding assays, nuclear magnetic resonance (NMR), isothermal titration calorimetry (ITC), or electron spin resonance (ESR) spectroscopy. Alternatively, the ability of a pore containing one or more of the mutant monomers to interact with a polynucleotide can be determined using any of the methods discussed above or below. Preferred assays are described in the examples.
[0379] One or more modifications may be further performed in the region approximately 44 to approximately 126 of SEQ ID NO: 2. Preferably, the one or more modifications are performed in any of the following regions: approximately 40 to approximately 125, approximately 50 to approximately 120, approximately 60 to approximately 110, and approximately 70 to approximately 100. If the one or more modifications are performed to improve polynucleotide capture, it is more preferably performed in any of the following regions: approximately 44 to approximately 103, approximately 68 to approximately 103, approximately 84 to approximately 103, approximately 44 to approximately 97, approximately 68 to approximately 97, or approximately 84 to approximately 97. If the one or more modifications are performed to improve polynucleotide recognition or differentiation, it is more preferably performed in any of the following regions: approximately 44 to approximately 109, approximately 44 to approximately 97, or approximately 48 to approximately 88. Preferably, the region is approximately positions 44 to 67 of SEQ ID NO: 2.
[0380] If the one or more modifications are intended to improve polynucleotide recognition or differentiation, then preferably, these modifications are performed in addition to one or more modifications for improving polynucleotide capture. This allows the pores formed by the mutant monomers to efficiently capture the polynucleotides, and then characterize the polynucleotides, such as estimating their sequences, as discussed below.
[0381] Protein nanopore modifications that alter the ability of protein nanopores to interact with polynucleotides, particularly enhancing their ability to capture and / or recognize or distinguish polynucleotides, are well documented in the art. Such modifications are disclosed, for example, in WO 2010 / 034018 and WO 2010 / 055307. Similar modifications can be made to the cytolysin monomers according to the invention.
[0382] Any number of modifications can be made, such as 1, 2, 5, 10, 15, 20, 30 or more modifications. Any one or more modifications can be made as long as the ability of the monomer to interact with the polynucleotide is altered. Suitable modifications include, but are not limited to, amino acid substitution, amino acid addition, and amino acid deletion. Preferably, the one or more modifications are one or more substitutions. This is discussed in more detail below.
[0383] The one or more modifications preferably (a) alter the steric hindrance effect of the monomer or preferably alter the steric hindrance effect of the region; (b) alter the net charge of the monomer or preferably alter the net charge of the region; (c) alter the ability of the monomer or preferably the region to bind to the hydrogen of the polynucleotide; (d) introduce or remove chemical groups that interact through a delocalized electron π system and / or (e) alter the structure of the monomer or preferably alter the structure of the region. The one or more modifications more preferably produce any combination of (a) to (e), such as (a) and (b); (a) and (c); (a) and (d); (a) and (e); (b) and (c); (b) and (d); (b) and (e); (c) and (d); (c) and (e); (d) and (e), (a), (b) and (c); (a), (b) and (d); (a), (b) and (e); (a), (c) and (d); (a), (c) and (d); (a), (c) and (d) (a), (d) and (e); (b), (c) and (d); (b), (c) and (e); (b), (d) and (e); (c), (d) and (e); (a), (b), (c) and (d); (a), (b), (c) and (e); (a), (b), (d) and (e); (a), (c), (d) and (e); (b), (c), (d) and (e); and (a), (b), (c) and (d).
[0384] For (a), the steric hindrance effect of the monomer can be increased or decreased. Any method of altering the steric hindrance effect can be used according to the invention. The steric hindrance of the monomer is increased by introducing bulky residues such as phenylalanine (F), tryptophan (W), tyrosine (Y), or histidine (H). The one or more modifications are preferably the introduction of one or more of F, W, Y, and H. Any combination of F, W, Y, and H can be introduced. The one or more of F, W, Y, and H can be introduced by addition. The one or more of F, W, Y, and H are preferably introduced by substitution. The appropriate positions for introducing such residues are discussed in more detail below.
[0385] Removing bulky residues such as phenylalanine (F), tryptophan (W), tyrosine (Y), or histidine (H) conversely reduces the steric hindrance of the monomer. The one or more modifications preferably involve the removal of one or more of F, W, Y, and H. Any combination of F, W, Y, and H can be removed. One or more of F, W, Y, and H can be removed by deletion. The one or more of F, W, Y, and H are preferably removed by substitution with residues having smaller side groups, such as serine (S), threonine (T), alanine (A), and valine (V).
[0386] For (b), the net charge can be changed in any way. Preferably, the net positive charge is increased or decreased. The net positive charge can be increased in any way. Preferably, the net positive charge is increased by introducing, preferably by substituting one or more positively charged amino acids and / or neutralizing, and most preferably by substituting one or more negatively charged amino acids.
[0387] Preferably, the net positive charge is increased by introducing one or more positively charged amino acids. This can be achieved by adding the one or more positively charged amino acids. Preferably, it is achieved by substitution of the one or more positively charged amino acids. The positively charged amino acids are amino acids having a net positive charge. The one or more positively charged amino acids can be naturally occurring or non-natural. The positively charged amino acids can be synthetic or modified. For example, modified amino acids having a net positive charge can be specifically designed for use in this invention. Various types of modifications to amino acids are well known in the art.
[0388] Preferred naturally occurring positively charged amino acids include, but are not limited to, histidine (H), lysine (K), and arginine (R). The one or more modifications are preferably the introduction of one or more of H, K, and R. Any number of H, K, and R and any combination thereof can be introduced. The one or more of H, K, and R can be introduced by addition. The one or more of H, K, and R are preferably introduced by substitution. The appropriate positions for introducing such residues are discussed in more detail below.
[0389] Methods for adding or substituting naturally occurring amino acids are well known in the field. For example, arginine (R) can be used to replace methionine (M) at the relevant position in the polynucleotide encoding the mutant monomer by replacing the codon for methionine (ATG) with the codon for arginine (CGT). The polynucleotide can then be expressed as discussed below.
[0390] Methods for adding or substituting non-naturally occurring amino acids are well known in the field. For example, non-naturally occurring amino acids can be introduced by including synthetic aminoacyl-tRNA in an IVTT system used for expression pores. Alternatively, non-naturally occurring amino acids can be introduced by expressing mutant monomers in *E. coli* that are auxotrophic for those specific amino acids in the presence of synthetic (i.e., non-naturally occurring) analogs of those amino acids. If the pores are generated using partial peptide synthesis, non-naturally occurring amino acids can also be generated via naked linking.
[0391] Any amino acid can be replaced by a positively charged amino acid. One or more uncharged amino acids, nonpolar amino acids, and / or aromatic amino acids can be replaced by one or more positively charged amino acids. Uncharged amino acids have no net charge. Suitable uncharged amino acids include, but are not limited to, cysteine (C), serine (S), threonine (T), methionine (M), asparagine (N), and glutamine (Q). Nonpolar amino acids have nonpolar side chains. Suitable nonpolar amino acids include, but are not limited to, glycine (G), alanine (A), proline (P), isoleucine (I), leucine (L), and valine (V). Aromatic amino acids have aromatic side chains. Suitable aromatic amino acids include, but are not limited to, histidine (H), phenylalanine (F), tryptophan (W), and tyrosine (Y). Preferably, one or more negatively charged amino acids are replaced by one or more positively charged amino acids. Suitable negatively charged amino acids include, but are not limited to, aspartic acid (D) and glutamic acid (E).
[0392] Preferred substitutions include, but are not limited to: replacing E with K, replacing M with R, replacing M with H, replacing M with K, replacing D with R, replacing D with H, replacing D with K, replacing E with R, replacing E with H, replacing N with R, replacing T with R, and replacing G with R. Most preferably, E is replaced with K.
[0393] Any number of positively charged amino acids can be introduced or substituted. For example, one, two, five, ten, fifteen, twenty, twenty, thirty or more positively charged amino acids can be introduced or substituted.
[0394] More preferably, the net positive charge is increased by neutralizing one or more negative charges. The one or more negatively charged amino acids can be neutralized by replacing one or more negatively charged amino acids with one or more uncharged amino acids, nonpolar amino acids, and / or aromatic amino acids. Removing the negative charge increases the net positive charge. The uncharged amino acids, nonpolar amino acids, and / or aromatic amino acids can be naturally occurring or non-natural. They can be synthetic or modified. Suitable uncharged amino acids, nonpolar amino acids, and aromatic amino acids have been discussed above. Preferred substitutions include, but are not limited to: substituting E with Q, substituting E with S, substituting E with A, substituting D with Q, substituting E with N, substituting D with N, substituting D with G, and substituting D with S.
[0395] It can replace any number of uncharged amino acids, nonpolar amino acids, and / or aromatic amino acids and any combination thereof. For example, it can replace 1, 2, 5, 10, 15, 20, 25, or 30 or more uncharged amino acids, nonpolar amino acids, and / or aromatic amino acids. Negatively charged amino acids can be replaced by the following amino acids: (1) uncharged amino acids; (2) nonpolar amino acids; (3) aromatic amino acids; (4) uncharged amino acids and nonpolar amino acids; (5) uncharged amino acids and aromatic amino acids; and (6) nonpolar amino acids and aromatic amino acids; or (7) uncharged amino acids, nonpolar amino acids, and aromatic amino acids.
[0396] The one or more negative charges can be neutralized by introducing one or more positively charged amino acids in the vicinity of one or more negatively charged amino acids, such as within one, two, three, or four amino acids, or adjacent to one or more negatively charged amino acids. Examples of positively and negatively charged amino acids have been discussed above. Positively charged amino acids can be introduced in any of the ways discussed above, such as by substitution.
[0397] Preferably, the net positive charge is reduced by introducing one or more negatively charged amino acids and / or neutralizing one or more positive charges. The manner in which this can be achieved will become clear from the above discussion on increasing net positive charge. All embodiments discussed above regarding increasing net positive charge are equally applicable to reducing net positive charge, except that the charge is changed in the opposite manner. Specifically, the one or more positively charged amino acids are preferably neutralized by replacing one or more positively charged amino acids with one or more non-charged amino acids, nonpolar amino acids, and / or aromatic amino acids, and / or by introducing one or more negatively charged amino acids in the vicinity of one or more positively charged amino acids, such as within one, two, three, or four of them, or adjacent to one or more positively charged amino acids.
[0398] Preferably, the net negative charge is increased or decreased. All the above-described embodiments discussed with reference to increasing or decreasing the net positive charge are equally applicable to decreasing or increasing the net negative charge.
[0399] Regarding (c), the ability of the monomer to bind hydrogen can be altered in any way. The introduction of serine (S), threonine (T), asparagine (N), glutamine (Q), tyrosine (Y), or histidine (H) increases the hydrogen-binding ability of the monomer. The one or more modifications are preferably the introduction of one or more of S, T, N, Q, Y, and H. Any combination of S, T, N, Q, Y, and H can be introduced. The introduction of one or more of S, T, N, Q, Y, and H can be achieved by addition. Preferably, the introduction of one or more of S, T, N, Q, Y, and H is achieved by substitution. The appropriate positions for introducing such residues are discussed in more detail below.
[0400] Removal of serine (S), threonine (T), asparagine (N), glutamine (Q), tyrosine (Y), or histidine (H) reduces the hydrogen-binding capacity of the monomer. The one or more modifications preferably involve the removal of one or more of S, T, N, Q, Y, and H. Any combination of S, T, N, Q, Y, and H can be removed. The one or more of S, T, N, Q, Y, and H can be removed by deletion. Preferably, the one or more of S, T, N, Q, Y, and H are removed by substitution with other amino acids that do not bind hydrogen well, such as alanine (A), valine (V), isoleucine (I), and leucine (L).
[0401] For (d), the introduction of aromatic residues such as phenylalanine (F), tryptophan (W), tyrosine (Y), or histidine (H) increases π-stacking in the monomer. The removal of aromatic residues such as phenylalanine (F), tryptophan (W), tyrosine (Y), or histidine (H) reduces π-stacking in the monomer. Such amino acids can be introduced or removed as discussed above with reference to (a).
[0402] Regarding (e), one or more modifications to alter the structure of the monomer can be made according to the invention. For example, one or more loop regions can be removed, shortened, or amplified. This typically facilitates the entry or exit of polynucleotides from the pore. The one or more loop regions can be on the cis-side, anti-side, or both sides of the pore. Alternatively, one or more regions of the amino-terminal and / or carboxyl-terminal of the pore can be amplified or deleted. This typically alters the pore size and / or charge.
[0403] Based on the above discussion, it will be clear that the introduction of certain amino acids will enhance the monomer's ability to interact with polynucleotides through more than one mechanism. For example, replacing E with H will not only increase the net positive charge according to (b) (by neutralizing the negative charge), but will also enhance the monomer's ability to bind hydrogen according to (c) hydrogen bonding.
[0404] The variant preferably includes substitutions at one or more of the following positions in SEQ ID NO: 2: M44, N46, N48, E50, R52, H58, D68, F70, E71, S74, E76, S78, Y79, S80, H81, S82, E84, E85, S86, Q87, S89, M90, E92, E94, E97, E102, H103, T104, T106, R115, Q117, N119, D121, and D126. The variants preferably include substitutions at positions 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, or 34. The variants preferably include substitutions at one or more of the following positions of SEQ ID NO: 2: D68, E71, S74, E76, S78, S80, S82, E84, E85, S86, Q87, S89, E92, E102, T104, T106, R115, Q117, N119, and D121. The variants preferably include substitutions at positions 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19 or 20.
[0405] The variant preferably includes substitutions at one or more of the following positions in SEQ ID NO: 2: (a) E84, E85, E92, E97, and D126; (b) E85, E97, and D126; or (c) E84 and E92. The amino acid substituted into the variant can be a naturally occurring derivative or a non-naturally occurring derivative. The amino acid substituted into the variant can be a D-amino acid. Each position listed above can be substituted with one of the following: asparagine (N), serine (S), glutamine (Q), arginine (R), glycine (G), tyrosine (Y), aspartic acid (D), leucine (L), lysine (K), or alanine (A).
[0406] The variant preferably includes at least one of the following mutations from SEQ ID NO: 2:
[0407] (a) Serine (S) at position 44;
[0408] (b) Serine (S) at position 46;
[0409] (c) Serine (S) at position 48;
[0410] (d) Serine (S) at position 52;
[0411] (e) Serine (S) at position 58;
[0412] (f) Serine (S) at position 68;
[0413] (g) Serine (S) at position 70;
[0414] (h) Serine (S) at position 71;
[0415] (i) Serine (S) at position 76;
[0416] (j) Serine (S) at position 79;
[0417] (k) Serine (S) at position 81;
[0418] (l) Serine (S), aspartic acid (D), or glutamine (Q) at position 84;
[0419] (m) serine (S) or lysine (K) at position 85;
[0420] (n) Serine (S) at position 87;
[0421] (o) Serine (S) at position 90;
[0422] (p) Asparagine (N) or glutamine (Q) at position 92;
[0423] (q) Serine (S) or asparagine (N) at position 94;
[0424] (r) serine (S) or asparagine (N) at position 97;
[0425] (s) serine at position 102;
[0426] (t) Serine (S) at position 103;
[0427] (u) Asparagine (N) or serine (S) at position 121;
[0428] (v) Serine (S) at position 50;
[0429] (w) Asparagine (N) or serine (S) at position 94;
[0430] (x) Asparagine (N) or serine (S) at position 97;
[0431] (y) serine (S) or asparagine (N) at position 121;
[0432] (z) Asparagine (N) or glutamine (Q) at position 126; and
[0433] (aa) Serine (S) or asparagine (N) at position 128.
[0434] The variant may contain any number of mutations (a) to (aa), such as 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, or 27 of the mutations. Preferred combinations of mutations are discussed below. The amino acid introduced into the variant may be a naturally occurring derivative or a non-naturally occurring derivative. The amino acid introduced into the variant may be a D-amino acid.
[0435] The variant preferably includes at least one of the following mutations from SEQ ID NO: 2:
[0436] (a) Serine (S) at position 68;
[0437] (b) Serine (S) at position 71;
[0438] (c) Serine (S) at position 76;
[0439] (d) Aspartic acid (D) or glutamine (Q) at position 84;
[0440] (e) Lysine (K) at position 85;
[0441] (f) Asparagine (N) or glutamine (Q) at position 92;
[0442] (g) Serine (S) at position 102;
[0443] (h) Asparagine (N) or serine (S) at position 121;
[0444] (i) Serine (S) at position 50;
[0445] (j) Asparagine (N) or serine (S) at position 94;
[0446] (k) Asparagine (N) or serine (S) at position 97; and
[0447] (l) Asparagine (N) or glutamine (Q) at position 126.
[0448] The variant may contain any number of mutations (a) to (l), such as 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, or 12 of the mutations. Preferred combinations of mutations are discussed below. The amino acid introduced into the variant may be a naturally occurring derivative or a non-naturally occurring derivative. The amino acid introduced into the variant may be a D-amino acid.
[0449] The variant may contain one or more additional modifications outside the region approximately 44 to approximately 126 of SEQ ID NO: 2, which, in combination with the modifications in the regions discussed above, enhance polynucleotide capture and / or enhance polynucleotide recognition or differentiation. Suitable modifications include, but are not limited to, substitutions at one or more of D35, E128, E135, E134, and E167. Specifically, removing the negative charge by substituting E at one or more of positions 128, 135, 134, and 167 enhances polynucleotide capture. E at one or more of these positions may be substituted in any of the manner discussed above. Preferably, all of E128, E135, E134, and E167 are substituted as discussed above. Preferably, A is used to substitute for E. In other words, the variant preferably includes one or more or all of E128A, E135A, E134A, and E167A. Another preferred substitution is D35Q.
[0450] In a preferred embodiment, the variant includes the following substitutions from SEQ ID NO: 2:
[0451] i. One or more of E84D and E85K, such as both;
[0452] ii. One or more of E84Q, E85K, E92Q, E97S, D126G and E167A, such as 2, 3, 4, 5 or 6;
[0453] iii. One or more of E92N, E94N, E97N, D121N, and D126N, such as two, three, four, or five;
[0454] iv. One or more of E92N, E94N, E97N, D121N, D126N and E128N, such as 2, 3, 4, 5 or 6;
[0455] v. one or more of E76S, E84Q, E85K, E92Q, E97S, D126G and E167A, such as 2, 3, 4, 5, 6 or 7;
[0456] vi. One or more of E84Q, E85K, E92Q, E97S, D126G, E167A and E50S, such as 2, 3, 4, 5, 6 or 7;
[0457] vii. One or more of E84Q, E85K, E92Q, E97S, D126G, E167A and E71S, such as 2, 3, 4, 5, 6 or 7;
[0458] viii. One or more of E84Q, E85K, E92Q, E97S, D126G, E167A and E94S, such as 2, 3, 4, 5, 6 or 7;
[0459] ix. One or more of E84Q, E85K, E92Q, E97S, D126G, E167A and E102S, such as 2, 3, 4, 5, 6 or 7;
[0460] x. One or more of E84Q, E85K, E92Q, E97S, D126G, E167A and E128S, such as 2, 3, 4, 5, 6 or 7;
[0461] xi. One or more of E84Q, E85K, E92Q, E97S, D126G, E167A and E135S, such as 2, 3, 4, 5, 6 or 7;
[0462] xii. One or more of E84Q, E85K, E92Q, E97S, D126G, E167A and D68S, such as 2, 3, 4, 5, 6 or 7;
[0463] xiii. One or more of E84Q, E85K, E92Q, E97S, D126G, E167A and D121S, such as 2, 3, 4, 5, 6 or 7;
[0464] xiv. One or more of E84Q, E85K, E92Q, E97S, D126G, E167A and D134S, such as 2, 3, 4, 5, 6 or 7;
[0465] One or more of xv.E84D, E85K and E92Q, such as two or three;
[0466] xvi. One or more of E84Q, E85K, E92Q, E97S, D126G and E135S, such as 1, 2, 3, 4, 5 or 6;
[0467] xvii. One or more of E85K, E92Q, E94S, E97S and D126G, such as 1, 2, 3, 4 or 5;
[0468] xviii. One or more of E76S, E85K, E92Q, E97S and D126G, such as 1, 2, 3, 4 or 5;
[0469] One or more of xix.E71S, E85K, E92Q, E97S and D126G, such as 1, 2, 3, 4 or 5;
[0470] One or more of xx.D68S, E85K, E92Q, E97S and D126G, such as 1, 2, 3, 4 or 5;
[0471] One or more of xxi.E85K, E92Q, E97S and D126G, such as 1, 2, 3 or 4;
[0472] xxii. One or more of E84Q, E85K, E92Q, E97S, H103S and D126G, such as 1, 2, 3, 4, 5 or 6;
[0473] xxiii. One or more of E84Q, E85K, M90S, E92Q, E97S and D126G, such as 1, 2, 3, 4, 5 or 6;
[0474] xxiv. one or more of E84Q, Q87S, E85K, E92Q, E97S and D126G, such as 1, 2, 3, 4, 5 or 6;
[0475] One or more of xxv.E84Q, E85S, E92Q, E97S and D126G, such as 1, 2, 3, 4 or 5;
[0476] One or more of xxvi.E84S, E85K, E92Q, E97S and D126G, such as 1, 2, 3, 4 or 5;
[0477] One or more of xxvii.H81S, E84Q, E85K, E92Q, E97S and D126G, such as 1, 2, 3, 4, 5 or 6;
[0478] xxviii.Y79S, E84Q, E85K, E92Q, E97S and D126G, such as 1, 2, 3, 4, 5 or 6;
[0479] One or more of xxix.F70S, E84Q, E85K, E92Q, E97S and D126G, such as 1, 2, 3, 4, 5 or 6;
[0480] One or more of xxx.H58S, E84Q, E85K, E92Q, E97S and D126G, such as 1, 2, 3, 4, 5 or 6;
[0481] One or more of xxxi.R52S, E84Q, E85K, E92Q, E97S and D126G, such as 1, 2, 3, 4, 5 or 6;
[0482] xxxii.N48S, E84Q, E85K, E92Q, E97S and D126G, such as 1, 2, 3, 4, 5 or 6;
[0483] xxxiii. One or more of N46S, E84Q, E85K, E92Q, E97S and D126G, such as 1, 2, 3, 4, 5 or 6;
[0484] xxxiv.M44S, E84Q, E85K, E92Q, E97S and D126G, such as 1, 2, 3, 4, 5 or 6;
[0485] One or more of xxxv.E92Q and E97S, such as both;
[0486] xxxvi. One or more of E84Q, E85K, E92Q and E97S, such as 1, 2, 3 or 4;
[0487] One or more of xxxvii.E84Q and E85K, such as both;
[0488] xxxviii. One or more of E84Q, E85K and D126G, such as 1, 2 or 3;
[0489] One or more of xxxix.E84Q, E85K, D126G and E167A, such as 1, 2, 3 or 4;
[0490] One or more of xl.E92Q, E97S and D126G, such as 1, 2 or 3;
[0491] One or more of xli.E84Q, E85K, E92Q, E97S and D126G, such as 1, 2, 3, 4 or 5;
[0492] xlii. One or more of E84Q, E85K, E92Q, E97S and E167A, such as 1, 2, 3, 4 or 5;
[0493] One or more of xliii.E84Q, E85K, E92Q, D126G and E167A, such as 1, 2, 3, 4 or 5;
[0494] One or more of xliv.E84Q, E85K, E97S, D126G and E167A, such as 1, 2, 3, 4 or 5;
[0495] One or more of xlv.E84Q, E92Q, E97S, D126G and E167A, such as 1, 2, 3, 4 or 5;
[0496] xlvi. One or more of E85K, E92Q, E97S, D126G and E167A, such as 1, 2, 3, 4 or 5;
[0497] One or more of xlvii.E84D, E85K and E92Q, such as 1, 2 or 3;
[0498] xlviii. One or more of E84Q, E85K, E92Q, E97S, D126G, E167A and D121S, such as 1, 2, 3, 4, 5, 6 or 7;
[0499] One or more of xlix.E84Q, E85K, E92Q, E97S, D126G, E167A and D68S, such as 1, 2, 3, 4, 5, 6 or 7;
[0500] l. One or more of E84Q, E85K, E92Q, E97S, D126G, E167A and E135S, such as 1, 2, 3, 4, 5, 6 or 7;
[0501] One or more of E84Q, E85K, E92Q, E97S, D126G, E167A and E128S, such as 1, 2, 3, 4, 5, 6 or 7;
[0502] One or more of E84Q, E85K, E92Q, E97S, D126G, E167A and E102S, such as 1, 2, 3, 4, 5, 6 or 7;
[0503] One or more of liii.E84Q, E85K, E92Q, E97S, D126G, E167A and E94S, such as 1, 2, 3, 4, 5, 6 or 7;
[0504] One or more of the following: liv.E84Q, E85K, E92Q, E97S, D126G, E167A and E71S, such as 1, 2, 3, 4, 5, 6 or 7;
[0505] One or more of the following: lv.E84Q, E85K, E92Q, E97S, D126G, E167A and E50S, such as 1, 2, 3, 4, 5, 6 or 7;
[0506] lvi. one or more of E76S, E84Q, E85K, E92Q, E97S, D126G and E167A, such as 1, 2, 3, 4, 5, 6 or 7;
[0507] lvii. one or more of E92N, E94N, E97N, D121N, D126N and E128N, such as 1, 2, 3, 4, 5 or 6;
[0508] lviii. One or more of E92N, E94N, E97N, D121N, and D126N, such as 1, 2, 3, 4, or 5; or
[0509] One or more of the following: E84Q, E85K, E92Q, E97S, D126G, and E167A, such as 1, 2, 3, 4, 5, or 6.
[0510] In the above text, the first letter refers to the amino acid that is substituted in SEQ ID NO:2, the number is the position in SEQ ID NO:2, and the second letter refers to the amino acid that will be used to replace the first one. Therefore, E84D means that aspartic acid (D) is used to replace glutamic acid (E) at position 84.
[0511] The variant may contain any number of substitutions from i to lix, such as 1, 2, 3, 4, 5, 6, or 7. The variant preferably contains all the substitutions shown in any of i to lix above.
[0512] In a preferred embodiment, the variant includes substitutions from any of i to xv described above. The variant may contain any number of substitutions from any of i to xv, such as 1, 2, 3, 4, 5, 6, or 7. Preferably, the variant includes all substitutions shown in any of i to xv described above.
[0513] If the one or more modifications are intended to improve the ability of monomers to recognize or distinguish polynucleotides, then preferably, in addition to the modifications discussed above that improve polynucleotide capture, such as E84Q, E85K, E92Q, E97S, D126G and E167A, the one or more modifications may also be performed.
[0514] The one or more modifications made to the identified region may involve replacing one or more amino acids in the region with amino acids at one or more corresponding positions present in homologs or parahomologs of cytosin. Four examples of homologs of cytosin are shown in SEQ ID NO: 14 to 17. The advantage of such substitutions is that they may result in mutant monomers that form pores, since homolog monomers also form pores. For example, mutations may be made at any one or more positions in SEQ ID NO: 2 that are different between SEQ ID NO: 2 and any one of SEQ ID NO: 14 to 17. Such mutations may be made by replacing the amino acid in SEQ ID NO: 2 with an amino acid at a corresponding position from SEQ ID NO: 14 to 17, preferably SEQ ID NO: 14 to 16. Alternatively, mutations at any of these positions may be substitutions made with any amino acid, or may be deletion or insertion mutations, such as substitutions, deletions, or insertions of 1 to 30 amino acids, such as 2 to 20, 3 to 10, or 4 to 8 amino acids. In addition to the mutations disclosed herein and those in the prior art, such as those disclosed in 2013 / 153359, the amino acids that are conserved or consistent between SEQ ID NO: 2 and all SEQ ID NO: 14 to 17, more preferably all SEQ ID NO: 14 to 16, are preferably conserved or present in variants of the invention. However, conserved mutations may be made at any one or more of these positions where SEQ ID NO: 2 is conserved or consistent between all SEQ ID NO: 14 to 17, or more preferably SEQ ID NO: 14 to 16.
[0515] This invention provides a cytolysin mutant monomer comprising any one or more amino acids described herein as having been substituted at a position in the structure of the cytolysin monomer corresponding to a specific position in SEQ ID NO: 2 at said specific position. The corresponding position can be determined using standard techniques in the art. For example, the PILEUP and BLAST algorithms mentioned above can be used to align the sequence of the cytolysin monomer with SEQ ID NO: 2 and thus identify the corresponding residues.
[0516] Mutant monomers typically retain the ability to form the same 3D structure as wild-type cytolysin monomers, such as the 3D structure of cytolysin monomers having the sequence SEQ ID NO: 2. The 3D structures of cytolysin monomers are known in the art and disclosed, for example, in De Colbis et al., *Structure*, 2012, Vol. 20, pp. 1498–1507. Mutant monomers typically retain the ability to form homo- and / or hetero-oligomeric pores with other cytolysin monomers. When present in pores, mutant monomers typically retain the ability to refold to form the same 3D structure as wild-type cytolysin monomers. The 3D structure of cytolysin monomers in cytolysin pores is shown in Figure 7 of this paper. In addition to the mutations described herein, any number of mutations can be made in the wild-type cytolysin sequence, such as 2 to 100, 3 to 80, 4 to 70, 5 to 60, 10 to 50, or 20 to 40, provided that the cytolysin mutant monomer retains one or more of the enhanced properties conferred upon it by the mutations of the present invention.
[0517] Typically, when a cytolysin monomer is assembled with other identical mutant monomers or with different cytolysin mutant monomers to form a pore, the cytolysin monomer will retain the ability to contribute two β-sheets to the barrel of the cytolysin pore.
[0518] The variant further preferably comprises one or more of E84Q / E85K / E92Q / E97S / D126G, or, where appropriate, all of E84Q / E85K / E92Q / E97S / D126G. "Where appropriate" means whether these positions are still present in the mutant monomer or have not been modified by different amino acids.
[0519] In addition to the specific mutations discussed above, variants may contain other mutations. These mutations do not necessarily enhance the ability of monomers to interact with polynucleotides. Mutations may promote, for example, expression and / or purification. Within the entire length of the amino acid sequence of SEQ ID NO: 2, the variant will preferably be at least 50% homologous to the sequence based on amino acid similarity or identity. More preferably, based on amino acid similarity or identity, the variant may be at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, and more preferably at least 95%, 97%, or 99% homologous to the amino acid sequence of SEQ ID NO: 2 throughout the entire sequence. Within an extension of 100 or more, such as 125, 150, 175, or 200 or more consecutive amino acids, there may be at least 80%, such as at least 85%, 90%, or 95% amino acid similarity or identity (“hard homology”).
[0520] Standard methods in the field can be used to determine homology. For example, the UWGCG package provides the BESTFIT program, which can be used to calculate homology, for example, using its default settings (Devereux et al. (1984), Nucleic Acids Research, Vol. 12, pp. 387–395). The PILEUP and BLAST algorithms can be used to calculate homology or to orient sequences (e.g., to identify equivalent residues or corresponding sequences, usually using their default settings), for example, as described in Altschul SF (1993), Journal of Molecular Evolution, Vol. 36, pp. 290–300; Altschul, SF et al. (1990), Journal of Molecular Biology, Vol. 215, pp. 403–410. Software for performing BLAST analyses is available from the National Center for Biotechnology Information. Information (http: / / www.ncbi.nlm.nih.gov / ) is publicly available. Similarity can be measured using pairwise identity or by applying a scoring matrix such as BLOSUM62 and converting it to equivalent identity. Because these represent functional changes rather than evolutionary changes, they can mask the location of intentionally mutated mutations when determining homology. Similarity can be determined more sensitively by applying position-specific scoring matrices, such as PSIBLAST, on a comprehensive database of protein sequences. Different scoring matrices can be used that reflect substitution frequencies within an amino acid chemical-physical property rather than an evolutionary timescale (e.g., charge).
[0521] The amino acid sequence of SEQ ID NO: 2 can be substituted with amino acids other than those discussed above, for example, up to 1, 2, 3, 4, 5, 10, 20, or 30 substitutions. Conservative substitution uses other amino acids with similar chemical structures, similar chemical properties, or similar side chain sizes to replace the amino acid. The introduced amino acid can have similar polarity, hydrophilicity, hydrophobicity, basicity, acidity, electroneutrality, or charge as the amino acid it replaces. Alternatively, conservative substitution can introduce another aromatic or aliphatic amino acid to replace a pre-existing aromatic or aliphatic amino acid. Conservative amino acid changes are well known in the art and can be selected based on the characteristics of the 20 major amino acids defined in Table 3 below. In cases where the amino acids have similar polarity, this can also be determined with reference to the hydrophilicity scale of the amino acid side chains in Table 4.
[0522] Table 3 - Chemical properties of amino acids
[0523]
[0524] Table 4 - Hydrophilicity Scale
[0525]
[0526]
[0527] Variants may include one or more substitutions outside the region specified above, wherein an amino acid is replaced by an amino acid at one or more corresponding positions in homologs and parahomologs of cytolysin. Four examples of homologs of cytolysin are shown in SEQ ID NO: 14 to 17.
[0528] Alternatively, one or more amino acid residues of the amino acid sequence of SEQ ID NO: 2 may be deleted from the variant described above. Up to 1, 2, 3, 4, 5, 10, 20, or 30 or more residues may be deleted.
[0529] The variant may contain a fragment of SEQ ID NO: 2. This fragment retains pore-forming activity. This can be determined as described above. The fragment length can be at least 50, 100, 150, 200, or 250 amino acids. This fragment can be used to generate the pores of the present invention. Since the region from approximately position 44 to approximately position 126 of SEQ ID NO: 2 can be modified by one or more deletions according to the present invention, the fragment does not necessarily contain the entire region. Therefore, the present invention contemplates fragments shorter than the length of the unmodified region. The fragment preferably includes the pore-forming domain of SEQ ID NO: 2. More preferably, the fragment includes the region from approximately position 44 to approximately position 126 of SEQ ID NO: 2 modified according to the present invention.
[0530] Alternatively or additionally, one or more amino acids may be added to the variants described above. An extension may be provided at the amino or carboxyl terminus of the amino acid sequence of the variant containing the fragment of SEQ ID NO:2. The length of the extension may be very short, for example, from 1 to 10 amino acids. Alternatively, the extension may be longer, for example, up to 50 or 100 amino acids. A carrier protein may be fused to the amino acid sequence according to the invention. Other fusion proteins are discussed in more detail below.
[0531] As discussed above, the variant is a polypeptide having an amino acid sequence different from that of SEQ ID NO: 2 and retaining its ability to form pores. The variant typically contains the pore-forming region of SEQ ID NO: 2, approximately positions 44 to 126, and this region is modified according to the invention as discussed above. It may contain fragments of this region, as discussed above. In addition to the modifications of the invention, the variant of SEQ ID NO: 2 may contain one or more additional modifications, such as substitution, addition, or deletion. These modifications are preferably located in the extensions of the variant corresponding to approximately positions 1 to 43 and approximately positions 127 to 297 of SEQ ID NO: 2 (i.e., outside the region modified according to the invention).
[0532] The mutant monomer can be modified, for example, by adding histidine residues (his tags), aspartic acid residues (asp tags), streptavidin tags, or flag tags, or by adding signal sequences that promote the secretion of mutant monomers from cells in which the polypeptide does not naturally contain such sequences, to aid in their identification or purification. An alternative approach to introducing gene tags is to chemically react the tags onto native or engineered sites on the pores. An example of this would be reacting a gel migration reagent with engineered cysteine residues on the pore exterior. This has been shown as a method for separating hemolysin heterooligomers (Chem Biol, July 1997, Vol. 4, No. 7, pp. 497–505).
[0533] Reveal tags can be used to label mutant monomers. Reveal tags can be any suitable tag that allows the well to be detected. Suitable tags include, but are not limited to, fluorescent molecules; radioactive isotopes, for example, 125 I, 35 S, enzymes, antibodies, antigens, polynucleotides, polyethylene glycol (PEG), peptides, and ligands such as biotin.
[0534] D-amino acids can also be used to generate mutant monomers. For example, mutant monomers can include a mixture of L-amino acids and D-amino acids. This is common practice in the field of production of such proteins or peptides.
[0535] The mutant monomer contains one or more specific modifications to promote interaction with polynucleotides. The mutant monomer may also contain other non-specific modifications, provided that these modifications do not interfere with pore formation. Various non-specific side-chain modifications are known in the art and can be performed on the side chains of the mutant monomer. Such modifications include, for example, reductive alkylation of amino acids by reaction with an aldehyde followed by reduction with NaBH4, amidation with methylacetylimine, or acylation with acetic anhydride.
[0536] Mutant monomers can be generated using standard methods known in the field. Monomers can be prepared synthetically or through recombination. For example, monomers can be synthesized via in vitro translation and transcription (IVTT). Appropriate methods for generating pore monomers are discussed in international applications PCT / GB09 / 001690 (published as WO 2010 / 004273), PCT / GB09 / 001679 (published as WO 2010 / 004265), or PCT / GB10 / 000133 (published as WO 2010 / 086603). Methods for inserting pores into membranes are discussed below.
[0537] The polynucleotide sequence encoding the mutant monomer can be derived or replicated using standard methods in the field. This sequence is discussed in more detail below. The polynucleotide sequence encoding the mutant monomer can be expressed in bacterial host cells using standard techniques in the field. The mutant monomer can be generated from a recombinant expression vector via in situ expression of the polypeptide in cells. The expression vector optionally carries an inducible promoter for controlling polypeptide expression.
[0538] Mutant monomers can be produced on a large scale after purification from the pore-generating organism using any protein liquid chromatography system, or after recombinant expression, as described below. Typical protein liquid chromatography systems include FPLC, AKTA systems, Bio-Cad systems, Bio-Rad biosystems, and Gilson HPLC systems. The mutant monomers can then be inserted into naturally occurring or artificial membranes for use according to the present invention. Methods for inserting pores into membranes are discussed below.
[0539] In some embodiments, the mutant monomer is chemically modified. The mutant monomer can be chemically modified in any manner and at any site. Preferably, the mutant monomer is chemically modified by attaching the molecule to one or more cysteine residues (cysteine linkage), attaching the molecule to one or more lysine residues, attaching the molecule to one or more unnatural amino acids, enzymatic modification of the epitope, or terminal modification. Suitable methods for performing such modifications are well known in the art. Suitable unnatural amino acids include, but are not limited to, 4-azido-L-phenylalanine (Faz) and Liu CC and Schultz PG, *Annu. Rev. Biochem.*, 2010, Vol. 79, pp. 413–444. Figure 1 Any amino acid numbered 1 to 71. Mutant monomers can be chemically modified by attaching any molecule. For example, mutant monomers can be chemically modified by attaching polyethylene glycol (PEG), nucleic acids such as DNA, dyes, fluorophores, or chromophores.
[0540] In some embodiments, the mutant monomer is chemically modified using a molecular adaptor that promotes the interaction between the pore, comprising the monomer, and the target analyte, target nucleotide, or target polynucleotide sequence. The presence of the adaptor improves the host-guest chemistry between the pore and the nucleotide or polynucleotide, and thereby enhances the sequencing capability of the pore formed from the mutant monomer. The principles of host-guest chemistry are well known in the art. The adaptor affects the physical or chemical properties of the pore, which enhances its interaction with the nucleotide or polynucleotide sequence. The adaptor may alter the charge of the pore's barrel or channel, or specifically interact with or bind to a nucleotide or polynucleotide, thereby promoting its interaction with the pore.
[0541] Molecular linkers are preferably cyclic molecules such as cyclodextrins, hybridizable species, DNA binders or intercalators, peptides or peptide analogs, synthetic polymers, aromatic planar molecules, positively charged small molecules, or small molecules capable of hydrogen binding.
[0542] The connector can be ring-shaped. The ring-shaped connector preferably has the same symmetry as the hole.
[0543] Integrators typically interact with analytes, nucleotides, or polynucleotides through host-guest chemistry. Integrators are generally capable of interacting with nucleotides or polynucleotides. An integrator comprises one or more chemical groups capable of interacting with nucleotides or polynucleotides. These one or more chemical groups preferably interact with the nucleotide or polynucleotide through non-covalent interactions, such as hydrophobic interactions, hydrogen binding, van der Waals forces, π-cation interactions, and / or electrostatic forces. The one or more chemical groups capable of interacting with nucleotides or polynucleotides are preferably positively charged. More preferably, the one or more chemical groups capable of interacting with nucleotides or polynucleotides include amino groups. The amino group may be attached to a primary, secondary, or tertiary carbon atom. Even more preferably, the integrator comprises an amino ring, such as a ring consisting of 6, 7, 8, or 9 amino groups. Most preferably, the integrator comprises a ring consisting of 6 or 9 amino groups. The protonated amino ring can interact with a negatively charged phosphate group in the nucleotide or polynucleotide.
[0544] The correct localization of the intransitive linker within the pore can be facilitated through host-guest chemistry between the intransitive linker and the pore containing the mutant monomer. The intransitive linker preferably comprises one or more chemical groups capable of interacting with one or more amino acids in the pore. More preferably, the intransitive linker comprises one or more chemical groups capable of interacting with one or more amino acids in the pore through non-covalent interactions such as hydrophobic interactions, hydrogen bonding, van der Waals forces, π-cation interactions, and / or electrostatic forces. The chemical groups capable of interacting with one or more amino acids in the pore are typically hydroxyl or amine groups. The hydroxyl group may be attached to a primary, secondary, or tertiary carbon atom. The hydroxyl group may form hydrogen bonds with uncharged amino acids in the pore. Any intransitive linker that promotes interaction between the pore and the nucleotide or polynucleotide can be used.
[0545] Suitable linkers include, but are not limited to, cyclodextrins, cyclic peptides, and cucurbiturils. Linkers are preferably cyclodextrins or derivatives thereof. A cyclodextrin or derivative thereof may be any cyclodextrin or derivative thereof disclosed in Elisev, AV and Schneider, HJ. (1994), *Journal of the American Chemical Society*, Vol. 116, pp. 6081-6088. Linkers are more preferably hepta-6-amino-β-cyclodextrin (am7-βCD), 6-monodeoxy-6-monoamino-β-cyclodextrin (am1-βCD), or hepta-(6-deoxy-6-guanidinyl)-cyclodextrin (gu7-βCD). The guanidinyl group in gu7-βCD has a much higher pKa than the primary amine in am7-βCD, and therefore carries a greater positive charge. This gu7-βCD adapter can be used to increase the residence time of nucleotides in the pore, increase the accuracy of the measured residual current, and increase the base detection rate at high temperatures or low data acquisition rates.
[0546] If a 3-(2-pyridinedithio)propionate succinimide (SPDP) crosslinking agent is used, as discussed in more detail below, the linker is preferably hepta(6-deoxy-6-amino)-6-N-mono(2-pyridyl)dithiopropionyl-β-cyclodextrin (am6amPDP1-βCD).
[0547] More suitable linkers include γ-cyclodextrin, which comprises eight sugar units (and thus has octet symmetry). γ-cyclodextrin may contain linker molecules or may be modified to include all or more of the modified sugar units used in the β-cyclodextrin examples discussed above.
[0548] The molecular adaptor is preferably covalently linked to the mutant monomer. The adaptor can be covalently linked to the pore using any method known in the art. Typically, the adaptor is attached by chemical linkage. If the molecular adaptor is attached via a cysteine linkage, then one or more cysteine residues have preferably been introduced into the mutant by substitution. The mutant monomer of the present invention may, of course, include cysteine residues at one or both of positions 272 and 283. The mutant monomer can be chemically modified by attaching the molecular adaptor to one or both of these cysteine residues. Alternatively, the mutant monomer can be chemically modified by attaching the molecule to one or more cysteine residues introduced at other positions or to non-natural amino acids such as FAz.
[0549] The reactivity of cysteine residues can be enhanced by modifying adjacent residues. For example, attaching a basic group to an arginine, histidine, or lysine residue can alter the pKa of the cysteine thiol group to make it more reactive. -The pKa of the cysteine residues can be protected by thiol protecting groups such as dTNB. These can react with one or more cysteine residues of the mutant monomer before attachment to the linker. The molecule can be directly attached to the mutant monomer. Preferably, the molecule is attached to the mutant monomer using a linker such as a chemical crosslinking agent or a peptide linker.
[0550] Suitable chemical crosslinking agents are well known in the art. Preferred crosslinking agents include 2,5-dioxopyrrolidone-1-yl ester of 3-(pyridin-2-yldisulfonyl)propionate, 2,5-dioxopyrrolidone-1-yl ester of 4-(pyridin-2-yldisulfonyl)butyrate, and 2,5-dioxopyrrolidone-1-yl ester of 8-(pyridin-2-yldisulfonyl)octanoate. The most preferred crosslinking agent is succinimide ester of 3-(2-pyridinedithio)propionate (SPDP). Typically, the molecule is covalently linked to the bifunctional crosslinking agent before the molecule / crosslinking agent complex is covalently linked to the mutant monomer, but it is also possible to covalently link the bifunctional crosslinking agent to the monomer before the bifunctional crosslinking agent / monomer complex is attached to the molecule.
[0551] Preferably, the joint is resistant to dithiothreitol (DTT). Suitable joints include, but are not limited to, iodoacetamide-based and maleimide-based joints.
[0552] In other embodiments, monomers may attach to polynucleotide-binding proteins. This forms a modular sequencing system that can be used in the methods of the present invention. Polynucleotide-binding proteins are discussed below.
[0553] Polynucleotide-binding proteins can be covalently linked to mutant monomers. Proteins can be covalently linked to pores using any method known in the art. The monomer and protein can be chemically fused or genetically fused. If the entire construct is expressed from a single polynucleotide sequence, the monomer and protein are genetically fused. Genetic fusion of monomers with polynucleotide-binding proteins is discussed in International Application No. PCT / GB09 / 001679 (published as WO 2010 / 004265).
[0554] If a polynucleotide-binding protein is attached via a cysteine linker, one or more cysteine residues have preferably been introduced into the mutant through substitution. This substitution typically occurs in a loop region that is lowly conserved among homologs, indicating tolerance to mutations or insertions. Therefore, it is suitable for attaching polynucleotide-binding proteins. This substitution is typically made at residues 1 to 43 and 127 to 297 of SEQ ID NO: 2. The reactivity of the cysteine residues can be enhanced by modifications as described above.
[0555] Polynucleotide-binding proteins can be directly attached to mutant monomers or attached via one or more adapters. Hybrid adapters described in International Application No. PCT / GB10 / 000132 (published as WO 2010 / 086602) can be used to attach polynucleotide-binding proteins to mutant monomers. Alternatively, peptide adapters can be used. Peptide adapters are amino acid sequences. The length, flexibility, and hydrophilicity of peptide adapters are generally designed so that they do not interfere with the function of the monomer and molecule. Preferred flexible peptide adapters are extensions of 2 to 20, such as 4, 6, 8, 10, or 16 serine and / or glycine residues. More preferred flexible adapters comprise (SG)1, (SG)2, (SG)3, (SG)4, (SG)5, and (SG)8, where S is serine and G is glycine. Preferred rigid adapters are extensions of 2 to 30, such as 4, 6, 8, 16, or 24 proline residues. More preferred rigid adapters comprise (P) 12 , where P is proline.
[0556] Mutant monomers can be chemically modified using molecular linkers and polynucleotide binding proteins.
[0557] Preparation of mutant cytolysin monomers
[0558] The present invention also provides a method for improving the ability of a cytolysin monomer comprising the sequence shown in SEQ ID NO: 2 to characterize polynucleotides. The method includes making one or more modifications and / or substitutions of the present invention in SEQ ID NO: 2. Any embodiments of the above-mentioned mutant cytolysin monomers and any embodiments discussed below for characterizing polynucleotides are equally applicable to this method of the present invention.
[0559] Construct
[0560] The present invention also provides a construct comprising two or more covalently linked monomers derived from cytolysin, wherein at least one of the monomers is a mutant cytolysin monomer of the present invention. The construct of the present invention retains its ability to form wells. One or more constructs of the present invention can be used to form wells for characterizing target analytes. One or more constructs of the present invention can be used to form wells for characterizing target polynucleotides, such as for sequencing target nucleotides. A construct may comprise 2, 3, 4, 5, 6, 7, 8, 9, or 10 or more monomers. The two or more monomers may be the same or different.
[0561] At least one monomer in the construct is a mutant monomer of the present invention. Two or more, three or more, four or more, five or more, six or more, seven or more, eight or more, nine or more, or ten or more monomers in the construct may be mutant monomers of the present invention. All monomers in the construct are preferably mutant monomers of the present invention. The mutant monomers may be the same or different. In a preferred embodiment, the construct includes two mutant monomers of the present invention.
[0562] The mutant monomers of the present invention in the construct are preferably of substantially the same or identical length. The barrels of the mutant monomers of the present invention in the construct are preferably of substantially the same or identical length. Length can be measured in the form of amino acid number and / or length units. The amino acid number of the mutant monomers of the present invention in the construct is preferably the same as the number of amino acids missing from positions 34 to 70 and / or positions 71 to 107 as described above.
[0563] Other monomers in the construct need not be mutant monomers of the present invention. For example, at least one monomer may include the sequence shown in SEQ ID NO: 2. At least one monomer in the construct may be a paralog or homolog of SEQ ID NO: 2. Suitable homologs are shown in SEQ ID NO: 14 to 17.
[0564] Alternatively, at least one monomer may comprise a variant of SEQ ID NO: 2 that is at least 50% homologous to SEQ ID NO: 2 throughout its entire sequence based on amino acid identity, but does not contain any of the specific mutations required for the mutant monomer of the present invention, or in said variant, an amino acid is not yet missing, as described above. More preferably, based on amino acid identity, the variant may be at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, and more preferably at least 95%, 97%, or 99% homologous to the amino acid sequence of SEQ ID NO: 2 throughout its entire sequence. The variant may be a fragment or any other variant discussed above. The constructs of the present invention may also comprise variants of SEQ ID NO: 14, 15, 16, or 17 that are at least 50% homologous to SEQ ID NO: 14, 15, 16, or 17 throughout its entire sequence based on amino acid identity, or at least any of the other homology levels mentioned above.
[0565] All monomers in the construct can be mutant monomers of the present invention. The mutant monomers can be the same or different. In a more preferred embodiment, the construct comprises two monomers, and at least one of the monomers is a mutant monomer of the present invention.
[0566] Monomers can be gene fusions. A monomer is a gene fusion if the entire construct is expressed from a single polynucleotide sequence. The coding sequences of monomers can be combined in any way to form a single polynucleotide sequence encoding the construct. Gene fusions are discussed in International Application No. PCT / GB09 / 001679 (published as WO 2010 / 004265).
[0567] Monomeric genes can be fused in any configuration. Monomers can be fused by their terminal amino acids. For example, the amino terminus of one monomer can be fused to the carboxyl terminus of another monomer.
[0568] Two or more monomers can be directly fused together. Preferably, a linker is used for gene fusion of monomers. The linker can be designed to restrict monomer mobility. Preferred linkers are amino acid sequences (i.e., peptide linkers). Any peptide linker discussed above can be used.
[0569] The length, flexibility, and hydrophilicity of peptide linkers are typically designed so as not to interfere with the function of monomers and molecules. Preferred flexible peptide linkers are those with 2 to 20 extensions, such as 4, 6, 8, 10, or 16 serine and / or glycine residues. More preferred flexible linkers comprise (SG)1, (SG)2, (SG)3, (SG)4, (SG)5, and (SG)8, where S is serine and G is glycine. Preferred rigid linkers are those with 2 to 30 extensions, such as 4, 6, 8, 16, or 24 proline residues. More preferred rigid linkers comprise (P) 12 , where P is proline.
[0570] In another preferred embodiment, the monomers are chemically fused. The monomers are chemically fused if, for example, they are chemically attached by a chemical crosslinking agent. Any of the chemical crosslinking agents discussed above can be used. The linker can be attached to one or more cysteine residues or non-natural amino acids such as Faz introduced into the mutant monomer. Alternatively, the linker can be attached to the end of one of the monomers in the construct. The monomers are typically linked by one or more of residues 1 to 43 and 127 to 297 of SEQ ID NO: 2.
[0571] If the construct contains different monomers, cross-linking of the monomers themselves can be prevented by making the concentration of the joints significantly higher than that of the monomers. Alternatively, a "lock and key" arrangement in which two joints are used can be employed. Only one end of each joint can react together to form a longer joint, and the other ends of the joint react with different monomers. Such joints are described in International Application No. PCT / GB10 / 000132 (published as WO 2010 / 086602).
[0572] The present invention also provides a method for generating the constructs of the present invention. The method includes: covalently linking at least one mutant cytolysin monomer of the present invention to one or more monomers derived from cytolysin. Any embodiments in the examples discussed above with reference to the constructs of the present invention are equally applicable to the method for generating the constructs.
[0573] Polynucleotides
[0574] This invention also provides polynucleotides encoding mutant monomers of the invention. The mutant monomer can be any of the mutant monomers discussed above. Based on nucleotide identity, the polynucleotide sequence preferably includes a sequence that is at least 50%, 60%, 70%, 80%, 90%, or 95% homologous to the sequence of SEQ ID NO: 1 throughout the entire sequence. Within an extension of 300 or more, for example 375, 450, 525, or 600 or more consecutive nucleotides, there may be at least 80%, for example at least 85%, 90%, or 95% nucleotide identity (“hard homology”). Homology can be calculated as described above. On the basis of degeneracy of the genetic code, the polynucleotide sequence may include a sequence different from SEQ ID NO: 1.
[0575] The present invention also provides a polynucleotide sequence encoding any construct of the gene fusion construct of the present invention. The polynucleotide preferably comprises two or more sequences as shown in SEQ ID NO: 1 or variants thereof as described above.
[0576] The polynucleotide sequence can be derived or replicated using standard methods in the field. Chromosomal DNA encoding wild-type cytosolic proteins can be extracted from pore-producing organisms such as *Eisenia fetida*. The gene encoding the pore monomer can be amplified using PCR involving specific primers. The amplified sequence can then be subjected to site-directed mutagenesis. Suitable site-directed mutagenesis methods are known in the field and involve, for example, combinatorial chain reactions. The polynucleotide encoding the constructs of the present invention can be prepared using well-known techniques, such as those described in *Molecular Cloning: A Laboratory Manual*, 3rd edition, Cold Spring Harbor Laboratory Press, Sambrook, J. and Russell, D. (2001).
[0577] The obtained polynucleotide sequence can then be incorporated into a recombinant reproducible vector, such as a cloning vector. The vector can then be used to replicate the polynucleotide in a compatible host cell. Therefore, a polynucleotide sequence can be prepared by introducing a polynucleotide into a reproducible vector, introducing the vector into a compatible host cell, and growing the host cell under conditions that induce vector replication. The vector can be recovered from the host cell. Suitable host cells for cloning polynucleotides are known in the art and are described in more detail below.
[0578] Polynucleotide sequences can be cloned into appropriate expression vectors. In these vectors, the polynucleotide sequence is typically operatively linked to a control sequence that enables the expression of the coding sequence via the host cell. Such expression vectors can be used to express pore subunits.
[0579] The term "operably ligated" refers to juxtaposition, where the described components are in a relationship that allows them to function in their intended manner. "Operably ligated" to a control sequence of a coding sequence is a ligation in a manner that enables expression of the coding sequence under conditions compatible with the control sequence. Multiple copies of the same or different polynucleotide sequences can be introduced into the vector.
[0580] The expression vector can then be introduced into a suitable host cell. Therefore, the mutant monomers or constructs of the present invention can be generated by inserting a polynucleotide sequence into an expression vector, introducing the vector into a compatible bacterial host cell, and growing the host cell under conditions that induce expression of the polynucleotide sequence. The monomers or constructs expressed in a recombinant manner can self-assemble into pores in the host cell membrane. Alternatively, the recombinant pores generated in this manner can be removed from the host cell and inserted into another membrane. When pores comprising at least two different subunits are generated, the different subunits can be expressed individually in different host cells as described above, removed from the host cell, and assembled into pores in separate membranes such as sheep erythrocyte membranes or liposomes containing sphingomyelin.
[0581] For example, cytolysin monomers can be oligomerized by adding a mixture of lipids including sphingomyelin and one or more of the following lipids and incubating the mixture at 30°C for 60 minutes, for example: phosphatidylserine; POPE; cholesterol; and Soy PC. The oligomers can be purified by any suitable method, such as by SDS-PAGE and gel purification as described in WO2013 / 153359.
[0582] The vector can be, for example, a plasmid, viral, or phage vector having a replication origin, an optional promoter for expressing the polynucleotide sequence, and optional regulators of the promoter. The vector may contain one or more optional marker genes, such as a tetracycline resistance gene. The promoter and other expression regulatory signals can be selected to be compatible with the host cell to which the expression vector is designed. T7, trc, lac, ara, or λ are commonly used. L Promoter.
[0583] Host cells typically express the pore subunit at high levels. Host cells transformed using a polynucleotide sequence can be selected to be compatible with the expression vector used for the transformed cells. Host cells are typically bacteria and preferably *Escherichia coli*. Any cell possessing a λDE3 lysogen, such as C41(DE3), BL21(DE3), JM109(DE3), B834(DE3), TUNER, Origami, and OrigamiB, can express vectors including the T7 promoter. In addition to the conditions listed above, cytolysin proteins can also be expressed using any of the methods described in *Proceedings of the National Academy of Sciences of the United States of America (Proc Natl Acad Sci USA)*, December 30, 2008, Vol. 105, No. 52, pp. 20647–20642.
[0584] hole
[0585] This invention also provides various wells. The wells of this invention are ideal for characterizing analytes. The wells of this invention are particularly ideal for characterizing polynucleotide sequences, such as sequencing polynucleotides, because they can distinguish different nucleotides with high sensitivity. The wells can be used to characterize nucleic acids such as DNA and RNA, including sequencing nucleic acids and recognizing single-base changes. The wells of this invention can even distinguish between methylated and unmethylated nucleotides. The basic resolution of the wells of this invention is very high. The wells show almost complete separation of all four DNA nucleotides. The wells can be further used to distinguish between deoxycytidine monophosphate (dCMP) and methyl-dCMP based on residence time in the well and current flowing through the well.
[0586] The pores of this invention can also distinguish different nucleotides under a range of conditions. Specifically, the pores will distinguish nucleotides under conditions favorable for characterizing polynucleotides, such as sequencing them. The degree to which the pores of this invention can distinguish different nucleotides can be controlled by changing the applied potential, salt concentration, buffer solution, temperature, and the presence of additives such as urea, betaine, and DTT. This allows for fine-tuning of the pore's function, especially during sequencing. This is discussed in more detail below. The pores of this invention can also be used to identify polynucleotide polymers based on their interaction with one or more monomers, rather than on a nucleotide-by-nucleotide basis.
[0587] The pores of the present invention can be isolated, substantially isolated, purified, or substantially purified. If the pores of the present invention are completely free of any other components, such as lipids or other pores, they are isolated or purified. If the pores are mixed with a carrier or diluent that will not interfere with their intended use, they are substantially isolated. For example, if the pores are present in the form of less than 10%, less than 5%, less than 2%, or less than 1% of other components such as lipids or other pores, they are substantially isolated or substantially purified. Alternatively, the pores of the present invention can exist in a lipid bilayer.
[0588] The pores of the present invention can exist as a single pore or a single pore. Alternatively, the pores of the present invention can exist as a homologous group or heterologous group of two or more pores, or as multiple groups of two or more pores.
[0589] Homologous oligomerization pores
[0590] The present invention also provides a homologous oligopore derived from cytosin, comprising a consistent mutant monomer of the present invention. The monomers are consistent in terms of their amino acid sequences. The homologous oligopore of the present invention is ideal for characterizing polynucleotides, such as sequencing them. The homologous oligopore of the present invention can possess any of the advantages discussed above. The advantages of specific homologous oligopores of the present invention are illustrated in the examples.
[0591] Homologous oligomer wells can contain any number of mutant monomers. The wells typically include two or more mutant monomers. Homologous oligomer wells can contain any number of mutant monomers. The wells typically include at least 6, at least 7, at least 8, at least 9, or at least 10 identical mutant monomers, such as 6, 7, 8, 9, or 10 mutant monomers. The wells preferably include eight or nine identical mutant monomers. The wells most preferably include nine identical mutant monomers. This number of monomers is referred to herein as a “sufficient number.”
[0592] One or more of the mutant monomers, such as 2, 3, 4, 5, 6, 7, 8, 9 or 10, are preferably chemically modified as discussed above or below.
[0593] One or more of the mutant monomers are preferably chemically modified as discussed above or below. In other words, as long as the amino acid sequence of each of the monomers is consistent, one or more of the chemically modified monomers (as well as other unmodified monomers) will not prevent the pores from becoming homooligomers.
[0594] A method for preparing cytosolic pores is described by Yamaji et al. in the Journal of Biochemistry, 1998, Vol. 273, No. 9, pp. 5300-5306.
[0595] Heterogeneous oligomerization pores
[0596] This invention also provides a heterooligomeric pore derived from cytosin, comprising at least one mutant monomer of this invention, wherein at least one of the monomers differs from the other monomers. The monomer differs from the other monomers in terms of its amino acid sequence. The heterooligomeric pore of this invention is ideal for characterizing polynucleotides, such as sequencing them. The heterooligomeric pore can be prepared using methods known in the art (e.g., *Protein Science*, July 2002, Vol. 11, No. 7, pp. 1813-1824).
[0597] The heterooligomeric pores contain sufficient monomers to form pores. The monomers can be of any type, including but not limited to wild-type. The pores typically comprise two or more monomers. The pores typically comprise at least 6, 7, 8, 9, or 10 monomers, such as 6, 7, 8, 9, or 10 monomers. The pores preferably comprise eight or nine monomers. The pores most preferably comprise nine monomers. This number of monomers is referred to herein as a “sufficient number.”
[0598] The pore comprises at least one monomer comprising the sequence shown in SEQ ID NO: 2, its paralog, its homolog, or a variant thereof, wherein the variant does not have the mutation required by the mutant monomer of the present invention or, in the variant, an amino acid is not yet missing, as described above. Suitable variants are any of the variants discussed above with reference to the constructs of the present invention, including SEQ ID NO: 2, 14, 15, 16, and 17 and their variants. In this embodiment, the remaining monomer is preferably a mutant monomer of the present invention.
[0599] In a preferred embodiment, the pore comprises (a) a mutant monomer of the present invention and (b) a number of consistent monomers sufficient to form the pore, wherein the mutant monomer in (a) is different from the consistent monomer in (b). The consistent monomer in (b) preferably comprises the sequence shown in SEQ ID NO: 2, its paralogs, its homologs, or variants thereof, said variants not having the mutation required by the mutant monomer of the present invention.
[0600] The heterologous oligomer pores of the present invention preferably comprise only one mutant cytolysin monomer of the present invention.
[0601] In another preferred embodiment, all monomers in the heterologous oligomer pores are mutant monomers of the present invention, and at least one of them is different from the other monomers.
[0602] The lengths of the mutant monomers of the present invention in the wells are preferably substantially the same or identical. The lengths of the barrels of the mutant monomers of the present invention in the wells are preferably substantially the same or identical. The length can be measured in the form of amino acid number and / or length units. The number of amino acids in the mutant monomers of the present invention in the wells is preferably the same as the number of amino acids missing from position 34 to position 70 and / or position 71 to position 107.
[0603] In all the embodiments discussed above, one or more of the mutant monomers are preferably chemically modified as discussed above or below. The presence of chemical modification on a monomer does not result in a pore being a heterooligomer. The amino acid sequence of at least one monomer must differ from the sequences of one or more of the other monomers. Methods for preparing the pores are discussed in more detail below.
[0604] Holes containing the building blocks
[0605] The present invention also provides a pore comprising at least one construct of the present invention. The construct of the present invention comprises two or more covalently linked monomers derived from cytosin, wherein at least one of the monomers is a mutant cytosin monomer of the present invention. In other words, the construct must contain more than one monomer. At least two of the monomers in the pore are in the form of the construct of the present invention. The monomers can be of any type.
[0606] A pore typically contains (a) a construct comprising two monomers and (b) a number of monomers sufficient to form a pore. The construct can be any of the constructs discussed above. The monomers can be any of the monomers discussed above, including the mutant monomers of this invention.
[0607] Another typical pore includes more than one construct of the present invention, such as two, three, or four constructs of the present invention. Such pores further include a sufficient number of monomers to form the pore. The monomers can be any of the mutant monomers discussed above. Further pores of the present invention include only constructs containing two monomers. Specific pores according to the present invention include several constructs each comprising two monomers. The constructs can oligomerize to form pores having a structure such that only one monomer from each construct contributes to the pore. Typically, the other monomers of the construct (i.e., monomers that do not form pores) will be located outside the pore.
[0608] Mutations can be introduced into the construct as discussed above. Mutations can be alternating, meaning the mutation is different for each monomer within the bimonomer construct, and the construct assembles into homologous oligomers, resulting in alternating modifications. In other words, monomers comprising MutA and MutB are fused and assembled to form AB:AB:AB:AB pores. Alternatively, mutations can be adjacent, meaning a consistent mutation is introduced into both monomers in the construct, and this is then oligomerized with monomers containing different mutations. In other words, monomers comprising MutA are fused, followed by oligomerization with monomers containing MutB to form AA:B:B:B:B:B:B.
[0609] One or more of the monomers of the invention in the pores containing the construct can be chemically modified as discussed above or below.
[0610] The chemically modified pores of the present invention
[0611] On the other hand, the present invention provides a chemically modified cytolysin pore comprising one or more mutant monomers, said mutant monomers being chemically modified such that the opening diameter of the barrel / channel of the assembled pore decreases, narrows, or contracts at one or more sites, such as two, three, four, or five sites, along the length of the barrel. The pore may comprise any number of monomers discussed above with reference to the homooligomeric and heterooligomeric pores of the present invention. The pore preferably comprises nine chemically modified monomers. The chemically modified pore may be homooligomers, as described above. In other words, all monomers in the chemically modified pore may have the same amino acid sequence and may be chemically modified in the same manner. The chemically modified pore may be heterooligomers, as described above. In other words, the pore may include (a) only one chemically modified monomer, (b) more than one, such as two, three, four, five, six, seven, or eight chemically modified monomers, wherein at least two of the chemically modified monomers, such as three, four, five, six, or seven, are different from each other, or (c) only chemically modified monomers (i.e., all monomers are chemically modified), wherein at least two of the chemically modified monomers, such as three, four, five, six, seven, eight, or nine, are different from each other. The monomers may be different from each other in terms of their amino acid sequence, their chemical modification, or both their amino acid sequence and their chemical modification. One or more chemically modified monomers may be any of the chemically modified monomers discussed above and / or below.
[0612] The present invention also provides a mutant cytolysin monomer chemically modified in any of the manner discussed below. The mutant monomer can be any of the mutant monomers discussed above or below. Therefore, the mutant monomers of the present invention, such as variants of SEQ ID NO: 2 including modifications at one or more of the following positions: K37, G43, K45, V47, S49, T51, H83, V88, T91, T93, V95, Y96, S98, K99, V100, I101, P108, P109, T110, S111, K112, and T114, or variants including barrel deletions of the above, can be chemically modified according to the present invention, as discussed below.
[0613] The mutant monomer can be chemically modified such that the diameter of the barrel of the assembled well is reduced or narrowed by any reduction factor depending on the size of the analyte to be transferred through the well. The width of the contraction zone will generally determine the extent to which the measurement signal is disrupted during analyte transfer due to, for example, the reduced ion flow through the well caused by the analyte. Greater signal disruption generally results in higher measurement sensitivity. Therefore, the contraction zone can be chosen to be slightly wider than the analyte to be transferred. For example, for the transfer of ssDNA, the width of the contraction zone can be selected from values in the range of 0.8 nm to 3.0 nm.
[0614] Chemical modifications can also determine the length of the contraction region, which in turn determines the number of polymeric units, such as nucleotides, contributing to the measurement signal. The nucleotides contributing to the current signal at any given time can be referred to as k-mers, where k is an integer and can be an integer or a fraction. In the case of measuring polynucleotides with four types of nucleobases, trimers will produce 4... 3 There are several potential signal levels. A larger k value will produce a larger number of signal levels. A short contraction region is generally desirable because it simplifies the analysis of the measured signal data.
[0615] Chemical modification is performed to preferably covalently link a chemical molecule to a mutant monomer or one or more mutant monomers. Any method known in the art can be used to covalently link a chemical molecule to a pore, a mutant monomer, or one or more mutant monomers. Chemical molecules are typically attached via chemical linkages.
[0616] Preferably, the mutant monomer or one or more mutant monomers are chemically modified by attaching the molecule to one or more cysteine residues (cysteine linkage), attaching the molecule to one or more lysine residues, attaching the molecule to one or more unnatural amino acids, or by enzymatic modification of the epitope. If the chemical modifier is attached via cysteine linkage, the one or more cysteine residues have preferably been introduced into the mutant by substitution. Suitable methods for performing such modifications are well known in the art. Suitable unnatural amino acids include, but are not limited to, 4-azido-L-phenylalanine (Faz) and Liu CC and Schultz PG, *Annu. Rev. Biochem.*, 2010, Vol. 79, pp. 413-444. Figure 1 Any amino acid numbered 1 to 71 in the amino acid spectrum.
[0617] Mutant monomers or one or more mutant monomers can be chemically modified by attaching any molecule that has the effect of reducing or shrinking the diameter of the assembly pore at any location or site. Mutant monomers can be chemically modified by attaching the following: (i) maleimides, such as 4-phenylazomaleinanil, 1,N-(2-hydroxyethyl)maleimide, N-cyclohexylmaleimide, 1,3-maleimide propionic acid, 1,1-4-aminophenyl-1H-pyrrole,2,5,dione, 1,1-4-hydroxyphenyl-1H-pyrrole,2,5,dione, N-ethylmaleimide. Imide, N-methoxycarbonylmaleimide, N-tert-butylmaleimide, N-(2-aminoethyl)maleimide, 3-maleimide-PROXYL, N-(4-chlorophenyl)maleimide, 1-[4-(dimethylamino)-3,5-dinitrophenyl]-1H-pyrrole-2,5-dione, N-[4-(2-benzimidazolyl)phenyl]maleimide, N-[4-(2-benzoxazolyl)phenyl]maleimide, N- (1-Naphthyl)maleimide, N-(2,4-dimethyl)maleimide, N-(2,4-difluorophenyl)maleimide, N-(3-chloro-p-tolyl)maleimide, 1-(2-amino-ethyl)-pyrrole-2,5-dione hydrochloride, 1-cyclopentyl-3-methyl-2,5-dihydro-1H-pyrrole-2,5-dione, 1-(3-aminopropyl)-2,5-dihydro-1H-pyrrole-2,5-dione hydrochloride, 3 -Methyl-1-[2-oxo-2-(piperazin-1-yl)ethyl]-2,5-dihydro-1H-pyrrole-2,5-dione hydrochloride, 1-benzyl-2,5-dihydro-1H-pyrrole-2,5-dione, 3-methyl-1-(3,3,3-trifluoropropyl)-2,5-dihydro-1H-pyrrole-2,5-dione, 1-[4-(methylamino)cyclohexyl]-2,5-dihydro-1H-pyrrole-2,5-dione trifluoroacetic acid, SMILES O=C1C=CC(=O)N1CC=2C=CN=CC2、SMILES O=C1C=CC(=O)N1CN2CCNCC2、1-Benzyl-3-methyl-2,5-dihydro-1H-pyrrole-2,5-dione、1-(2-fluorophenyl)-3-methyl-2,5-dihydro-1H-pyrrole-2,5-dione、N-(4-phenoxyphenyl)maleimide、N-(4-nitrophenyl)maleimide;(ii) Iodoacetamides, such as 3-(2-iodoacetamido)-PROXYL, N-(cyclopropylmethyl)-2-iodoacetamide, 2-iodo-N-(2-phenylethyl)acetamide, 2-iodo-N-(2,2,2-trifluoroethyl)acetamide, N-(4-acetylphenyl)-2-iodoacetamide, N-(4-(aminosulfonyl)phenyl)-2-iodoacetamide, N-(1,3-benzothiazolyl-2-yl)-2-iodoacetamide. Amides, N-(2,6-diethylphenyl)-2-iodoacetamide, N-(2-benzoyl-4-chlorophenyl)-2-iodoacetamide; (iii) Bromoacetamides: such as N-(4-(acetamido)phenyl)-2-bromoacetamide, N-(2-acetylphenyl)-2-bromoacetamide, 2-bromo-N-(2-cyanophenyl)acetamide, 2-bromo-N-(3-(trifluoromethyl)phenyl)acetamide, N-(2-benzoyl)acetamide, etc. 2-Bromo-N-(4-fluorophenyl)-3-methylbutyramide, N-benzyl-2-bromo-N-phenylpropionamide, N-(2-bromo-butyryl)-4-chlorobenzenesulfonamide, 2-bromo-N-methyl-N-phenylacetamide, 2-bromo-N-phenethyl-acetamide, 2-adamantane-1-yl-2-bromo-N-cyclohexyl-acetamide, 2-bromo-N-(2-methylphenyl)butyramide, acetyl-p-bromoaniline; (iv) Disulfides, such as: ALDRITHIOL-2, ALDRITHIOL-4, isopropyl disulfide, 1-(isobutyldithioalkyl)-2-methylpropane, dibenzyl disulfide, 4-aminophenyl disulfide, 3-(2-pyridyldithio)propionic acid, 3-(2-pyridyldithio)propionic acid hydrazide, 3-(2-pyridyldithio)propionic acid N-succinimide ester, am6amPDP1-βCD;
[0618] And (v) thiols, such as 4-phenylthiazolyl-2-thiol, Pulpald, 5,6,7,8-tetrahydro-quinazolin-2-thiol.
[0619] Mutant monomers, or one or more mutant monomers, can be chemically modified by attaching polyethylene glycol (PEG), nucleic acids such as DNA, dyes, fluorophores, or chromophores. In some embodiments, the mutant monomers, or one or more mutant monomers, are chemically modified using a molecular adaptor that promotes the interaction between the pore containing the monomer and the target analyte, target nucleotide, or target polynucleotide sequence. The presence of the adaptor improves the host-guest chemistry of the pore and the nucleotide or polynucleotide, and thereby enhances the sequencing capability of the pores formed from the mutant monomer.
[0620] The mutant monomer or one or more mutant monomers can be chemically modified at any location by attaching any molecule that has the effect of reducing or shrinking the opening diameter of the assembly pore. K37, V47, S49, T55, S86, E92, E94. More preferably, the mutant monomer can be chemically modified at positions E92 and E94 by attaching any molecule that has the effect of reducing or shrinking the opening diameter of the assembly pore. In one embodiment, the mutant monomer or one or more mutant monomers are chemically modified by attaching molecules to one or more cysteine residues (cysteine linkages) at these locations.
[0621] The reactivity of cysteine residues can be enhanced by modifying adjacent residues. For example, attaching a basic group to an arginine, histidine, or lysine residue can alter the pKa of the cysteine thiol group to make it more reactive. - The pKa of the cysteine residue can be protected by thiol protecting groups such as dTNB. These can react with one or more cysteine residues of the mutant monomer prior to attachment linker.
[0622] The molecule can be directly attached to the mutant monomer or one or more mutant monomers. Preferably, the molecule is attached to the mutant monomer using a connector such as a chemical crosslinking agent or a peptide linker. Suitable chemical crosslinking agents are well known in the art. Preferred crosslinking agents include 2,5-dioxopyrrolidine-1-yl ester of 3-(pyridin-2-yldisulfonyl)propionate, 2,5-dioxopyrrolidine-1-yl ester of 4-(pyridin-2-yldisulfonyl)butyrate, and 2,5-dioxopyrrolidine-1-yl ester of 8-(pyridin-2-yldisulfonyl)octanoate. The most preferred crosslinking agent is succinimide ester of 3-(2-pyridinedithio)propionate (SPDP). Typically, the molecule is covalently linked to the bifunctional crosslinking agent before the molecule / crosslinking agent complex is covalently linked to the mutant monomer, but it is also possible to covalently link the bifunctional crosslinking agent to the monomer before the bifunctional crosslinking agent / monomer complex is attached to the molecule.
[0623] Preferably, the joint is resistant to dithiothreitol (DTT). Suitable joints include, but are not limited to, iodoacetamide-based and maleimide-based joints.
[0624] The pores modified in this way exhibit the following specific advantages: (i) improved readhead clarity, (ii) improved base discrimination, and (iii) increased range, i.e., improved signal-to-noise ratio.
[0625] New readheads can be introduced or existing readheads can be modified by using chemical molecules to modify specific locations within the readhead. The physical size of the readhead can be significantly altered due to the size of the modified molecule. Similarly, the properties of the readhead can be changed due to the chemical properties of the modified molecule. The combination of these two effects has been shown to result in readheads with improved resolution and better base discrimination. Not only are the relative contributions of the signal to different bases at different locations altered, but readhead positions at extreme points exhibit much less discrimination, meaning their contribution to the signal is greatly reduced, and therefore the length of the K-mer measured at a given time is shorter. This clearer readhead simplifies the process of deconvolving the K-mer from the original signal.
[0626] The hole of the present invention is generated
[0627] The present invention also provides a method for generating the pores of the present invention. The method comprises: allowing at least one mutant monomer of the present invention or at least one construct of the present invention to oligomerize with a sufficient number of mutant cytolysin monomers of the present invention, constructs of the present invention, cytolysin monomers, or monomers derived from cytolysin to form pores. If the method relates to preparing homooligomeric pores of the present invention, all monomers used in the method are mutant cytolysin monomers of the present invention having the same amino acid sequence. If the method relates to preparing heterooligomeric pores of the present invention, at least one of the monomers is different from the other monomers.
[0628] Typically, monomers are expressed in host cells as described above, removed from host cells, and assembled into pores in separate membranes such as sheep erythrocyte membranes or liposomes containing sphingomyelin.
[0629] For example, cytolysin monomers can be oligomerized by adding a mixture of lipids including sphingomyelin and one or more of the following lipids and incubating the mixture at 30°C for 60 minutes, for example: phosphatidylserine; POPE; cholesterol; and Soy PC. The oligomers can be purified by any suitable method, such as by SDS-PAGE and gel purification as described in WO2013 / 153359.
[0630] Any of the embodiments discussed above with reference to the holes of the present invention are equally applicable to the method of generating holes.
[0631] Methods for characterizing analytes
[0632] This invention provides a method for characterizing a target analyte. The method includes: contacting the target analyte with a pore of the invention, causing the target analyte to move through the pore. The pore can be any of the pores discussed above. Then, as the analyte moves relative to the pore, one or more properties of the target analyte are measured using standard methods known in the art. Preferably, one or more properties of the target analyte are measured as the analyte moves through the pore. Steps (a) and (b) are preferably performed with an applied potential across the pore. As discussed in more detail below, the applied potential typically induces the formation of a complex between the pore and a polynucleotide-binding protein. The applied potential can be a voltage potential. Alternatively, the applied potential can be a chemielectric potential. An example of this operation is the use of a salt gradient across an amphiphilic layer. Salt gradients are disclosed in Holden et al., *Journal of the American Chemical Society*, July 11, 2007, Vol. 129, No. 27, pp. 8650–8655.
[0633] The method of the present invention is used to characterize a target analyte. The method is used to characterize at least one analyte. The method may involve characterizing two or more analytes. The method may include characterizing any number of analytes, such as 2, 5, 10, 15, 20, 30, 40, 50, 100, or more analytes.
[0634] The target analyte is preferably a metal ion, inorganic salt, polymer, amino acid, peptide, polypeptide, protein, nucleotide, oligonucleotide, polynucleotide, dye, bleach, drug, diagnostic agent, recreational drug, explosive, or environmental pollutant. The method may involve characterizing two or more analytes of the same type, such as two or more proteins, two or more nucleotides, or two or more drugs. Alternatively, the method may involve characterizing two or more analytes of different types, such as one or more proteins, one or more nucleotides, and one or more drugs.
[0635] The target analyte can be secreted from the cell. Alternatively, the target analyte can be an analyte present inside the cell, so that the analyte must be extracted from the cell before the present invention can be performed.
[0636] The analyte is preferably an amino acid, peptide, polypeptide, and / or protein. The amino acid, peptide, polypeptide, or protein may be naturally occurring or non-natural. The polypeptide or protein may contain synthetic or modified amino acids. Various types of modifications to amino acids are known in the art. Suitable amino acids and their modifications are described above. For the purposes of this invention, it should be understood that the target analyte can be modified by any method available in the art.
[0637] Proteins can be enzymes, antibodies, hormones, growth factors, or growth-regulating proteins, such as cytokines. Cytokines can be selected from: interleukins, preferably IFN-1, IL-1, IL-2, IL-4, IL-5, IL-6, IL-10, IL-12, and IL-13; interferons, preferably IL-γ; and other cytokines, such as TNF-α. Proteins can be bacterial proteins, fungal proteins, viral proteins, or parasite-derived proteins.
[0638] The target analyte is preferably a nucleotide, oligonucleotide, or polynucleotide. A nucleotide typically contains a nucleobase, a sugar, and at least one phosphate group. The nucleobase is usually heterocyclic. Nucleobases include, but are not limited to, purines and pyrimidines, and more specifically, adenine, guanine, thymine, uracil, and cytosine. The sugar is usually a pentose sugar. Nucleotide sugars include, but are not limited to, ribose and deoxyribose. Nucleotides are usually ribonucleotides or deoxyribonucleotides. Nucleotides typically contain monophosphate, diphosphate, or triphosphate. The phosphate may be attached to the 5' or 3' side of the nucleotide.
[0639] Nucleotides include, but are not limited to: adenosine monophosphate (AMP), adenosine diphosphate (ADP), adenosine triphosphate (ATP), guanosine monophosphate (GMP), guanosine diphosphate (GDP), guanosine triphosphate (GTP), thymidine monophosphate (TMP), thymidine diphosphate (TDP), thymidine triphosphate (TTP), uridine monophosphate (UMP), uridine diphosphate (UDP), uridine triphosphate (UTP), cytidine monophosphate (CMP), cytidine diphosphate (CDP), cytidine triphosphate (CTP), 5-methylcytidine monophosphate, 5-methylcytidine diphosphate, 5-methylcytidine triphosphate, 5-hydroxymethylcytidine monophosphate, 5-hydroxymethylcytidine diphosphate, 5-hydroxymethylcytidine triphosphate, cyclic adenosine monophosphate (cAMP), cyclic guanosine monophosphate (cGMP), deoxyadenosine monophosphate (dAMP), and deoxyadenosine diphosphate (dAMP). ADP), deoxyadenosine triphosphate (dATP), deoxyguanosine monophosphate (dGMP), deoxyguanosine diphosphate (dGDP), deoxyguanosine triphosphate (dGTP), deoxythymidine monophosphate (dTMP), deoxythymidine diphosphate (dTDP), deoxythymidine triphosphate (dTTP), deoxyuridine monophosphate (dUMP), deoxyuridine diphosphate (dUDP), deoxyuridine triphosphate (dUTP), deoxycytidine monophosphate (dCMP), deoxycytidine diphosphate (dCDP), and deoxycytidine triphosphate (dCTP), 5-methyl-2'-deoxycytidine monophosphate, 5-methyl-2'-deoxycytidine diphosphate, 5-methyl-2'-deoxycytidine triphosphate, 5-hydroxymethyl-2'-deoxycytidine monophosphate, 5-hydroxymethyl-2'-deoxycytidine diphosphate, and 5-hydroxymethyl-2'-deoxycytidine triphosphate. The nucleotide is preferably selected from AMP, TMP, GMP, UMP, dAMP, dTMP, dGMP, or dCMP. The nucleotide may be baseless (i.e., lacking a nucleobase). The nucleotide may contain additional modifications. Specifically, suitable modified nucleotides include, but are not limited to, 2'-aminopyrimidines (such as 2'-aminocytidine and 2'-aminouridine), 2'-hydroxypurines (such as 2'-fluoropyrimidines (2'-fluorocytidine and 2'-fluorouridine), hydroxypyrimidines (such as 5'-α-P-boraneuridine), 2'-O-methylnucleotides (such as 2'-O-methyladenosine, 2'-O-methylguanosine, 2'-O-methylcytidine and 2'-O-methyluridine), and 4'-thiopyrimidines (such as 4'-thiouridine and 4'-thiocytidine), and the nucleotides have modifications to the nucleobases (such as 5-pentynyl-2'-deoxyuridine, 5-(3-aminopropyl)uridine, and 1,6-diaminohexyl-N-5-carbamoylmethyluridine).
[0640] Oligonucleotides are short nucleotide polymers that typically have 50 or fewer nucleotides, such as 40 or fewer, 30 or fewer, 20 or fewer, 10 or fewer, or 5 or fewer nucleotides. Oligonucleotides can include any nucleotide discussed below, including baseless and modified nucleotides. The methods of the present invention are preferably used to characterize target polynucleotides. Polynucleotides, such as nucleic acids, are macromolecules comprising two or more nucleotides. Polynucleotides or nucleic acids can include any combination of any nucleotides. Nucleotides can be naturally occurring or artificial. One or more nucleotides in the target polynucleotide can be oxidized or methylated. One or more nucleotides in the target polynucleotide can be damaged. For example, polynucleotides can include pyrimidine dimers. Such dimers are commonly associated with UV-induced damage and are a major cause of skin melanoma. One or more nucleotides in the target polynucleotide can be modified, for example, by labeling or tagging. Suitable labeling is described below. Target polynucleotides can include one or more spacers.
[0641] The above defines nucleotides. Nucleotides present in polynucleotides include, but are not limited to: adenosine monophosphate (AMP), guanosine monophosphate (GMP), thymidine monophosphate (TMP), uridine monophosphate (UMP), cytosine monophosphate (CMP), cyclic adenosine monophosphate (cAMP), cyclic guanosine monophosphate (cGMP), deoxyadenosine monophosphate (dAMP), deoxyguanosine monophosphate (dGMP), deoxythymidine monophosphate (dTMP), deoxyuridine monophosphate (dUMP), and deoxycytidine monophosphate (dCMP). The nucleotides are preferably selected from AMP, TMP, GMP, CMP, UMP, dAMP, dTMP, dGMP, dCMP, and dUMP.
[0642] Nucleotides can be baseless (i.e., lacking nucleobases).
[0643] Nucleotides in polynucleotides can attach to each other in any way. Nucleotides are typically attached by their sugar and phosphate groups, as in nucleic acids. Nucleotides can also be linked by their nucleobases, as in pyrimidine dimers.
[0644] Polynucleotides can be single-stranded or double-stranded. At least a portion of the polynucleotide is preferably double-stranded. A single-stranded polynucleotide may have one or more primers for hybridization with it, and thus includes one or more short regions of the double-stranded polynucleotide. The primers may be the same type of polynucleotide as the target polynucleotide or may be a different type of polynucleotide.
[0645] Polynucleotides can be nucleic acids, such as deoxyribonucleic acid (DNA) or ribonucleic acid (RNA). A polynucleotide may include an RNA strand hybridized to a DNA strand. A polynucleotide can be any synthetic nucleic acid known in the field, such as peptide nucleic acid (PNA), glycerol nucleic acid (GNA), threonine nucleic acid (TNA), locked nucleic acid (LNA), or other synthetic polymers with nucleotide side chains.
[0646] This method can be used to characterize the entire target polynucleotide or only a portion thereof. The target polynucleotide can have any length. For example, the length of the polynucleotide can be at least 10, at least 50, at least 100, at least 150, at least 200, at least 250, at least 300, at least 400, or at least 500 nucleotide pairs. The length of the polynucleotide can be 1000 or more nucleotide pairs, 5000 or more nucleotide pairs, or 100,000 or more nucleotide pairs.
[0647] Target analytes such as target polynucleotides are present in any suitable sample. This invention is typically performed on samples known to contain or suspected of containing target analytes such as target polynucleotides. Alternatively, this invention can be performed on samples to confirm the identity of the presence of one or more target polynucleotides in the sample as one or more known or anticipated target analytes.
[0648] The sample can be a biological sample. The invention can be performed in vitro on samples obtained or extracted from any organism or microorganism. The organism or microorganism is typically an archaea, prokaryotic, or eukaryotic microorganism and generally belongs to one of the following five kingdoms: plant, animal, fungi, prokaryotes, and protists. The invention can be performed in vitro on samples obtained or extracted from any virus. The sample is preferably a fluid sample. The sample typically includes bodily fluids from a patient. The sample can be urine, lymph, saliva, mucus, or amniotic fluid, but is preferably blood, plasma, or serum. Typically, the sample is derived from a human, but alternatively it can be derived from another mammal, such as a commercially raised animal like a horse, cattle, sheep, or pig, or alternatively, a pet such as a cat or dog. Alternatively, plant-derived samples can typically be obtained from economic crops such as cereals, legumes, fruits, or vegetables, such as wheat, barley, oats, rapeseed, corn, soybeans, rice, rhubarb, bananas, apples, tomatoes, potatoes, grapes, tobacco, kidney beans, lentils, sugarcane, cocoa, and cotton.
[0649] The sample can be a non-biological sample. Non-biological samples are preferably fluid samples. Examples of non-biological samples include surgical fluids, water such as drinking water, seawater, or river water, and reagents used in laboratory testing.
[0650] Samples are typically processed prior to analysis, for example, by centrifugation or filtration to remove unwanted molecules or cell membranes such as red blood cells. Measurements can then be performed immediately after sample acquisition. Samples are also typically stored, preferably at below -70°C, prior to analysis.
[0651] Pores are typically present in membranes. Any membrane can be used according to the invention. Suitable membranes are well known in the art. The membrane preferably comprises sphingomyelin. The membrane is preferably an amphiphilic layer. An amphiphilic layer is a layer formed of amphiphilic molecules, such as phospholipids, having at least one hydrophilic portion and at least one lipophilic or hydrophobic portion. The amphiphilic molecules can be synthetic or naturally occurring. Non-naturally occurring amphiphiles and amphiphiles forming monolayers are known in the art and include, for example, block copolymers (Gonzalez-Perez et al., Langmuir, 2009, Vol. 25, pp. 10447-10450). A block copolymer is a polymeric material in which two or more monomer subunits are polymerized together to produce a single polymer chain. Block copolymers typically have properties contributed by each monomer subunit. However, block copolymers can have unique properties not possessed by polymers formed by individual subunits. Block copolymers can be engineered such that one of the monomer subunits is hydrophobic (i.e., lipophilic) in an aqueous medium, while one or more other subunits are hydrophilic. In this case, the block copolymer can possess amphiphilic properties and can form structures that mimic biological membranes. Block copolymers can be diblock (consisting of two monomer subunits), but can also be constructed from more than two monomer subunits to form more complex arrangements that exhibit amphiphilic behavior. Copolymers can be triblock, tetrablock, or pentablock copolymers.
[0652] The amphiphilic layer can be a monolayer or a bilayer. It is typically a planar lipid bilayer or a supporting bilayer.
[0653] Amphiphilic layers are typically lipid bilayers. Lipid bilayers serve as models of cell membranes and act as excellent platforms for a range of experimental studies. For example, lipid bilayers can be used for in vitro studies of membrane proteins via single-channel recording. Alternatively, lipid bilayers can be used as biosensors to detect the presence of a range of substances. A lipid bilayer can be any type of lipid bilayer. Suitable lipid bilayers include, but are not limited to, flat lipid bilayers, supported bilayers, or liposomes. Preferably, a flat lipid bilayer is used. Suitable lipid bilayers are disclosed in the following publications: International Application No. PCT / GB08 / 000563 (published as WO2008 / 102121), International Application No. PCT / GB08 / 004127 (published as WO 2009 / 077734), and International Application No. PCT / GB2006 / 001057 (published as WO 2006 / 100484).
[0654] Methods for forming lipid bilayers are known in the art. Suitable methods are disclosed in the examples. Lipid bilayers are typically formed by the method of Montal and Mueller (Proceedings of the National Academy of Sciences of the United States of America, 1972, Vol. 69, pp. 3561-3566), in which a lipid monolayer is carried on an aqueous / air interface through an opening perpendicular to either side of the interface.
[0655] The Montal and Mueller method is popular because it is cost-effective and a relatively straightforward method for forming good-quality lipid bilayers suitable for protein pore insertion. Other common methods for bilayer formation include tip immersion, bilayer brushing, and patch clamping.
[0656] In a preferred embodiment, a lipid bilayer is formed as disclosed in International Application No. PCT / GB08 / 004127 (published as WO 2009 / 077734). In another preferred embodiment, the membrane is a solid layer. The solid layer is not of biological origin. In other words, the solid layer is not derived from or isolated from a biological environment, such as an organism or cell, or a synthetically manufactured version of a biologically available structure. The solid layer can be formed from both organic and inorganic materials, including but not limited to: microelectronic materials; insulating materials such as Si3N4, Al2O3, and SiO; organic and inorganic polymers such as polyamides; and plastics such as... Alternatively, it can be an elastomer, such as a two-component addition-cured silicone rubber; and glass. The solid layer can be formed from a single atomic layer, such as graphene, or a layer only a few atoms thick. Suitable graphene layers are disclosed in International Application No. PCT / US2008 / 010637 (published as WO 2009 / 035647).
[0657] The methods are typically performed using: (i) an artificial amphiphilic layer comprising pores; (ii) a separated, naturally occurring lipid bilayer comprising pores; or (iii) a cell having pores inserted therein. The methods are typically performed using artificial amphiphilic layers such as artificial lipid bilayers. The layers may include other transmembrane and / or intramembrane proteins and other molecules besides pores. Suitable equipment and conditions are discussed below. The methods of the invention are typically performed in vitro.
[0658] Analytes such as target polynucleotides can be coupled to the membrane. This can be accomplished using any known method. If the membrane is an amphiphilic layer such as a lipid bilayer (as discussed in detail above), the analyte, such as the target polynucleotide, is preferably coupled to the membrane via a polypeptide present in the membrane or a hydrophobic anchor present in the membrane. The hydrophobic anchor is preferably a lipid, fatty acid, sterol, carbon nanotube, or amino acid.
[0659] Analytes such as target polynucleotides can be directly coupled to the membrane. Alternatively, analytes such as target polynucleotides are preferably coupled to the membrane via linkers. Preferred linkers include, but are not limited to, polymers such as polynucleotides, polyethylene glycol (PEG), and peptides. If the polynucleotide is directly coupled to the membrane, some data will be lost because characterization cannot continue to the end of the polynucleotide due to the distance between the membrane and the pore interior. Using linkers allows for complete processing of the polynucleotide. Linkers can be attached to the polynucleotide at any location. Preferably, the linker attaches to the polynucleotide at the tail polymer.
[0660] Coupling can be stable or transient. For some applications, the transient nature of the coupling is preferred. If a stable coupling molecule is directly attached to the 5' or 3' end of a polynucleotide, some data will be lost because the characterization run cannot continue to the end of the polynucleotide due to the distance between the bilayer and the pore interior. If the coupling is transient, the polynucleotide can be fully processed when the coupling ends randomly become free of the bilayer. The chemical groups that form stable or transient connections with the membrane are discussed in more detail below. Analytes such as target polynucleotides can be transiently coupled to amphiphilic layers such as lipid bilayers using cholesterol or fatty acyl chains. Any fatty acyl chain with a length of 6 to 30 carbon atoms, such as hexadecanoic acid, can be used.
[0661] In a preferred embodiment, an analyte, such as a target polynucleotide, is coupled to the amphiphilic layer. Various tethering strategies have previously been used to perform the coupling of analytes, such as target polynucleotides, to synthetic lipid bilayers. This information is summarized in Table 5 below:
[0662] Table 5
[0663]
[0664] Polynucleotides can be functionalized using modified phosphorous amides in synthetic reactions. These modified phosphorous amides are readily compatible with the addition of reactive groups such as thiols, cholesterol, lipids, and biotin groups. These different attachment chemistry methods provide a range of attachment options for polynucleotides. Each different modifying group tethers the polynucleotide in a slightly different manner, and the coupling is not always permanent, thus allowing for varying residence times for the polynucleotide to couple to the bilayer. The advantages of transient coupling have been discussed above.
[0665] Polynucleotide coupling can also be achieved through a variety of other means, provided that a receptive group is added to the polynucleotide. Adding reactive groups to either end of DNA has been previously reported. A thiol group can be added to the 5' end of ssDNA using polynucleotide kinases and ATPγS (Grant, GP and PZQin (2007), “A facile method for attaching nitroxide spinlabels at the 5' terminus of nucleic acids,” Nucleic Acids Res, Vol. 10, No. 35, p. e77). Terminal deoxynucleotidyltransferases can be used to incorporate modified oligonucleotides into the 3' end of ssDNA to add a more diverse set of chemical groups, such as biotin, thiols, and fluorophores (Kumar, A., P. Tchen et al. (1988), “Nonradioactive labelling of synthetic oligonucleotide probes with terminal deoxynucleotidyltransferase”). Analytical Biochemistry (Volume 169, Issue 2, pp. 376-382).
[0666] Alternatively, the reactive group can be viewed as an addition of a short piece of DNA complementary to the DNA already coupled to the bilayer, allowing attachment to be achieved via hybridization. Ligation of short ssDNA fragments using T4 RNA ligase I has been reported (Troutt, AB, MGMcHeyzer-Williams et al. (1992), “Ligation-anchored PCR: a simple amplification technique with single-sided specificity,” Proceedings of the National Academy of Sciences of the United States of America, Vol. 89, No. 20, pp. 9823-9825). Alternatively, ssDNA or dsDNA can be ligated to native dsDNA, and the two strands can then be separated by heat or chemical denaturation. For native dsDNA, a piece of ssDNA can be added to one or both ends of the bilayer, or dsDNA can be added to one or both ends. Then, when the duplex melts, if ssDNA is used for ligation or modification at the 5' and 3' ends, each single strand will have either a 5' or 3' modification; or if dsDNA is used for ligation, each single strand will have both 5' and 3' modifications. If the polynucleotide is a synthetic strand, coupling chemistry can be incorporated during the chemical synthesis of the polynucleotide. For example, primers with reactive groups attached to them can be used to synthesize polynucleotides.
[0667] A common technique for amplifying genomic DNA segments is the use of polymerase chain reaction (PCR). Here, using two synthetic oligonucleotide primers, numerous copies of the same DNA segment can be produced, where for each copy, the 5' end of each strand in the duplex will be a synthetic polynucleotide. By using antisense primers with reactive groups such as cholesterol, thiols, biotin, or lipids, each copy of the amplified target DNA will contain a reactive group for coupling.
[0668] The pores used in the methods of this invention are the pores of this invention (i.e., pores comprising at least one mutant monomer of this invention or at least one construct of this invention). The pores can be chemically modified in any of the ways discussed above. Preferably, the pores are modified using covalent linkers capable of interacting with the target analyte, as discussed above.
[0669] The method is preferably used to characterize a target polynucleotide, and step (a) comprises: contacting the target polynucleotide with the pore and a polynucleotide-binding protein, wherein the polynucleotide-binding protein controls the movement of the target polynucleotide through the pore. The polynucleotide-binding protein can be any protein capable of binding to a polynucleotide and controlling its movement through the pore. In the art, determining whether a polynucleotide-binding protein binds to a polynucleotide is straightforward. Polynucleotide-binding proteins typically interact with polynucleotides and modify at least one property of the polynucleotide. Polynucleotide-binding proteins can modify polynucleotides by cleaving them to form individual nucleotides or shorter nucleotide chains such as dinucleotides or trinucleotides. The portion can modify the polynucleotide by orienting or moving it to a specific location, i.e., controlling its movement.
[0670] Polynucleotide-binding proteins are preferably polynucleotide-processing enzymes. Polynucleotide-processing enzymes are polypeptides capable of interacting with polynucleotides and modifying at least one of their properties. The enzyme can modify polynucleotides by cleaving them to form individual nucleotides or shorter nucleotide chains such as dinucleotides or trinucleotides. The enzyme can also modify polynucleotides by orienting or moving them to a specific location. Polynucleotide-binding proteins typically include a polynucleotide-binding domain and a catalytic domain. Polynucleotide-processing enzymes do not need to exhibit enzymatic activity, as long as they can bind to the target sequence and control its movement through the pore. For example, the enzyme can be modified to remove its enzymatic activity, or it can be used under conditions that prevent it from acting as an enzyme. Such conditions are discussed in more detail below.
[0671] The polynucleotide processing enzyme is preferably derived from an autolysin. More preferably, the polynucleotide processing enzyme used in the enzyme construct is derived from a member of any of the following enzyme classification (EC) groups: 3.1.11, 3.1.13, 3.1.14, 3.1.15, 3.1.16, 3.1.21, 3.1.22, 3.1.25, 3.1.26, 3.1.27, 3.1.30, and 3.1.31. The enzyme may be any enzyme disclosed in International Application No. PCT / GB10 / 000133 (published as WO 2010 / 086603).
[0672] Preferred enzymes are polymerases, exonucleases, helicases, and topoisomerases, such as gyrases. Suitable enzymes include, but are not limited to, exonuclease I (SEQ ID NO: 6) from *E. coli*, exonuclease III (SEQ ID NO: 8) from *E. coli*, RecJ (SEQ ID NO: 10) from thermophilic bacteria, and bacteriophage λ exonuclease (SEQ ID NO: 12) and its variants. The three subunits comprising the sequence shown in SEQ ID NO: 10, or variants thereof, interact to form a trimer exonuclease. The enzyme may be Phi29 DNA polymerase (SEQ ID NO: 4) or a variant thereof. The enzyme may be a helicase or derived from a helicase. Typical helicases are Hel308, RecD, or XPD, for example, Hel308 Mbu (SEQ ID NO: 13) or a variant thereof.
[0673] The enzyme is preferably derived from helicases, such as Hel308 helicase, TraI helicase or TrwC helicase, RecD helicase, XPD helicase or Dda helicase. The helicase may be any of the helicases, modified helicases, or helicase constructs disclosed in the following international applications: PCT / GB2012 / 052579 (published as WO 2013 / 057495); PCT / GB2012 / 053274 (published as WO 2013 / 098562); PCT / GB2012 / 053273 (published as WO2013098561); PCT / GB2013 / 051925 (published as WO 2014 / 013260); PCT / GB2013 / 051924 (published as WO 2014 / 013259); PCT / GB2013 / 051928 (published as WO 2014 / 013262); and PCT / GB2014 / 052736.
[0674] The helicase preferably comprises the sequence shown in SEQ ID NO: 18 (Dda) or a variant thereof. The variant may differ from the native sequence in any of the ways discussed below with respect to transmembrane pores. Preferred variants of SEQ ID NO: 18 include: (a) E94C and A360C; or (b) E94C, A360C, C109A, and C136A, and then optionally (ΔM1)G1G2 (i.e., M1 is deleted and then G1 and G2 are added).
[0675] Variants of SEQ ID NO: 4, 6, 8, 10, 12, 13, or 18 are enzymes whose amino acid sequences differ from those of SEQ ID NO: 4, 6, 8, 10, 12, 13, or 18 and which retain polynucleotide binding capacity. Variants may contain modifications that promote polynucleotide binding and / or promote the activity of polynucleotides at high salt concentrations and / or at room temperature.
[0676] Within the entire length of the amino acid sequence of SEQ ID NO: 4, 6, 8, 10, 12, 13, or 18, the variant will preferably be at least 50% homologous to the sequence based on amino acid identity. More preferably, based on amino acid identity, the variant polypeptide may be at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, and more preferably at least 95%, 97%, or 99% homologous to the amino acid sequence of SEQ ID NO: 4, 6, 8, 10, 12, 13, or 18 throughout the entire sequence. Within an extension of 200 or more, such as 230, 250, 270, or 280 or more consecutive amino acids, there may be at least 80%, such as at least 85%, 90%, or 95% amino acid identity (“hard homology”). Homology is determined as described above. Variants may differ from the wild-type sequence in any of the ways discussed above with reference to SEQ ID NO: 2. The enzyme may be covalently linked to the well, as discussed above.
[0677] There are two main strategies for sequencing polynucleotides using nanopores: strand sequencing and exonuclease sequencing. The method of this invention can involve either strand sequencing or exonuclease sequencing.
[0678] In strand sequencing, DNA is transferred through nanopores by applying or resisting an applied potential. Exonucleases that act progressively or gradually on double-stranded DNA can be used on the cis side of the pore to feed the remaining single strand at the applied potential or on the trans side at the opposite potential. Similarly, helicases that unwind double-stranded DNA can be used in a similar manner. Polymerases can also be used. There are also sequencing applications that require strand transfer against an applied potential, but the DNA must first be “captured” by the enzyme at the opposite or no potential. With the potential then switched back after binding, the strand will pass through the pore in a cis-to-trans configuration and remain in an elongated conformation by current. Single-stranded DNA exonucleases or single-stranded DNA-dependent polymerases can act as molecular motors to pull the recently transferred single strand back through the pore in a stepwise, controlled manner from trans to cis against the applied potential.
[0679] In one embodiment, a method for characterizing a target polynucleotide involves contacting the target sequence with a pore and a helicase. Any helicase can be used in this method. The helicase can function relative to the pore in two modes. First, the method is preferably performed using a helicase such that the helicase controls the movement of the target sequence through the pore using a field generated by an applied voltage. In this mode, the 5' end of the DNA is first captured in the pore, and the enzyme controls the movement of the DNA into the pore, causing the target sequence to be passed through the pore using the field until it is finally transferred through to the opposite side of the bilayer. Alternatively, the method is preferably performed such that the helicase controls the movement of the target sequence through the pore against a field generated by an applied voltage. In this mode, the 3' end of the DNA is first captured in the pore, and the enzyme controls the movement of the DNA through the pore, causing the target sequence to be pulled out of the pore against the applied field until it is finally pushed back to the cis side of the bilayer.
[0680] In exonuclease sequencing, an exonuclease releases individual nucleotides from one end of a target polynucleotide, and these individual nucleotides are recognized as discussed below. In another embodiment, the method for characterizing the target polynucleotide involves contacting the target sequence with a well and an exonuclease. Any of the exonucleases discussed above can be used in this method. The enzyme can be covalently linked to the well, as discussed above.
[0681] Exonucleases are enzymes that typically latch onto one end of a polynucleotide and digest the sequence one nucleotide at a time from that end. Exonucleases can digest polynucleotides in a 5' to 3' orientation or a 3' to 5' orientation. The polynucleotide end to which the exonuclease binds is typically determined by selecting enzymes used in the field and / or by methods known in the field. A hydroxyl group or cap structure at either end of the polynucleotide can often be used to prevent or facilitate exonuclease binding to a specific end of the polynucleotide.
[0682] The method involves contacting a polynucleotide with an exonuclease, such that the nucleotide is digested from the end of the polynucleotide at a rate that allows for the characterization or recognition of a certain proportion of the nucleotides, as discussed above. Methods for performing this operation are well known in the art. For example, Edman degradation is used for the sequential digestion of single amino acids from the ends of a polypeptide, allowing the amino acid to be identified using high-performance liquid chromatography (HPLC). Homology methods can be used in this invention.
[0683] Exonucleases typically function at a slower rate than the optimal rate of wild-type exonucleases. Suitable rates of exonuclease activity in the methods of this invention include the following digestion rates: 0.5 to 1000 nucleotides per second, 0.6 to 500 nucleotides per second, 0.7 to 200 nucleotides per second, 0.8 to 100 nucleotides per second, 0.9 to 50 nucleotides per second, or 1 to 20 or 10 nucleotides per second. These rates are preferably 1, 10, 100, 500, or 1000 nucleotides per second. Suitable rates of exonuclease activity can be achieved in various ways. For example, a variant exonuclease with a reduced optimal activity rate can be used according to the invention.
[0684] The method of the present invention relates to measuring one or more properties of a target analyte, such as a target polynucleotide. The method may involve measuring two, three, four, five, or more properties of the target analyte. For the target polynucleotide, the one or more properties are preferably selected from: (i) the length of the target polynucleotide; (ii) the identity of the target polynucleotide; (iii) the sequence of the target polynucleotide; (iv) the secondary structure of the target polynucleotide; and (v) whether the target polynucleotide is modified. Any combination of (i) to (v) can be measured according to the present invention.
[0685] For (i), the length of the polynucleotide can be measured using the number of interactions between the target polynucleotide and the pore.
[0686] Regarding (ii), polynucleotide identity can be measured in several ways. Polynucleotide identity can be measured either by combining the measurement of the target polynucleotide's sequence with or without measuring the target polynucleotide's sequence. The former is straightforward; the polynucleotide is sequenced and thus identified. The latter can be performed in several ways. For example, the presence of a specific motif in the polynucleotide can be measured (without measuring the remaining sequence of the polynucleotide). Alternatively, the measurement of specific electrical and / or optical signals in the method can identify the target polynucleotide as originating from a specific source.
[0687] For (iii), the sequence of the polynucleotide can be determined as previously described. Suitable sequencing methods, especially those using electrical measurements, are described in Stoddart D et al., Proceedings of the National Academy of Sciences, 2012, Vol. 106, No. 19, pp. 7702-7707; Lieberman KR et al., Journal of the American Chemical Society, 2010, Vol. 132, No. 50, pp. 17961-17972; and international application WO 2000 / 28312.
[0688] For (iv), secondary structure can be measured in a variety of ways. For example, if the method involves electrical measurements, changes in residence time or changes in the current flowing through the pore can be used to measure secondary structure. This allows for the differentiation of regions of single-stranded and double-stranded polynucleotides.
[0689] For (v), the presence of any modification can be measured. The method preferably includes determining whether the target polynucleotide has been modified by methylation, oxidation, or damage using one or more proteins or by one or more markers, tags, or spacers. Specific modifications will induce specific interactions with the pore, which can be measured using the methods described below. For example, methylcytosine and cytosine can be distinguished based on the current flowing through the pore during the interaction of the pore with each nucleotide.
[0690] This invention also provides a method for estimating the sequence of a target polynucleotide. This invention further provides a method for sequencing a target polynucleotide.
[0691] Various types of measurements can be performed. This includes, but is not limited to, electrical measurements and optical measurements. Possible electrical measurements include: current measurements, impedance measurements, tunneling measurements (Ivanov AP et al., Nano Letters, January 12, 2011, Vol. 11, No. 1, pp. 279-285), and FET measurements (International...).
[0692] (WO 2005 / 124888). Appropriate optical methods involving fluorescence measurements are disclosed in the *Journal of the American Chemical Society*, 2009, Vol. 131, pp. 1652 and 1653. Optical measurements can be combined with point measurements (Soni GV et al., *Rev Sci Instrum*, January 2010, Vol. 81, No. 1, p. 014301). Measurements can be transmembrane current measurements, such as measurements of ion currents flowing through pores.
[0693] Electrical measurements can be performed using standard single-channel recording equipment as described in the following documents:
[0694] Stoddart D et al., *Proceedings of the National Academy of Sciences of the United States of America (Proc Natl Acad Sci)*, 2012, Vol. 106, No. 19, pp. 7702-7707; Lieberman KR et al., *Journal of the American Chemical Society (J Am Chem Soc)*, 2010, Vol. 132, No. 50, pp. 17961-17972; and international applications.
[0695] WO-2000 / 28312. Alternatively, a multi-channel system, such as those described in the following documents, can be used for electrical measurements:
[0696] International application WO-2009 / 077734 and international application WO-2011 / 067559.
[0697] In a preferred embodiment, the method includes:
[0698] (a) Contacting a target polynucleotide with the pore and polynucleotide-binding protein of the present invention, such that the target polynucleotide moves through the pore, and the binding protein controls the movement of the target polynucleotide through the pore; and
[0699] (b) Measure the current passing through the pore as the polynucleotide moves relative to the pore, wherein the current indicates one or more properties of the target polynucleotide and thereby characterizes the target polynucleotide.
[0700] The method can be performed using any device suitable for studying membrane / pore systems with pores inserted into the membrane. The method can also be performed using any device suitable for transmembrane pore sensing. For example, the device includes a chamber containing an aqueous solution and a barrier dividing the chamber into two sections. The barrier has openings in which a pore-containing membrane is formed.
[0701] The method can be performed using the apparatus described in International Application No. PCT / GB08 / 000562 (WO 2008 / 102120).
[0702] The method may involve measuring the current flowing through a pore as an analyte, such as a target polynucleotide, moves relative to the pore. Therefore, the device may also include circuitry capable of applying a potential and measuring electrical signals across the membrane and pore. The method can be performed using patch clamps or voltage clamps. The method preferably involves the use of voltage clamps.
[0703] The method of the present invention may involve measuring the current passing through a pore as an analyte, such as a target polynucleotide, moves relative to the pore. Suitable conditions for measuring the ionic current passing through a transmembrane protein pore are known in the art and disclosed in examples. The method is typically performed using a voltage applied across the membrane and to the pore. The voltage used is typically from +2V to -2V, and typically from -400mV to +400mV. The voltage used is preferably within a range having a lower limit selected from -400mV, -300mV, -200mV, -150mV, -100mV, -50mV, -20mV, and 0mV, and the upper limit is independently selected from +10mV, +20mV, +50mV, +100mV, +150mV, +200mV, +300mV, and +400mV. The voltage used is more preferably in the range of 100mV to 240mV, and most preferably in the range of 120mV to 220mV. The ability to distinguish different nucleotides through the pore can be increased by using a higher applied potential.
[0704] The method is typically performed in the presence of any charge carriers, such as metal salts, e.g., alkali metal salts; or halide salts, e.g., chloride salts such as alkali metal chloride salts. The charge carriers may comprise ionic liquids or organic salts, such as tetramethylammonium chloride, trimethylphenylammonium chloride, phenyltrimethylammonium chloride, or 1-ethyl-3-methylimidazolium chloride. In the exemplary apparatus discussed above, the salt is present in an aqueous solution within the chamber. Potassium chloride (KCl), sodium chloride (NaCl), or cesium chloride (CsCl) is typically used. KCl is preferred. The salt concentration may be saturated. The salt concentration may be 3 M or lower, and is typically 0.1 M to 2.5 M, 0.3 M to 1.9 M, 0.5 M to 1.8 M, 0.7 M to 1.7 M, 0.9 M to 1.6 M, or 1 M to 1.4 M. The salt concentration is preferably 150 mM to 1 M. The method is preferably performed using a salt concentration of at least 0.3 M, such as at least 0.4 M, at least 0.5 M, at least 0.6 M, at least 0.8 M, at least 1.0 M, at least 1.5 M, at least 2.0 M, at least 2.5 M, or at least 3.0 M. High salt concentrations provide a high signal-to-noise ratio and allow identification of currents indicating the presence of nucleotides against a background of normal current fluctuations.
[0705] The method is typically performed in the presence of a buffer solution. In the exemplary apparatus discussed above, the buffer solution is present in an aqueous solution within the chamber. Any buffer solution can be used in the method of the present invention. Typically, the buffer solution is HEPES. Another suitable buffer solution is Tris-HCl buffer. The method is typically performed at the following pH values: 4.0 to 12.0, 4.5 to 10.0, 5.0 to 9.0, 5.5 to 8.8, 6.0 to 8.7, or 7.0 to 8.8, or 7.5 to 8.5. The pH used is preferably about 7.5.
[0706] The method can be performed at the following temperatures: 0°C to 100°C, 15°C to 95°C, 16°C to 90°C, 17°C to 85°C, 18°C to 80°C, 19°C to 70°C, or 20°C to 60°C. The method is typically performed at room temperature. Optionally, the method can be performed at a temperature that supports enzyme function, such as about 37°C.
[0707] The method is typically performed in the presence of free nucleotides or free nucleotide analogs and enzyme cofactors that promote the action of polynucleotide-binding proteins such as helicases or exonucleases. Free nucleotides can be one or more of any of the individual nucleotides discussed above. Free nucleotides include, but are not limited to: adenosine monophosphate (AMP), adenosine diphosphate (ADP), adenosine triphosphate (ATP), guanosine monophosphate (GMP), guanosine diphosphate (GDP), guanosine triphosphate (GTP), thymidine monophosphate (TMP), thymidine diphosphate (TDP), thymidine triphosphate (TTP), uridine monophosphate (UMP), uridine diphosphate (UDP), uridine triphosphate (UTP), cytidine monophosphate (CMP), cytidine diphosphate (CDP), cytidine triphosphate (CTP), cyclic adenosine monophosphate (cAMP), cyclic guanosine monophosphate (cGMP), and monophosphate. The free nucleotides are preferably selected from AMP, TMP, GMP, CMP, UMP, dAMP, dTMP, dGMP, and dCMP. The free nucleotide is preferably adenosine triphosphate (ATP). The enzyme cofactor is a factor that allows the helicase to function. The enzyme cofactor is preferably a divalent metal cation. The divalent metal cation is preferably Mg. 2+ Mn2+ Ca 2+ or Co 2+ The most preferred enzyme cofactor is Mg. 2+ .
[0708] The target polynucleotide can be contacted with the pore and the polynucleotide-binding protein in any order. Preferably, when contacting the target polynucleotide with the polynucleotide-binding protein and the pore, the target polynucleotide first forms a complex with the polynucleotide-binding protein. When a voltage is applied across the pore, the target polynucleotide / protein complex forms a complex with the pore and controls the movement of the polynucleotide through the pore.
[0709] Methods for identifying individual nucleotides
[0710] The present invention also provides a method for characterizing individual nucleotides. In other words, the target analyte is an individual nucleotide. The method includes: contacting the nucleotide with a pore of the present invention, such that the nucleotide interacts with the pore; and measuring the current passing through the pore during the interaction and thereby characterizing the nucleotide. Therefore, the present invention relates to nanopore sensing of individual nucleotides. The present invention also provides a method for identifying individual nucleotides, comprising: measuring the current passing through the pore during the interaction and thereby determining the identity of the nucleotide. Any of the pores discussed above can be used. Preferably, the pores are chemically modified using molecular adapters, as discussed above.
[0711] If the current flows through the pore in a manner specific to nucleotides (i.e., if a unique current associated with the analyte is detected flowing through the pore), then nucleotides are present. If the current does not flow through the pore in a manner specific to nucleotides, then nucleotides are not present.
[0712] This invention can be used to distinguish nucleotides based on the different effects of nucleotides with the same structure on the current passing through a pore. Individual nucleotides can be identified at the single-molecule level based on the amplitude of the nucleotide current when the nucleotide interacts with the pore. This invention can also be used to determine the presence of specific nucleotides in a sample. Furthermore, this invention can be used to measure the concentration of specific nucleotides in a sample.
[0713] Pores are typically present within the membrane. The method can be performed using any of the suitable membrane / pore systems described above.
[0714] A single nucleotide is a single nucleotide. A single nucleotide is a nucleotide that is not bound to another nucleotide or polynucleotide by a nucleotide bond. A nucleotide bond involves the binding of one of the phosphate groups of a nucleotide to the sugar group of another nucleotide. A single nucleotide is typically a nucleotide that is bound by a nucleotide bond to another polynucleotide consisting of at least 5, at least 10, at least 20, at least 50, at least 100, at least 200, at least 500, at least 1000, or at least 5000 nucleotides. For example, single nucleotides have been digested from polynucleotide sequences of target analytes such as DNA or RNA chains. The method of the present invention can be used to identify any nucleotide. A nucleotide can be any nucleotide discussed above.
[0715] Nucleotides can be derived from the digestion of nucleic acid sequences such as ribonucleic acid (RNA) or deoxyribonucleic acid (DNA). Nucleic acid sequences can be digested using any method known in the field. Suitable methods include, but are not limited to, methods using enzymes or catalysts. Catalytic digestion of nucleic acids is disclosed in Deck et al., Inorganic Chemistry, 2002, Vol. 41, pp. 669–677.
[0716] Individual nucleotides from a single polynucleotide can be sequentially contacted with the pore to sequence all or part of the polynucleotide. Sequencing of polynucleotides has been discussed in more detail above.
[0717] Nucleotides can be made to contact the pores on both sides of the membrane. Nucleotides can be introduced into the pores from both sides of the membrane. Nucleotides can be made to contact the sides of the membrane, allowing the nucleotide to pass through the pore to the other side of the membrane. For example, the nucleotide can be made to contact the end of the pore, which allows ions or small molecules such as nucleotides to enter the barrel or channel of the pore in their native environment, allowing the nucleotide to pass through the pore. In this case, the nucleotide interacts with the pore and / or the adaptor as it crosses the pore barrel or channel. Alternatively, the nucleotide can contact the sides of the membrane, which allow the nucleotide to interact with the pore through or bind to the adaptor, dissociate from the pore, and remain on the same side of the membrane. The present invention provides pores in which the positions of the adaptors are fixed. Therefore, the nucleotide is preferably contacted with the end of the pore, which allows the adaptor to interact with the nucleotide.
[0718] Nucleotides can interact with pores in any manner and at any site. As discussed above, nucleotides preferably bind reversibly to pores via or by binding to an adaptor. Most preferably, nucleotides bind reversibly to pores via or by binding to an adaptor as they cross the membrane. Nucleotides can also reversibly bind to the barrel or channel of the pore via or by binding to an adaptor as they cross the membrane.
[0719] During the interaction between the nucleotide and the pore, the nucleotide affects the current flowing through the pore in a manner specific to that nucleotide. For example, a particular nucleotide will reduce the current flowing through the pore to a specific extent over a specific average time period. In other words, the current flowing through the pore is unique to a particular nucleotide. Control experiments can be performed to determine the effect of a particular nucleotide on the current flowing through the pore. The results obtained by performing the method of the present invention on a test sample can then be compared with the results derived from such control experiments to identify a particular nucleotide in the sample or to determine the presence of a particular nucleotide in the sample. The frequency with which the current flowing through the pore is affected in a manner indicative of a particular nucleotide can be used to determine the concentration of that nucleotide in the sample. The ratio of different nucleotides within the sample can also be calculated. For example, the ratio of dCMP to methyl-dCMP can be calculated.
[0720] The method may involve using any of the equipment, samples, or conditions discussed above.
[0721] Methods for forming sensors
[0722] The present invention also provides a method for forming a sensor for characterizing a target polynucleotide. The method includes forming a complex between a pore of the invention and a polynucleotide-binding protein, such as a helicase or exonuclease. The complex can be formed by contacting the pore and the protein in the presence of the target polynucleotide and then applying a potential across the pore. The applied potential can be a chemielectric potential or a voltage potential as described above. Alternatively, the complex can be formed by covalently linking the pore to the protein. Methods for covalent linking are known in the art and are disclosed, for example, in International Applications PCT / GB09 / 001679 (published as WO 2010 / 004265) and PCT / GB10 / 000133 (published as WO 2010 / 086603). The complex is a sensor for characterizing a target polynucleotide. The method preferably includes forming a complex between a pore of the invention and a helicase. Any of the embodiments discussed above are equally applicable to this method.
[0723] The present invention also provides a sensor for characterizing target polynucleotides. The sensor comprises a complex between the pores of the present invention and a polynucleotide-binding protein. Any embodiments discussed above are equally applicable to the sensor of the present invention.
[0724] Reagent test kit
[0725] The present invention also provides a kit for characterizing target polynucleotides, such as sequencing them. The kit includes (a) the wells of the present invention and (b) a membrane. The kit preferably further includes polynucleotide-binding proteins such as helicases or exonucleases. Any of the embodiments discussed above are also applicable to the kit of the present invention.
[0726] The kit of the present invention may additionally include one or more other reagents or instruments that enable any of the embodiments mentioned above to be performed. Such reagents or instruments include one or more of the following: one or more suitable buffer solutions (aqueous solutions), devices for obtaining samples from a subject (such as containers or instruments including needles), devices for amplifying and / or expressing polynucleotide sequences, membranes, voltage clamps, or patch clamps as defined above. The reagents may be present in the kit in a dry state, such that a fluid sample resuspends the reagents. The kit may also optionally include instructions for use of the methods of the present invention or details regarding which patients the methods may be used for. The kit may optionally include nucleotides.
[0727] equipment
[0728] The present invention also provides an apparatus for characterizing target polynucleotides in a sample, such as sequencing them. The apparatus may include (a) a plurality of wells of the present invention and (b) a plurality of polynucleotide-binding proteins, such as helicases or exonucleases. The apparatus may be any conventional apparatus for analyte analysis, such as an array or a chip.
[0729] The array or chip typically contains multiple wells in a membrane, such as a block copolymer membrane, each well having a single nanopore inserted within it. The array can be integrated within an electronic chip.
[0730] The device preferably includes:
[0731] A sensor device that can support multiple pores and operate to use pores and proteins to perform polynucleotide characterization or sequencing.
[0732] - At least one reservoir for holding materials to be used for characterization or sequencing;
[0733] - A fluid dynamics system configured to controllably supply material from the at least one storage tank to the sensor device; and
[0734] - Multiple containers for holding corresponding samples, the fluid dynamics system being configured to selectively supply the samples from the containers to the sensor device.
[0735] The device may be any of the devices described in International Application No. PCT / GB10 / 000789 (published as WO 2010 / 122293), International Application No. PCT / GB10 / 002206 (published as WO 2011 / 067559) or International Application No. PCT / US99 / 25679 (published as WO 00 / 28312).
[0736] The following examples illustrate the present invention.
[0737] Example 1
[0738] This example describes how helicase-T4Dda-E94C / C109A / C136A / A360C (SEQ ID NO: 18 with mutant E94C / C109A / C136A / A360C) was used to control DNA migration through multiple different mutant cytolysin nanopores. All tested nanopores exhibited changes in current during DNA transfer through the nanopores. The tested mutant nanopores exhibited: 1) increased range; 2) reduced noise; 3) improved signal-to-noise ratio; 4) increased capture compared to mutant control nanopores; or 5) altered readhead size compared to baseline.
[0739] Materials and methods
[0740] DNA construct preparation
[0741] Exchange 70 μL of T4 Dda-E94C / C109A / C136A / A360C buffer (using a Zeba column) into 70 μL of 1x KOAc buffer containing 2 mM EDTA.
[0742] • Add 70 μL of the T4 Dda-E94C / C109A / C136A / A360C buffer exchange mixture to 70 μL of 2 μM DNA adaptor (for sequence details, please see [link to relevant documentation]). Figure 5 The samples were then mixed and incubated at room temperature for 5 minutes.
[0743] Add 1 μL of 140 mM TMAD and mix the sample, then incubate at room temperature for 60 minutes. This sample is referred to as Sample A. Then, take 2 μL of the sample aliquot for analysis by Agilent.
[0744] HS / ATP steps
[0745] Mix the reagents listed in the table below and incubate at room temperature for 25 minutes. This sample is referred to as Sample B.
[0746] reagents volume final Sample A (500 nM) 139 220nM 2x Hs buffer (100mM Hepes, 2M KCl, pH 8) 150 1x <![CDATA[600mM MgCl2]]> 7 14mM 100mM rATP 4.2 14mM final 300.2
[0747] SPRI purification
[0748] Add 1.1 mL of SPRI beads to sample B, then mix the sample and incubate for 5 minutes.
[0749] • Precipitate the beads and remove the supernatant. Then wash the beads with 50 mM Tris.HCl, 2.5 M NaCl, and 20% PEG8000.
[0750] • Elute sample C in 70 μL of 10 mM Tris.HCl and 20 mM NaCl.
[0751] 10 kbλC was ligated to the adaptor using an enzyme.
[0752] • Incubate the reagents in the table below at 20°C for 10 minutes in a thermal cycler.
[0753]
[0754]
[0755] The reaction mixture (1 x 500 μl aliquots) was then treated as follows: SPRI purification was performed using 200 μl of 20% SPRI beads; washing was performed in 750 μl of wash buffer 1; and elution was performed in 125 μl of elution buffer 1. The final DNA sequence (SEQ ID NO: 24) was hybridized to DNA. This sample was designated Sample D.
[0756] Components of the ligation buffer (5x)
[0757] reagents volume final 1M Tris.HCl pH8 15 150mM <![CDATA[1M MgCl2]]> 5 50mM 100mM ATP 5 5mM 40% PEG 8000 75 30% total 100uL
[0758] Components of Wash Buffer 1
[0759] reagents volume final water 1100 1M Tris.HCl pH8 100 50mM 5M NaCl 300 750mM 40% PEG 8000 500 10% total 2000uL
[0760] Components of elution buffer 1
[0761] reagents volume final water 906.7 Up to 1000uL 0.5M CAPS pH10 80 40mM 3M KCl 13.3 40mM total 1000uL
[0762] Electrophysiological experiments
[0763] Electrometry was obtained from single cytolysin nanopores inserted into a block copolymer in buffer (25 mM potassium phosphate buffer, 150 mM potassium ferrocyanide (II), 150 mM potassium ferricyanide (III), pH 8.0). After reaching the single pores inserted into the block copolymer, buffer (2 mL, 25 mM potassium phosphate buffer, 150 mM potassium ferrocyanide (II), 150 mM potassium ferricyanide (III), pH 8.0) was then passed through the system to remove any excess cytolysin nanopores. Then, 150 μL of 500 mM KCl, 25 mM potassium phosphate, pH 8.0 was passed through the system. After 10 minutes, 150 μL of 500 mM KCl, 25 mM potassium phosphate, pH 8.0 was flowed through the system, and then a premixture of T4 Dda-E94C / C109A / C136A / A360C, DNA, and fuel (MgCl2, ATP) (total 150 μL, sample D) was flowed into the single nanopore experimental system. Experiments were performed at 180 mV, and helicase-controlled DNA migration was monitored.
[0764] result
[0765] Several different nanopores were investigated to determine the effect of mutations on transmembrane pore regions. The mutant pores studied are listed below along with baseline nanopores (baseline pores 1 to 4) for comparison. Several different parameters were investigated to identify the improved nanopores: 1) the mean noise of the signal (where noise equals the standard deviation of all events in the strand, calculated across all strands), which would be lower than the baseline in the improved nanopores; 2) the mean current range, a measure of the range of current levels within the signal, which would be higher than the baseline in the improved nanopores; 3) the mean signal-to-noise ratio (SNR) referenced in the table, which is the SNR across all strands (mean current range divided by the mean noise of the signal), which would be higher than the baseline in the improved nanopores; 4) the DNA capture rate, which would be higher than the baseline in the improved nanopores; and 5) the read head size, which could be increased or decreased in the improved nanopores depending on the read head size of the baseline.
[0766] Each table below contains relevant data for the corresponding baseline nanopore. Table 6 = mutant 1, Table 7 = mutant 2, Table 8 = mutant 3, and Table 9 = mutant 10, which are then compared with the mutant pores.
[0767] Cytolysin mutant 1 = Cytolysin-(E84Q / E85K / E92Q / E97S / D126G)9 (SEQ ID NO: 2 with the mutation E84Q / E85K / E92Q / E97S / D126G). (Baseline 1)
[0768] Cytolysin mutant 2 = Cytolysin-(E84Q / E85K / E92Q / E94D / E97S / D126G)9 (SEQ ID NO: 2 with the mutation E84Q / E85K / E92Q / E94D / E97S / D126G). (Baseline 2)
[0769] Cytolysin mutant 3 = Cytolysin-(E84Q / E85K / E92Q / E94Q / E97S / D126G)9 (SEQ ID NO: 2 with the mutation E84Q / E85K / E92Q / E94Q / E97S / D126G). (Baseline 3)
[0770] Cytolysin mutant 4 = Cytolysin-(E84Q / E85K / S89Q / E92Q / E97S / D126G)9 (SEQ ID NO: 2 with the mutation E84Q / E85K / S89Q / E92Q / E97S / D126G).
[0771] Cytolysin mutant 5 = Cytolysin-(E84Q / E85K / T91S / E92Q / E97S / D126G)9 (SEQ ID NO: 2 with the mutation E84Q / E85K / T91S / E92Q / E97S / D126G).
[0772] Cytolysin mutant 6 = Cytolysin-(E84Q / E85K / E92Q / E97S / S98Q / D126G)9 (SEQ ID NO: 2 with the mutant E84Q / E85K / E92Q / E97S / S98Q / D126G).
[0773] Cytolysin mutant 7 = Cytolysin-(E84Q / E85K / E92Q / E97S / V100S / D126G)9 (SEQ ID NO: 2 with the mutant E84Q / E85K / E92Q / E97S / V100S / D126G).
[0774] Cytolysin mutant 8 = Cytolysin-(E84Q / E85K / E92Q / E94D / E97S / S80K / D126G)9 (SEQ ID NO: 2 with the mutant E84Q / E85K / E92Q / E94D / E97S / S80K / D126G).
[0775] Cytolysin mutant 9 = Cytolysin-(E84Q / E85K / E92Q / E94D / E97S / T106R / D126G)9 (SEQ ID NO: 2 with the mutation E84Q / E85K / E92Q / E94D / E97S / T106R / D126G).
[0776] Cytolysin mutant 10 = Cytolysin-(E84Q / E85K / E92Q / E94D / E97S / T106K / D126G)9 (SEQ ID NO: 2 with the mutation E84Q / E85K / E92Q / E94D / E97S / T106K / D126G). (Baseline 4)
[0777] Cytolysin mutant 11 = Cytolysin-(E84Q / E85K / E92Q / E94D / E97S / T104R / D126G)9 (SEQ ID NO: 2 with the mutation E84Q / E85K / E92Q / E94D / E97S / T104R / D126G).
[0778] Cytolysin mutant 12 = Cytolysin-(E84Q / E85K / E92Q / E94D / E97S / T104K / D126G)9 (SEQ ID NO: 2 with the mutation E84Q / E85K / E92Q / E94D / E97S / T104K / D126G).
[0779] Cytolysin mutant 13 = Cytolysin-(S78N / E84Q / E85K / E92Q / E94D / E97S / D126G)9 (SEQ ID NO: 2 with the mutation S78N / E84Q / E85K / E92Q / E94D / E97S / D126G).
[0780] Cytolysin mutant 14 = Cytolysin-(S82N / E84Q / E85K / E92Q / E94D / E97S / D126G)9 (SEQ ID NO: 2 with the mutation S82N / E84Q / E85K / E92Q / E94D / E97S / D126G).
[0781] Cytolysin mutant 15 = Cytolysin-(E76N / E84Q / E85K / E92Q / E94Q / E97S / D126G)9 (SEQ ID NO: 2 with the mutation E76N / E84Q / E85K / E92Q / E94Q / E97S / D126G).
[0782] Cytolysin mutant 16 = Cytolysin-(E76S / E84Q / E85K / E92Q / E94Q / E97S / D126G)9 (SEQ ID NO: 2 with the mutant E76S / E84Q / E85K / E92Q / E94Q / E97S / D126G).
[0783] Cytolysin mutant 17 = Cytolysin-(E84Q / E85K / E92Q / E94Q / Y96D / D97S / T106K / D126G)9 (SEQ ID NO: 2 with the mutant E84Q / E85K / E92Q / E94Q / Y96D / D97S / T106K / D126G).
[0784] Cytolysin mutant 18 = Cytolysin-(K45D / E84Q / E85K / E92Q / E94K / D97S / T106K / D126G)9 (SEQ ID NO: 2 with the mutant K45D / E84Q / E85K / E92Q / E94K / D97S / T106K / D126G).
[0785] Cytolysin mutant 19 = Cytolysin-(K45R / E84Q / E85K / E92Q / E94D / D97S / T106K / D126G)9 (SEQ ID NO: 2 with the mutant K45R / E84Q / E85K / E92Q / E94D / D97S / T106K / D126G).
[0786] Cytolysin mutant 20 = Cytolysin-(D35N / E84Q / E85K / E92Q / E94D / D97S / T106K / D126G)9 (SEQ ID NO: 2 with the mutation D35N / E84Q / E85K / E92Q / E94D / D97S / T106K / D126G).
[0787] Cytolysin mutant 21 = Cytolysin-(K37N / E84Q / E85K / E92Q / E94D / D97S / T106K / D126G)9 (SEQ ID NO: 2 with the mutant K37N / E84Q / E85K / E92Q / E94D / D97S / T106K / D126G).
[0788] Cytolysin mutant 22 = Cytolysin-(K37S / E84Q / E85K / E92Q / E94D / D97S / T106K / D126G)9 (SEQ ID NO: 2 with the mutant K37S / E84Q / E85K / E92Q / E94D / D97S / T106K / D126G).
[0789] Cytolysin mutant 23 = Cytolysin-(E84Q / E85K / E92D / E94Q / D97S / T106K / D126G)9 (SEQ ID NO: 2 with the mutation E84Q / E85K / E92D / E94Q / D97S / T106K / D126G).
[0790] Cytolysin mutant 24 = Cytolysin-(E84Q / E85K / E92E / E94Q / D97S / T106K / D126G)9 (SEQ ID NO: 2 with the mutation E84Q / E85K / E92E / E94Q / D97S / T106K / D126G).
[0791] Cytolysin mutant 25 = Cytolysin-(K37S / E84Q / E85K / E92Q / E94D / D97S / T104K / T106K / D126G)9 (SEQ ID NO: 2 with the mutant K37S / E84Q / E85K / E92Q / E94D / D97S / T104K / T106K / D126G).
[0792] Cytolysin mutant 26 = Cytolysin-(E84Q / E85K / M90I / E92Q / E94D / E97S / T106K / D126G)9 (SEQ ID NO: 2 with the mutant E84Q / E85K / M90I / E92Q / E94D / E97S / T106K / D126G).
[0793] Cytolysin mutant 27 = Cytolysin-(K45T / V47K / E84Q / E85K / E92Q / E94D / E97S / T106K / D126G)9 (SEQ ID NO: 2 with the mutant K45T / V47K / E84Q / E85K / E92Q / E94D / E97S / T106K / D126G).
[0794] Cytolysin mutant 28 = Cytolysin-(T51K / E84Q / E85K / E92Q / E94D / E97S / T106K / D126G)9 (SEQ ID NO: 2 with the mutant T51K / E84Q / E85K / E92Q / E94D / E97S / T106K / D126G).
[0795] Cytolysin mutant 29 = Cytolysin-(K45Y / S49K / E84Q / E85K / E92Q / E94D / E97S / T106K / D126G)9 (SEQ ID NO: 2 with the mutant K45Y / S49K / E84Q / E85K / E92Q / E94D / E97S / T106K / D126G).
[0796] Cytolysin mutant 30 = Cytolysin-(S49L / E84Q / E85K / E92Q / E94D / E97S / T106K / D126G)9 (SEQ ID NO: 2 with the mutation S49L / E84Q / E85K / E92Q / E94D / E97S / T106K / D126G).
[0797] Cytolysin mutant 31 = Cytolysin-(E84Q / E85K / V88I / M90A / E92Q / E94D / E97S / T106K / D126G)9 (SEQ ID NO: 2 with the mutant E84Q / E85K / V88I / M90A / E92Q / E94D / E97S / T106K / D126G).
[0798] Cytolysin mutant 32 = Cytolysin-(K45N / S49K / E84Q / E85K / E92D / E94N / E97S / T106K / D126G)9 (SEQ ID NO: 2 with the mutant K45N / S49K / E84Q / E85K / E92D / E94N / E97S / T106K / D126G).
[0799] Cytolysin mutant 33 = Cytolysin-(K45N / V47K / E84Q / E85K / E92D / E94N / E97S / T106K / D126G)9 (SEQ ID NO: 2 with the mutant K45N / V47K / E84Q / E85K / E92D / E94N / E97S / T106K / D126G).
[0800]
[0801] Table 6
[0802]
[0803] Table 7
[0804]
[0805] Table 8
[0806]
[0807]
[0808] Table 9
[0809] Read head analysis
[0810] For cytolysin mutants 1 and 10, we obtained models of the expected ionic current distributions for all possible 9-meric polynucleotides. These models may include the mean and standard deviation of the current distribution for each 9-meric polynucleotide.
[0811] We examined and compared the structures of the models obtained for cytosin mutants 1 and 10. (See attached figures) Figure 1 and Figure 2An example of this comparison is provided. In the case of each model (i.e., cytosin 1 or 10), we combine the means of the distributions of all 9-mers of the form A, x_2, x_3, x_4, x_5, x_6, x_7, x_8, x_9, where x_{i} represents any polynucleotide selected from {A, C, G, T}, and this combination is applied to take the median mean. This median mean is repeated for all nucleotides {A, C, G, T} at position 1 and for all positions, such that when the nucleotide is present at any of the 9 positions of the 9-mer, we obtain 36 medians encoding the median effect for each nucleotide.
[0812] Figure 1 (cytolysin mutant 1) and Figure 2 (Cytolysin mutant 2) plotted these medians from two different wells. Figure 1 and Figure 2 The diagram illustrates the degree of differentiation between all bases at each position in the read head. The greater the differentiation, the greater the difference in current contribution levels between those specific positions. If a position is not part of the read head, the current contribution at that position will be similar for all four bases. Figure 2 (Cytolysin mutant 10) shows similar current contributions for all four bases at positions 6 to 8 of the read head. Figure 1 (Cytolysin mutant 1) The similar current contribution of all four bases at any position in the read head is not shown. Therefore, the read head of cytolysin mutant 10 is shorter than that of cytolysin mutant 1. The shorter read head may be advantageous because fewer bases contribute to the signal at any given time, which can lead to improved base recall accuracy.
[0813] Example 2
[0814] This example describes a scheme for producing chemically modified assembly holes with reduced diameter barrels / channels.
[0815] First, the monomeric cytolysin sample (approximately 10 μmol) was reduced to ensure maximum reactivity of cysteine residues, thus ensuring efficient coupling. The monomeric cytolysin sample (approximately 10 μmol) was incubated with 1 mM dithiothreitol (DTT) for 5 to 15 minutes. Cell debris and suspended aggregates were then precipitated by centrifugation at 20,000 rpm for 10 minutes. The soluble fraction was then recovered and buffer-exchanged using a Zeba spinning column (ThermoFisher) with a molecular weight cutoff of 7 kDa to 1 mM Tris, 1 mM EDTA, pH 8.0.
[0816] The molecule to be attached (e.g., 2-iodo-N-(2,2,2-trifluoroethyl)acetamide) is dissolved in a suitable solvent, typically DMSO, to a concentration of 100 mM. This is added to a buffer-exchanged cytosin monomer sample to a final concentration of 1 mM. The resulting solution is incubated at 30 °C for 2 hours. The modified sample (100 μL) is then oligomerized by adding 20 μL of a 5-lipid mixture (phosphatidylserine (0.325 mg / mL): POPE (0.55 mg / mL): cholesterol (0.45 mg / mL): Soy PC (0.9 mg / mL): sphingomyelin (0.275 mg / mL)) from Encapsula Nanosciences. The sample is incubated at 30 °C for 60 minutes. The sample was then subjected to SDS-PAGE and purified from the gel as described in International Application No. PCT / GB2013 / 050667 (published as WO2013 / 153359).
[0817] Example 3
[0818] This example compares a chemically modified assembled cytolysin pore (cytolysin-with 2-iodo-N-(2,2,2-trifluoroethyl)acetamide attached via E94C, comprising (E84Q / E85K / E92Q / E94C / E97S / T106K / D126G / C272A / C283A)9 (with SEQ ID E84Q / E85K / E92Q / E94C / E97S / T106K / D126G / C272A / C283A with mutation E84Q / E85K / E92Q / E94C / E97S / T106K / D126G / C272A / C283A) with a reduced diameter barrel / channel. NO: 2) and cytolysin-(E84Q / E85K / E92Q / E94D / E97S / T106K / D126G / C272A / C283A)9 (SEQ ID NO: 2 with the mutation E84Q / E85K / E92Q / E94D / E97S / T106K / D126G / C272A / C283A).
[0819] Materials and methods
[0820] The DNA construct was prepared as described in Example 1. Electrophysiological experiments were performed as described in Example 1.
[0821] result
[0822] Electrophysiological experiments showed that the chemically modified assembly pore (cytolysin-containing 2-iodo-N-(2,2,2-trifluoroethyl)acetamide attached via E94C, consisting of (E84Q / E85K / E92Q / E94C / E97S / T106K / D126G / C272A / C283A)9 (with SEQ ID E84Q / E85K / E92Q / E94C / E97S / T106K / D126G / C272A / C283A) with mutation E84Q / E85K / E92Q / E94C / E97S / T106K / D126G / C272A / C283A) NO:2) shows a median range of 21 pA, which is greater than that of cytolysin-(E84Q / E85K / E92Q / E94D / E97S / T106K / D126G / C272A / C283A)9, which shows a median range of 12 pA. This increase in median range provides a larger current space for resolving k-mers.
[0823] Figure 3 (cytolysin-(E84Q / E85K / E92Q / E94D / E97S / T106K / D126G / C272A / C283A)9) and Figure 4 ((Cytolysin-containing 2-iodo-N-(2,2,2-trifluoroethyl)acetamide with E94C attachment (E84Q / E85K / E92Q / E94C / E97S / T106K / D126G / C272A / C283A)9 (with mutant E84Q / E85K / E92Q / E94C / E97S / T106K / D126G / C272A / C283A) SEQ ID NO: 2) shows a plot of the median as described in Example 1. In the Figure 4 and Figure 3 When making comparisons, the relative contributions of the signals to different bases at different positions have changed. Figure 4 The extreme readhead positions (positions 7 to 8) show much less differentiation, meaning they contribute significantly less to the signal and therefore the length of the K-mer measured at a given time is shorter. This shorter readhead may be advantageous because fewer bases contribute to the signal at any given time, which can lead to improved accuracy in base recall.
[0824] Cytolysin-containing 2-iodo-N-(2-phenylethyl)acetamide (E84Q / E85S / E92C / E94D / E97S / T106K / D126G / C272A / C283A)9 (SEQ ID with mutant E84Q / E85S / E92C / E94D / E97S / T106K / D126G / C272A / C283A) NO: 2) and cytolysin-9 (E84Q / E85S / E92C / E94D / E97S / T106K / D126G / C272A / C283A) with 1-benzyl-2,5-dihydro-1H-pyrrole-2,5-dione attached via E92C (SEQ ID NO: 2 with mutant E84Q / E85S / E92C / E94D / E97S / T106K / D126G / C272A / C283A) were subjected to experiments similar to those described in Example 3.
Claims
1. An apparatus for characterizing a target polynucleotide, the apparatus comprising: (a) a plurality of mutant lysenins pores, wherein each pore comprises a monomer comprising a variant of the sequence set forth in SEQ ID NO: 2, wherein the mutation of each variant is K45R / E84Q / E85K / E92Q / E94D / D97S / T106K / D126G or K45T / V47K / E84Q / E85K / E92Q / E94D / E97S / T106K / D126G; and (b) a plurality of polynucleotide binding proteins.
2. The apparatus of claim 1, wherein the plurality of mutant lysenins pores comprises a homo-oligomeric pore consisting of monomers comprising a variant of the sequence set forth in SEQ ID NO: 2, the mutation of which is K45R / E84Q / E85K / E92Q / E94D / D97S / T106K / D126G or K45T / V47K / E84Q / E85K / E92Q / E94D / E97S / T106K / D126G.
3. The apparatus of claim 1, wherein the plurality of mutant lysenins pores comprises a hetero-oligomeric pore, each hetero-oligomeric pore comprising at least one monomer comprising a variant of the sequence set forth in SEQ ID NO: 2, the mutation of which is K45R / E84Q / E85K / E92Q / E94D / D97S / T106K / D126G or K45T / V47K / E84Q / E85K / E92Q / E94D / E97S / T106K / D126G.
4. The apparatus of claim 1, wherein the polynucleotide binding proteins comprise a helicase.
5. A transmembrane protein pore comprising at least one lysenin monomer having the sequence set forth in SEQ ID NO: 2 with a mutation of K45R / E84Q / E85K / E92Q / E94D / D97S / T106K / D126G or K45T / V47K / E84Q / E85K / E92Q / E94D / E97S / T106K / D126G.
6. The pore of claim 5, wherein the pore is a hetero-oligomeric pore.
7. The pore of claim 5, wherein the pore is a homo-oligomeric pore.
Citation Information
Patent Citations
Mutant pores
CN116514944A
A miniature support for thin films containing single channels or nanopores and methods for using same
WO2000028312A1
Suspended carbon nanotube field effect transistor
WO2005124888A1
Deliver of molecules to a li id bila
WO2006100484A2
Lipid bilayer sensor system
WO2008102120A1