Mutant pore
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- VLAAMS INTERUNIVERSITAIR INST VOOR BIOTECHNOLOGIE VZW
- Filing Date
- 2022-08-25
- Publication Date
- 2026-08-04
Smart Images

Figure 0007900223000043 
Figure 0007900223000044 
Figure 0007900223000045
Abstract
Description
[Technical Field]
[0001] This invention relates to novel protein pores and their uses. More specifically, this invention relates to nucleic acids Regarding bio-nanopores for sequencing and molecular sensing applications.
[0002] This invention relates to a variant of CsgG. This invention relates to the detection of analytes using CsgG. This also relates to naming and characterization. [Background technology]
[0003] Protein pores are transmembrane polypeptides that contain ions and certain molecules within their membranes. It is a complex that forms a channel through which the offspring can pass. The minimum diameter of this channel is usually is nanometer (10 -9 The range is (meters), and therefore these polypeptides Some of them are called "nanopores."
[0004] Nanopores have great potential as biosensors. When a gas potential is applied, ions flow through the channel. This flow of ions is called an electric current. It can be measured. Suitable electrical measurement techniques using a single-channel recording device include, for example, For example, see International Publication No. 2000 / 28312 pamphlet and D. Stoddart et al., Proc. This is described in Natl. Acad. Sci., 2010, 106, 7702-7. Multichannel recording techniques are, for example, For example, it is described in the international publication pamphlet No. 2009 / 077734.
[0005] Molecules translated by the pore or that bind within or near the pore pass through the channel. It acts to reduce the ion flow by obstructing the on-current. The degree of ion flow reduction measured by this method indicates the magnitude of interference within or near the pore. Therefore, the measured current is used as a measure of the magnitude and degree of interference to the channel. It can be used in this way. The change in current is due to molecules or parts of molecules entering or near the pore. It can be used to identify binding (molecular sensing), or for certain systems In this process, the identity of molecules present within a pore is determined based on their size (nucleic acid sequencing). It can be used for ).
[0006] The "strand sequencing" method, which uses biological nanopores to sequence nucleic acids, is publicly known. Therefore, by passing a single polynucleotide chain through a nanopore, individual nucleotides can be separated. The change in current measured when the above bases temporarily pass through the nanopore channels. The determination is made by chemical analysis. This method is significantly more effective than conventional nucleic acid sequencing methods. It results in time and cost savings.
[0007] Previously reported protein nanopores, e.g., mutant MspA (Manrao et al., Nat (ure Biotechnology, 2012, 30(4), 349-353) and alpha-hemolysin nanopore (Nat. Nano Technol., 2009, 4(4), 265-70) describes nucleic acids using a "strand sequencing" approach. It is used in sequencing. Similarly, other pores are used in protein sensing, for example. , alpha-hemolysin (J Am Chem Soc, 2012, 134(5), 2781-7) and ClyA (Am. Chem. Soc. Nano. 2014, 8(12), 12826-35)(J. Am. Chem. Soc, 2013, 135(36), 13456-63) It is also applicable.
[0008] Lack of prior art, particularly for molecular sensing applications and, for example, nucleic acid sequencing applications. Novel nanopores overcome the lack of prior art in terms of optimizing pore dimensions and characteristics. It is still needed.
[0009] Nanopore sensing allows for the observation of individual binding or interaction events between analyte molecules and receptors. This is an approach to sensing that relies on nanopore sensors. Nanopore sensors are nanometer-sized A single pore is placed within an insulating film, and voltage-dependent ion transport through that pore is observed in the presence of analyte molecules. It can be created by measuring below. The true nature of the analyte is its unique current signal. In particular, this can be revealed by the duration and extent of current interruption and the fluctuations in the current level. It can be done.
[0010] Rapid and inexpensive nucleic acid (e.g., DNA or RNA) sequencing technologies are now widely used. It is needed for a wide range of applications. Existing technologies are time-consuming, expensive, and The main reason is that these technologies rely on amplification techniques to produce large amounts of nucleic acids, and large amounts are used for signal detection. The requirement for specialized fluorescent chemicals lies in nanopore sensing. By reducing the amount of reagents and other components, it provides rapid and inexpensive nucleic acid sequencing. It's possible.
[0011] Two of the essential elements of nucleic acid sequencing using nanopore sensing are, 1) Control of nucleic acid movement through pores and (2) Movement of nucleic acid polymers through pores This is the identification of nucleotides. Traditionally, to perform nucleotide identification, nucleic acids were used with hemolysin. We have passed it through the variant. This has yielded current signs, and such current signs are It has been proven to be sequence-dependent. When using hemolysin pores, a large number of nucleotides This also contributes to the observed current, making a direct correlation between the observed current and polynucleotides difficult. It has been revealed.
[0012] The current range for nucleotide recognition was improved by mutations in the hemolysin pore, but the nucleus If the difference in current between rheotides can be further improved, the sequencing system will be better. This will result in higher performance. In addition, when nucleic acids are moved through the pore, some of the electricity The flow state has been observed to show greater variability. Some mutant hemolysin pores are different from others. It has also been proven that these show larger fluctuations. These fluctuations in state indicate sequence-specific information While it may exist, to simplify the system, it is possible to create pores with small fluctuations. Desirable. It is also desirable to reduce the number of nucleotides that contribute to the observed current. [Overview of the Initiative] [Problems that the invention aims to solve]
[0013] The inventors identified the structure of the bacterial amyloid secretion channel CsgG. Nell is a transmembrane oligomator that forms channels with a minimum diameter of approximately 0.9 nm. It is protein. Due to this structure of CsgG nanopores, CsgG nanopores are proteins. Suitable for sensing applications, particularly nucleic acid sensing. CsgG polypeptide The modified version of the channel further enhances its suitability for such specific applications. It can help to raise the level.
[0014] CsgG pores are advantageous for DNA sensing applications due to their structure, unlike ClyA or It is more advantageous than existing protein pores such as alpha-hemolysin. CsgG pores are ClyA It has a more favorable aspect ratio and, as a result, contains a shorter transmembrane channel than ClyA. SgG pores also have wider channel openings compared to alpha-hemolysin pores. For specific applications, such as facilitating enzyme binding for nucleic acid sequencing applications, Yes, it is possible. In these embodiments, it is possible to use an enzyme and a reading head (with the narrowest pore portion and fixed Minimizing the length of the nucleic acid strand portion located between the (i) and, as a result, directing the readout signal It can also be increased. Narrow internal constriction of the CsgG pore channel affects nucleic acid sequencing. In embodiments of the present invention that include [a specific component], the migration of single-stranded DNA is also facilitated. This narrowing is adjacent to [another component]. The juxtaposition of the tyrosine residue at position 51 (Tyr51) in the protein monomer, as well as at position 56 and the phenylalanine and asparagine residues at position 55 (Phe56 and As It is made up of two ring-shaped rings formed by n55). Changing the dimensions of this constriction This is possible. ClyA allows the passage of double-stranded DNA, which is not currently used in sequencing. It has a much wider internal constriction. The alpha-hemolysin pore has a width of 1.3 nm. It has internal constriction, but also features a 2nm wide beta barrel with an additional reading head. ru. [Means for solving the problem]
[0015] In a first aspect, the present invention relates to a method for molecular sensing, a) A CsgG biopore formed by at least one CsgG monomer is provided within the insulating layer. Steps to kick b) Apply an electrical potential through the insulating layer to confirm the flow of current through the biopore. Standing steps, c) The step of bringing the CsgG biopore into contact with the test substrate, and d) Step of measuring the flow of electric current through the biopore. Regarding methods including
[0016] Typically, the insulating layer is a film such as a lipid bilayer. In the embodiment, the current passing through the pore is insulating The soluble ions are transported by the flow of soluble ions from the first side of the layer to the second side of the insulating layer.
[0017] In embodiments of the present invention, molecular sensing is analyte detection. In certain embodiments, The method for detecting the analyte is to produce when the test substrate is absent after step (d). By reducing the current passing through the biopore compared to the current passing through the body pore, the presence of the test substrate is reduced. This includes further steps to determine presence.
[0018] In an alternative embodiment of the present invention, molecular sensing is nucleic acid sequencing. Typically, The type of nucleic acid sequenced by the above method is DNA or RNA. In a specific embodiment of the present invention, the CsgG biopore is further accessory protein They are adapted to conform to the DNA. Typically, additional accessory proteins are added to the DNA. RNA polymerase, isomerase, topoisomerase, gyrase, telomerase Nucleic acid processing, selected from the group consisting of exonucleases and helicases. It is an enzyme.
[0019] In embodiments of the present invention, the CsgG biopore is a modified CsgG pore, and the modified CsgG The pore is a monomer field in at least one of the CsgG monomers that forms a CsgG pore. It has at least one modification to the viable Escherichia coli (E-coli) CsgG polypeptide sequence. In this invention, the same modification is applied to all CsgG monomers that form CsgG pores. In a specific embodiment, the modified CsgG monomer is located at positions 38-63 according to SEQ ID NOs: 4-388 It has a polypeptide sequence.
[0020] In a second aspect, the present invention relates to a modified CsgG molecule comprising at least one CsgG monomer. A body pore with only one channel constriction having a diameter in the range of 0.5 nm to 1.5 nm. Regarding modified CsgG biopores that do not possess the CsgG monomer polypeptide sequence. The modification occurs at positions 38-63. Preferably, from Tyr51, Asn55 and Phe56. The selected position has a modification. In a particular embodiment, the position Tyr51 has a modification, or the position Both Asn55 and Phe56 have modifiers.
[0021] In embodiments of the present invention, modification of the CsgG monomer is the substitution of a naturally occurring amino acid. A group consisting of deletions of naturally occurring amino acids and modifications of naturally occurring amino acid side chains. Selected from the following. Preferably, the steric hindrance of the unmodified amino acid is reduced or eliminated by the modification. In a specific embodiment, at least one CsgG monomer in the pore is SEQ ID NO: 4~ It has a polypeptide sequence at positions 38-63, indicated by 388.
[0022] In a third aspect, the present invention relates to at least the modified CsgG biopore of the second aspect of the present invention. This concerns an isolated polypeptide encoding a single CsgG monomer.
[0023] In a fourth aspect, the present invention encodes an isolated polypeptide according to a third aspect of the present invention. Regarding isolated nucleic acids.
[0024] In a fifth aspect, the present invention is a) Insulating layer, b) CsgG biopores within the insulating layer, c) A device for measuring the electric current passing through a biopore and Regarding biosensors, including those mentioned above.
[0025] In a particular embodiment, the CsgG biopore within the biosensor is in a second aspect of the present invention. This is a modified CsgG biological pore.
[0026] In a sixth aspect, the present invention relates to the use of CsgG biopores for biosensing applications. Furthermore, the applications of this biosensing technology include analyte detection or nucleic acid sequencing. In the embodiment of the 6th aspect, nucleic acid sequencing is DNA sequencing or RN This is A-sequencing.
[0027] Surprisingly, the inventors have found that CsgG and its novel variants are polynucleotides This invention demonstrates that it can potentially be used to characterize any analyte. A mutant CsgG monomer that interacts with analytes such as polynucleotides. A variant CsgG monoma with one or more modifications to enhance its abilities. - Regarding this. Surprisingly, the inventors have found that pores containing novel mutant monomers are polynes The ability to interact with analytes such as creotides has been improved, and therefore, polynucleotides We also demonstrated an improvement in the ability to estimate the characteristics of analytes, such as the sequence of tides. Mutant pore This showed a remarkable improvement in nucleotide recognition. In detail, the mutant pores had different nuclei. Increased current range to make identification between occlusions easier, and increased signal-to-noise ratio. It showed a remarkable reduction in state fluctuations. In addition, polynucleotides move through the pore. When this is happening, the number of nucleotides contributing to the current decreases. This is because polynucleotides pass through the pore. The direct relationship between the electric current observed when nucleotides are moving and polynucleotides. This makes identification easier. In addition, mutant pores may also show increased throughput. In other words, it is more likely to interact with analytes such as polynucleotides. This makes it easier to characterize the analytes using this method. Mutant pores can enter the membrane more easily. It is possible.
[0028] Therefore, the present invention relates to mutant Cs containing the sequence variant shown in SEQ ID NO: 390. gG monomer, wherein the variant is located at one of positions Y51, N55, and F56. The invention provides a variant CsgG monomer that includes multiple modifications.
[0029] Therefore, the present invention relates to mutant Cs containing the sequence variant shown in SEQ ID NO: 390. gG monomers, wherein the variant CsgG mono includes one or more of the following: Provides a MAR: (i) one or more mutations at the following positions (i.e., the following positions) Mutations in one or more of the following: N40, D43, E44, S54, S57, Q62 , R97, E101, E124, E131, R142, T150 and R192, (ii ) Y51 / N55, Y51 / F56, N55 / F56 or Y51 / N55 / F56 Mutations in (iii) Q42R or Q42K, (iv) K49R, (v) N102 R, N102F, N102Y or N102W, (vi) D149N, D149Q if (viii) D149R, (vii) E185N, E185Q or E185R D195N, D195Q or D195R, (ix)E201N, E201Q or E201R, (x)E203N, E203Q or E203R, and (xi) position F48, K49, P50, Y51, P52, A53, S54, N55, F56 and S5 One or more 7s are missing.
[0030] The present invention - A construct comprising two or more covalently bonded CsgG monomers, wherein the monomers are small A construct in which at least one of the mutant monomers of the present invention is present, -Polynucleotides encoding mutant monomers or constructs of the present invention - Derived from CsgG containing the same mutant monomer or the same construct of the present invention Homo oligomer pore, - At least one mutant monomer of the present invention or at least one construct of the present invention heterooligomeric pores derived from CsgG, - A method for determining the presence, absence, or one or more characteristics of a target analyte, a) The target analyte is moved relative to the CsgG pore or its variant, and the target analyte moves relative to the pore. The steps of bringing them into contact, and b) Take one or more measurements while the analyte is moving through the pore, and This is a step to determine the presence or absence of the analyte, or one or more of its characteristics. Methods including - A method for forming a sensor for characterizing a target polynucleotide, wherein Csg It forms a complex with a G pore or its variant and a polynucleotide-binding protein, and Therefore, the method includes the step of forming a sensor for characterizing a target polynucleotide. law, - A sensor for characterizing a target polynucleotide, which is a CsgG pore or A sensor containing a complex of a variant and a polynucleotide-binding protein, - CsgG point for determining the presence, absence, or one or more features of the target analyte A or the use of its variant, - A kit for characterizing a target analyte, comprising (a) CsgG pore or its mutation A kit containing (b) the body and membrane components, - An apparatus for characterizing a target analyte in a sample, comprising (a) a plurality of CsgG pores or an apparatus comprising (b) those variants and multiple membranes, - A method for characterizing a target polynucleotide, a) Polynucleotides are converted into CsgG pores or their variants, polymerases and labeled nucleotides. Rheotide and phosphate-labeled chemical species are sequentially converted to target polynucleotides by polymerase. The phosphate species is brought into contact with the nucleotides so that it can be attached to them, and each nucleotide contains a target specific to it. The steps that are, b) Use a pore to detect phosphate-labeled chemical species, thereby identifying polynucleotides. Steps to mark Methods including, - A method for producing mutant monomers or constructs of the present invention, wherein a suitable host The polynucleotide of the present invention is expressed in cells, thereby creating the mutant monomer of the present invention. - A method including the step of producing a construct. We also offer this. [Brief explanation of the drawing]
[0031] [Figure 1] This is a side cross-sectional view showing the structure of the channel conformation of the CsgG nonomer in ribbon and surface displays. [Figure 2] This is a cross-sectional view showing the measurement of CsgG channel constriction (i.e., the pore reading head in a nanopore sensing application) and associated diameter. [Figure 3] This is a structural motif diagram that contributes to pore constriction, including three stacked concentric side chain layers: Tyr51, Asn55, and Phe56. [Figure 4] This figure shows the sequence homology of CsgG homologs, including multiple sequence alignments of CsgG-like proteins (SEQ ID NOs: 442-448). The selected sequences were chosen from a monophyletic clade of the entire phylogenetic tree of CsgG-like sequences (not shown) and represent a typical diagram of sequence diversity. Secondary structure elements of β-strands or α-helices are shown as arrows or bars, respectively, and these are based on the crystal structure of E. coli CsgG. For their importance, residues equivalent to E. coli Thr51, Asn55, and Phe56 are highlighted with arrows. These residues form the internal constriction of the pore, i.e., the pore read head in the context of nanopore sensing applications. [Figure 5] This graph shows representative single-channel current recordings (a) and conductance histograms (b) of CsgG reconstituted into a planar phospholipid bilayer and measured in electric fields of +50mV (n=33) or -50mV (n=13). [Figure 6] This graph shows single-channel current recordings of PPB-reconstructed CsgG at +50mV or -50mV, and with supplementation of increased CsgE concentrations. The horizontal scale bar is at 0pA. [Figure 7]a. This figure shows a raw negative stained EM image of C8E4 / LDAO solubilized CsgG. The arrows indicate different particle populations shown in the size exclusion profile indicated in g: (I) aggregates of CsgG nonamers, (II) CsgG octadecamers, and (III) CsgG nonamers. Scale bar, 20 nm. b. Representative class-mean images of top and side views of the oligomerization states shown. c. Autocorrelation function graph in the rotational direction of LDAO solubilized CsgG in the top view, showing 9 degrees of symmetry. d. Raw negative stained EM image of CsgGC1S. The arrows indicate hexadecamer (IV) and octamer (V) particles observed in size exclusion chromatography in g. e. Representative class-mean image of the side view of the CsgGC1S oligomer. No top view was observed for this construct. f, g: Table showing the elution volume (EV), calculated molecular weight (MWcalc), predicted molecular weight (MWCsgG) corresponding to the oligomerization state of CsgG (CsgGn), and particle symmetry observed by size exclusion chromatography of CsgGC1S and CsgG particles, as shown by size exclusion chromatography. g: Size exclusion chromatograms of CsgGC1S (black) and C8E4 / LDAO solubilized CsgG (gray) performed on Superdex 200 10 / 300 GL (GE Healthcare). h, i: Top and side views of the ribbon representation of the crystallized oligomers, showing the D8 hexadecamer of CsgGC1S (h) and the D9 octadecamer of membrane-extracted CsgG (i). One protomer is shown in iridescent colors from the N-terminus (blue) to the C-terminus (red). Two C8 octamers (CgGC1S) or C9 nonamers (CsgG) that form tail-tail dimers trapped in the crystal are shown in blue and yellowish-brown. r and θ represent the radius and interprotomer rotation, respectively. [Figure 8]This figure shows the electron density map of CsgGC1S at 2.8 Å, calculated using an experimental SAD phase averaged and density corrected by NCS, and displayed as a 1.5σ contour. The map shows the channel construction region (CL: showing a single protomer) and is superimposed on the final refined model. CsgGC1S is a mutant CsgG in which the N-terminal Cys, i.e., Cys1, of the mature CsgG sequence is replaced with Ser, thereby lacking lipid modification by the E. coli (LOL) pathway. This results in a soluble homooctamer oligomer that exists in a prepore conformation opposite to that of the membrane-targeted homonomerpore formed by native lipid-modified CsgG (see Figure 42) (Figure 43). [Figure 9] Figure 9a and Figure 9b show a top view (top view) and a side view (side view) of a CsgG constriction model by a polyalanine chain traversing an elongated conformational channel from the C-terminus to the N-terminus. To clarify the model solvation of the polyalanine chain present in Figure 9b, the C-loop is removed and shown in Figure 9c (showing solvent molecules within 10 Å of the entire polyalanine chain). [Figure 10] This figure shows CsgG from Escherichia coli (E. coli). [Figure 11] This diagram shows the dimensions of CsgG. [Figure 12] This figure shows a 1G rearrangement at 10 Å / nanosecond. A large barrier exists for guanine to enter the F56 ring of CsgG-Eco. *=G enters the F56 ring. A=G terminates the interaction with the 56 ring. B=G terminates the interaction with the 55 ring. C=G terminates the interaction with the 51 ring. [Figure 13] This figure shows ssDNA transposition at 100 Å / nanosecond. A greater force is required to pull the DNA into the constriction of CsgG-Eco. [Figure 14] This figure shows ssDNA transposition at 10 Å / nanosecond. Both the CsgG-F56A-N55S and CsG-F56A-N55S-Y51A mutants have a lower barrier to ssDNA transposition. [Figure 15]This graph shows mutant pores that exhibit an increased range compared to the wild type (WT). [Figure 16] This graph shows mutant pores that exhibit an increased range compared to the wild type (WT). [Figure 17] This graph shows mutant pores that exhibit an increased range compared to the wild type (WT). [Figure 18] This graph shows mutant pores that exhibit increased throughput compared to the wild type (WT). [Figure 19] This graph shows mutant pores that exhibit increased throughput compared to the wild type (WT). [Figure 20] This graph shows mutant pores that exhibit increased insertions compared to the wild type (WT). [Figure 21] This graph shows mutant pores that exhibit increased insertions compared to the wild type (WT). [Figure 22] This figure shows DNA construct X used in Example 18. The region labeled 1 corresponds to 30 SpC3 spacers. The region labeled 2 corresponds to Sequence ID No. 415. The region labeled 3 corresponds to 4 iSp18 spacers. The region labeled 4 corresponds to Sequence ID No. 416. The region labeled 5 corresponds to 4 5-nitroindoles. The region labeled 6 corresponds to Sequence ID No. 417. The region labeled 7 corresponds to Sequence ID No. 418. The region labeled 8 corresponds to Sequence ID No. 419, with 4 iSp18 spacers (the region labeled 9) bound to its 3' end. The opposite end of the iSp18 spacer was a 3' cholesterol tether (labeled 10). The region labeled 11 corresponds to 4 SpC3 spacers. [Figure 23]This graph shows a chromatographic trace as an example of Strep trap (GE Healthcare) purification of CsgG protein (x-axis = elution volume (mL), Y-axis = absorbance (mAu)). Samples were loaded in 25 mM Tris, 150 mM NaCl, 2 mM EDTA, and 0.01% DDM and eluted with 10 mM desthiobiotin. The elution peak where CsgG protein is eluted is indicated as E1. [Figure 24] This photograph shows a typical SDS-PAGE visualization example of CsgG protein after initial Strep purification. A 4-20% TGX gel (Bio Rad) was run in 1×TGS buffer at 300V for 22 minutes. The gel was stained with Sypro Ruby stain. Lanes 1-3 show the main elution peak containing CsgG protein (indicated as E1 in Figure 23), indicated by the arrows. Lanes 4-6 correspond to the elution fraction of the tail of the main elution peak (indicated as E1 in Figure 23) containing contaminants. M indicates the molecular weight marker used, which was Novex Sharp Unstained (unit = kD). [Figure 25] This graph shows an example of a size exclusion chromatogram (SEC) of a CsgG protein (120 mL S200, GE Healthcare, x-axis = elution volume (mL), y-axis = absorbance (mAu)). SEC was performed after Strep purification and heating of the protein sample. The running buffer for SEC was 25 mM Tris, 150 mM NaCl, 2 mM EDTA, 0.01% DDM, 0.1% SDS, pH 8.0, and the column was eluted at a rate of 1 mL / min. Traces labeled X show absorbance at 220 nm, and traces labeled Y show absorbance at 280 nm. Peaks marked with an asterisk were collected. [Figure 26]This photograph shows a typical example of SDS-PAGE visualization of CsgG protein after SEC. A 4-20% TGX gel (Bio Rad) was electrophoresed in 1×TGS buffer at 300V for 22 minutes, and then the gel was stained with Sypro Ruby stain. Lane 1 shows the CsgG protein sample after Strep purification and heating, but before SEC. Lanes 2-8 show the fractions collected from the entire peak eluting approximately 48mL-60mL (peak center = 55mL) in Figure 25, indicated by stars in Figure 25. M indicates the molecular weight marker used, which was Novex Sharp Unstained (unit = kD). The bar corresponding to the CsgG-Eco pore is indicated by an arrow. [Figure 27] This graph shows mutant pores that exhibit an increased range compared to the wild type (WT). [Figure 28] This graph shows mutant pores that exhibit an increased range compared to the wild type (WT). [Figure 29] This graph shows mutant pores that exhibit an increased range compared to the wild type (WT). [Figure 30] This graph shows mutant pores that exhibit an increased range compared to the wild type (WT). [Figure 31] This graph shows mutant pores that exhibit an increased range compared to the wild type (WT). [Figure 32] This graph shows mutant pores that exhibit an increased range compared to the wild type (WT). [Figure 33] This graph shows mutant pores that exhibit an increased range compared to the wild type (WT). [Figure 34] This graph shows mutant pores that exhibit an increased range compared to the wild type (WT). [Figure 35] This graph shows mutant pores that exhibit an increased range compared to the wild type (WT). [Figure 36] This graph shows mutant pores that exhibit an increased range compared to the wild type (WT). [Figure 37] This graph shows mutant pores that exhibit an increased range compared to the wild type (WT). [Figure 38] This graph shows mutant pores that exhibit an increased range compared to the wild type (WT). [Figure 39] This graph shows mutant pores that exhibit an increased range compared to the wild type (WT). [Figure 40] These are snapshot images of the enzyme (T4 Dda-(E94C / C109A / C136A / A360C) (sequence number 412, which has the mutation E94C / C109A / C136A / A360C and then (ΔM1)G1G2) on the pore (CsgG-Eco-(Y51T / F56Q)-StrepII(C))9 (which has the mutation Y51T / F56Q, StrepII(C) is sequence number 435, and binds to C-terminal pore mutant number 20, sequence number 390) taken at 0 and 20 nanoseconds in the simulations (experiments 1 to 3). [Figure 41] These are snapshot images of the enzyme (T4 Dda-(E94C / C109A / C136A / A360C) (sequence number 412, which has the mutation E94C / C109A / C136A / A360C and then (ΔM1)G1G2) on the pore (CsgG-Eco-(Y51T / F56Q)-StrepII(C))9 (which has the mutation Y51T / F56Q, StrepII(C) is sequence number 435, and binds to C-terminal pore mutant number 20, sequence number 390) taken at 30 and 40 nanoseconds in the simulations (experiments 1 to 3). [Figure 42] This figure shows the X-ray structure of CsgGC1S in prepore conformation. a. Ribbon diagram of the CsgGC1S monomer, shown in iridescent colors from blue to red from the N-terminus to the C-terminus. Secondary structural elements are shown according to ABD-like folding, and further N-terminal and C-terminal α-helices, as well as the elongated loop connecting β1 and α1, are labeled αN, αC, and C-loop (CL), respectively. b. Side view of the CsgGC1SC8 octamer, where multiple subunits are identified by color, but the orientation and color of a single subunit are the same as in a. [Figure 43]This figure shows the structure of the channel conformation of CsgG. a. Amide I region (1,700–1,600 cm⁻¹) of ATR-FTIR spectra of CsgGC1S (blue) and membrane-extracted CsgG (red). b. TM1 and TM2 sequences (SEQ ID NOs. 449 and 450) (residues facing the bilayer, shown in blue), as well as Congo red bonds of E. coli (E. coli) BW25141ΔcsgG complemented by wild-type csgG (WT), an empty vector, or csgG lacking the underlined fragments of TM1 or TM2. Data are representative of three biological replicate experiments. c. Overlay of CsgG monomers with prepore (light blue; TM1 is pink, TM2 is purple) and channel conformation (yellowish-brown; TM1 is green, TM2 is orange). CL, C-loop. d, e, Side view (d) and cross-sectional view (e) of the CsgG nonamer in ribbon and surface representation; helix 2, core domain, and TM hairpin are shown in blue, light blue, and yellowish-brown, respectively. A single protomer is shown in the same colors as in Figure 42a. The magenta sphere indicates the location of Leu2. OM, outer membrane. [Figure 44] This figure shows CsgG channel constriction. a. Cross-sectional view of the CsgG channel constriction and its diameter excluding the solvent. b. The constriction consists of three stacked concentric side chain layers: Tyr51, Asn55, and Phe56 from the peripheral side after Phe48. c. Topology of the CsgG channel. d. Congo red binding of csgG with csgG (WT), an empty vector, or the constricted mutant shown, complemented by csgG(WT). Data are representative of six biological replicate experiments. e, f. Representative single-channel current recordings (e) and conductance histograms (f) of CsgG reconstituted in a planar phospholipid bilayer and measured in an electric field of +50mV (n=33) or -50mV (n=13). [Figure 45]This is a model diagram of the CsgG transport mechanism. a. Native PAGE of CsgE(E), CsgG(G), and CsgG(E+G) supplemented with excess CsgE shows the formation of the CsgG-CsgE complex (EG*). Data are representative of 7 experiments, including 4 protein batches. b. SDS-PAGE of CsgE(E), CsgG(G), and EG* complexes recovered from native PAGE. Data are representative of 2 replicate experiments. m. Molecular weight markers. c. Selected class-mean images of CsgG-CsgE particles. From left to right: Top and side views visualized by cryo-EM, as well as a comparison of negatively stained side views with CsgG nonamers. d. Cryo-EM mean images of CsgE particles in top and oblique side views. Autocorrelation in the rotational direction shows 9 symmetries. e. Three-dimensional reconstruction of CsgG-CsgE (resolution 24 Å, 1,221 single particles) shows nonamer particles containing CsgG (blue) and further densities assigned as CsgE nonamers (orange). f. Single-channel current recording of PPB-reconstructed CsgG supplemented with increased CsgE concentrations at +50 mV or -50 mV. Bars on the horizontal scale are located at 0 pA. g. Proposed CsgG-mediated protein secretion model. It has been proposed that CsgG and CsgE form a secretion complex that captures CsgA (discussed in Figure 54), generating an entropy potential above the channel. Following the capture of CsgA at channel constriction, DS-regulated Brownian diffusion facilitates the progressive translocation of polypeptides across the outer membrane. [Figure 46]This diagram shows the Curli biosynthesis pathway in Escherichia coli (E. coli). The large curli subunit CsgA (light green) is secreted from the cell as a soluble monomer protein. The small curli subunit CsgB (dark green) associates with the outer membrane (OM) and acts as a nucleating agent for converting CsgA from a soluble protein to an amyloid deposit. CsgG (orange) assembles into an oligomer-specific transposition channel in the outer membrane. CsgE (purple) and CsgF (light blue) form soluble accessory proteins necessary for the transport and deposition of productive CsgA and CsgB. CsgC forms a putative oxidoreductase of unknown function. All curli proteins have putative Sec signaling sequences for transport across the cytoplasmic (inner) membrane (IM). [Figure 47]This figure shows the oligomerization states of CsgG and CsgGC1S in solution as analyzed by size exclusion chromatography and negative staining electron microscopy. a. Raw negative staining EM image of C8E4 / LDAO solubilized CsgG. Arrows indicate different particle populations shown in the size exclusion profile indicated in g: (I) aggregates of CsgG nonamers, (II) CsgG octadecamers, and (III) CsgG nonamers. Scale bar, 20 nm. b. Representative class-mean images of top and side views of the oligomerization states shown. c. Autocorrelation function graph in the rotation direction of LDAO solubilized CsgG in the top view, showing 9 degrees of symmetry. d. Raw negative staining EM image of CsgGC1S. Arrows indicate hexadecamer (IV) and octamer (V) particles observed by size exclusion chromatography indicated in g. e. Representative class-mean image of the side view of CsgGC1S oligomers. The top view was not observed for this construct. Tables f and g show the elution volume (EV), calculated molecular weight (MWcalc), predicted molecular weight (MWCsgG) corresponding to the oligomerization state of CsgG (CsgGn) of CsgGC1S and CsgG particles observed by size exclusion chromatography, as well as the symmetry of the particles observed by negative staining EM and X-ray crystallography. g is a size exclusion chromatogram of CsgGC1S (black) and C8E4 / LDAO solubilized CsgG (gray) performed on a Superdex 200 10 / 300 GL (GE Healthcare). h and i are top and side views of the ribbon representation of the crystallized oligomers, showing the D8 hexadecamer of CsgGC1S (h) and the D9 octadecamer of membrane-extracted CsgG (i). One protomer is shown in iridescent colors from the N-terminus (blue) to the C-terminus (red). Two C8 octamers (CgGC1S) or C9 nonamers (CsgG) that form a tail-tail dimer trapped in the crystal are shown in blue and yellowish-brown. r and h indicate radius and interprotomer rotation, respectively. [Figure 48]This figure shows a comparison of CsgG and its structural homolog, and interprotomeral contact in CsgG. Both a and b are ribbon diagrams of the CsgGC1S monomer (e.g., CsgG in prepore conformation) (a) and the nucleotide-binding domain-like domain of TolB (b) (PDB2hqs), both exhibiting iridescence from the N-terminus (blue) to the C-terminus (red). Common secondary structural elements are shown equally. c. From left to right, CsgGC1S (gray) superimposed on the rare lipoprotein B of Xanthomonas campestris (PDB 2r76, pink), the virtual lipoprotein DUF330 of Shewanella oneidensis (PDB 2iqi, pink), and Escherichia coli TolB (PDB 2hqs, shown in pink and yellow for the N-terminus and β-propeller domain, respectively). CsgG-specific structural elements are displayed and colored as shown in the upper left panel. d and e. Ribbon diagrams of two adjacent protomers found in the CsgG structure viewed along the bilayer plane from either the outside (c) or inside (d) of the oligomer. One protomer is shown in iridescent colors (dark blue to red) from the N-terminus to the C-terminus. The second protomer is shown in light blue (core domain), blue (helix 2), and yellowish-brown (TM domain). Four main oligomerization interfaces are evident: the b6-b39 main chain interaction within the β-barrel, the constriction loop (CL), the side-chain packing of helix 1 (α1) to b1-b3-b4-b5, and the helix-helix packing of helix 2 (α2). The 18-residue N-terminal loop connecting the lipid anchor (magenta sphere indicates the position of Ca in Leu2) to the N-terminal helix (αN) also appears to wrap around the two adjacent protomers. The protruding position of the lipid anchor is expected to be opposite to the TM1 and TM2 hairpins of the +2 protomer (not shown for clarity). [Figure 49]This shows an assay of Cys accessibility of selected surface residues in CsgG oligomers. a-c, Ribbon representation of CsgG nonomers shown in peripheral (a), lateral (b), and extracellular figures. One protomer is shown in iridescent colors from the N-terminus (blue) to the C-terminus (red). Cysteine substitutions are shown, with equivalent positions of S atoms shown as spheres, and colored according to their accessibility to MAL-PEG (5,000 Da) labeling on the outer membrane of E. coli (E. coli). d, Western blot of MAL-PEG reaction samples analyzed by SDS-PAGE, showing a 5 kDa increase due to MALPEG binding of introduced cysteine. Accessible (11 and 111), moderately accessible (1), and inaccessible (2) sites are shown in green, orange, and red in a-e, respectively. In Arg97 and Arg110, a second chemical species is present at 44 kDa, corresponding to a fraction of the protein where both introduced and native cysteine become labeled. The data are representative of four independent experiments from biological replicate experiments. e. Side view of the dimer-forming interface in the D9 octadecamer present in the X-ray structure. Labeling of introduced cysteine at the dimer-forming interface or inside the lumen of the D9 particle. In membrane-bound CsgG, these residues are accessible to MAL-PEG, demonstrating that the D9 particle is an artifact of a concentrated solution of membrane-extracted CsgG and that the C9 complex forms physiologically relevant chemical species. Residues in the C-terminal helix (aC; Lys242, Asp248, and His255) were found to be inaccessible or poorly accessible, indicating that aC may form further contact with the E. coli (E. coli) cell envelope, possibly the peptidoglycan layer. [Figure 50]This figure shows a molecular dynamics simulation of CsgG constriction using a model polyalanine chain. a and b are top view (a) and side view (b) of the CsgG constriction modeled by a polyalanine chain traversing a channel of elongated conformation, shown from the C-terminus to the N-terminus. The passage of substrates within the CsgG transporter is not sequence-specific in itself (References 16, 23). For clarity, a polyalanine chain was used to model the estimated interaction of the passing polypeptide chain. The model region consists of nine concentric CsgG C-loops, each containing 47–58 residues. The side chains lining the inside of the constriction are shown as sticks, with Phe51 in dull blue, Asn55 (amide clamp) in cyan, and Phe48 and Phe56 (Φ-clamp) in light and dark orange, respectively. N, O, and H atoms (only hydroxyl or side-chain amide H atoms are shown) are shown in blue, red, and white, respectively. The C, N, O, and H atoms of the polyalanine chain are shown in green, blue, red, and white, respectively. Solvent molecules (water) within 10 Å of polyalanine residues in constriction (residues labeled 11-15) are shown as red dots. Model solvation of the polyalanine chain (showing solvent molecules within 10 Å of the entire polyalanine chain), similarly located in c and b, with the C-loop removed for clarity. At the highest points of the amide-clamp and Φ-clamp, the solvation of the polyalanine chain is reduced to a single water shell connecting the peptide backbone and the amide-clamp side chain. Most of the side chains in the Tyr51 ring are rotated toward the solvent compared to their inward-facing center-facing position observed in the CsgG (and CsgGC1S) X-ray structure. The model is the result of a 40-nanosecond explicit solvent molecular dynamics simulation of all atoms according to GROMACS reference 53, using the AMBER99SB-ILDN54 force field and restricting the positions of the Cα atoms at the terminal residues of the C-loop (Gln47 and Thr58). [Figure 51]This figure shows sequence conservation in CsgG homologs. a) Surface representation of CsgG nonamares viewed from the peripheral environment (far left), lateral view (left center), extracellular environment (right center), or cross-sectional lateral view (far right), colored according to sequence similarity (from low to high conservation score, yellow to blue). The figure shows that the highest sequence conservation regions map to the entrance of the peripheral vestibule, the lateral view of the antral loop, and the luminal surface of the TM domain. b) Multiple sequence alignment of CsgG-like lipoproteins. The selected sequences were chosen from a monophyletic clade of the entire phylogenetic tree of CsgG-like sequences (not shown) to provide a representative figure of sequence diversity. Based on the CsgG crystal structure of Escherichia coli (E. coli), secondary structure elements are shown as arrows or bars with respect to β-strands or α-helices, respectively. c and d show cross-sectional lateral views of the CsgG protomer in secondary structure representation (c) and the CsgG nonomer in surface representation (d), both in gray, with three consecutive blocks exhibiting high sequence conservation shown in red (HCR1), blue (HCR2), and yellow (HCR3). HCR1 and HCR2 form the antral lateral surfaces of the constriction loop. HCR3 corresponds to helix 2, located at the entrance of the peripheral vestibule. Within the constriction, Phe56 is 100% conserved, but Asn55 can be conservatively substituted for Ser or Thr, with a small polar side chain that can act, for example, as a hydrogen bond donor / acceptor. The concentric side chain ring (Tyr51) at the exit of the constriction is not conserved. The presence of the Phe ring at the entrance of the constriction is topologically similar to the Phe427 ring (called the Φ-clamp) of the Bacillus anthrax protective antigen PA63 (reference 20), which has been shown to catalyze polypeptide capture and passage. MST of toxB superfamily proteins reveals a conserved motif D(D / Q)(F)(S / N)S at the highest point of the Phe ring. This is similar to the S(Q / N / T)(F)ST motif found in curli-like transporters.Although the atomic decomposition structure of PA63 in pore conformation is not yet available, available structures suggest that a conserved hydrogen bond donor / acceptor (Ser / Asn428) may follow the Phe ring as a subsequent concentric ring in the dislocation channel (note that the orientation of the elements is opposite in both transporters). [Figure 52]This graph shows the single-channel current analysis of CsgG and CsgG:CsgE pores. a) Under negative electric field potential, CsgG pores exhibit two conductance states. The upper left and upper right panels show representative single-channel current traces of the normal (measured at +50, 0, and -50 mV) and low-conductance (measured at 0, +50, and -50 mV) states, respectively. No conversion between the two states was observed during the total observation time (n=22), indicating that the conductance state has a long lifetime (timescale from seconds to minutes). The lower left panel shows current histograms of the normal and low-conductance CsgG pores acquired at +50 and -50 mV (n=33). The lower right panel shows the IV curves of CsgG pores with normal and low conductance. Data represent the mean and standard deviation from at least four independent records. The properties or physiological presence of the low-conductance type are unknown. b. Electrophysiology of CsgG channels titrated with the accessory factor CsgE. The plot shows the fractionation of the open, intermediate, and closed channels as a function of CsgE concentration. The open and closed states of CsgG are illustrated in Figure 45f. When the CsgE concentration is increased above 10 nM, the CsgG pore closes. This effect occurs at +50 mV (left) and -50 mV (right), ruling out the possibility that the pore blockage is caused by electrophoresis of CsgE (calculated pI 4.7) into the CsgG pore. The infrequent (5%) intermediate state has almost half the conductance of the open channel. This may represent incomplete closure of the CsgG channel induced by CsgE. Alternatively, this could represent the transient formation of CsgG dimers caused by the binding of residual CsgG monomers in the electrolyte solution to pores embedded in the membrane. The three-state fractions were obtained from histogram analysis of single-channel current traces at all time points. From the histograms, peak areas for up to three states were obtained, and the fraction for a given state was obtained by dividing the corresponding peak area by the sum of all other states in the recording. Under negative electric field potential, the two open-conductance states are identified in the same way as the findings for CsgG (see a).Since both open-state channel variants are blocked by higher CsgE concentrations, the “open-state” trace in b represents a combination of both conductance types. The data in the plot represent the mean and standard deviation from three independent recordings. c, crystal structure, size exclusion chromatography, and EM show that the washing agent extracted CsgG pores form non-native tail-tail stacked dimers (e.g., two nonamers as D9 particles, Figure 47) at higher protein concentrations. These dimers can also be observed in single-channel recordings. The upper panel shows single-channel current traces of CsgG pores stacked at +50, 0, and -50 mV (left to right). The lower left panel shows current histograms of dimer CsgG pores recorded at +50 mV and -50 mV. The experimental conductances of +16.2±1.8 and -16.0±3.0 pA (n=15) at +50mV and -50mV are close to the theoretically calculated value of 23 pA, respectively. The lower right panel shows the IV curve of the stacked CsgG pore. The data represent the mean and standard deviation from six independent records. d. The ability of CsgE to couple and block the stacked CsgG pore was tested electrophysiologically. Single-channel current traces of the stacked CsgG pore in the presence of 10 or 100 nM CsgE at +50mV (top) and -50mV (bottom) are shown. The current traces indicate that pore closure of the stacked CsgG dimer does not occur by otherwise saturated concentrations of CsgE. These findings are in good agreement with the fact that the CsgG-CsgE contact region maps to the mouth portion of the helix 2 and CsgG peripheral cavity, as identified by EM and site-directed mutagenesis (Figures 45 and 52). [Figure 53]This figure shows CsgE oligomers and CsgG-CsgE complexes. a. Size exclusion chromatography of CsgE (Superose 6, 16 / 600; running buffer 20 mM Tris-HCl, pH 8, 100 mM NaCl, 2.5% glycerol) shows equilibrium of two oligomer states 1 and 2, with an apparent molecular mass ratio of 9.16:1. Examination of the negative stained EM of peak 1 reveals nine copies of CsgE and separate CsgE particles that fit in size (five representative class-mean images are shown in the inset in increasing order of tilt angle). b. Selected class-mean images of CsgE oligomers observed in the top view of the cryo-EM and their autocorrelation in the direction of rotation indicate the presence of C9 symmetry. c. FSC analysis of the CsgG-CsgE cryo-EM model. Three-dimensional reconstruction achieved a resolution of 24 Å when determined by FSC with a correlation threshold of 0.5, using 125 classes corresponding to 1,221 particles. d. Superposition of CsgG nonamers observed in CsgG-CsgE cryo-EM density and X-ray structure. The superposition is shown as translucent density (left) or as a side view as a cross-section. e. Congo red linkage of Escherichia coli (E. coli) BW25141ΔcsgG complemented with wild-type csgG (WT), empty vector (ΔcsgG), or csgG helix 2 variants (1 amino acid substitutions represented by single-letter codes). Data are representative of four biological replicate experiments. f. Effect of bile salt toxicity on Escherichia coli (E. coli) LSR12 complemented with csgG (WT), or on csgG with (1) csgE complemented or uncomplemented, or (2) different helix 2 variants. A series of 10-fold dilutions starting with 107 bacteria were spotted onto McConkey agar plates. Expression of CsgG pores in the outer membrane leads to increased bile salt sensitivity, which can be blocked by co-expression of CsgE (n=6, 3 biological replications, 2 replicates each). g. Cross-sectional view of the X-ray structure of CsgG in molecular surface representation. CsgG mutants that do not show effect against Congo red binding or toxicity are shown in blue, and mutants that interfere with rescue by bile salt-sensitive CsgE are shown in red. [Figure 54]This figure shows the assembly and substrate recruitment of the CsgG secretory complex. It has been proposed that the curli transporter CsgG and the soluble secretory cofactor CsgE form a secretory complex with a 9:9 stoichiometry that houses a 24,000 ų chamber, which captures the CsgA substrate and facilitates its entropy-driven diffusion across the outer membrane (OM; see text and Figure 45). From a theoretical standpoint, three hypothetical pathways (a-c) can be imagined regarding the substrate recruitment and assembly of the secretory complex. a. The "catch-and-cap" mechanism involves the binding of CsgA to the apoCsgG transposition channel (1), which changes the conformation of the apoCsgG transposition channel, exposing a high-affinity binding platform for CsgE binding (2). The binding of CsgE causes capping of the substrate cage. When CsgA is secreted, CsgG reverts to its low-affinity conformation, and the dissociation and release of CsgE from the secretory channel occurs for a new secretory cycle. b. In the “dock and trap” mechanism, peripheral CsgA is initially captured by CsgE (1), and CsgE becomes a high-affinity complex that docks in the CsgG rearrangement pore (2), trapping CsgA within the secretory complex. The binding of CsgA can be direct to either a CsgE oligomer or a CsgE monomer, which is then oligomerized and binds to CsgG. When CsgA is secreted, CsgE reverts to its low-affinity conformation and dissociates from the secretory channel. c. CsgG and CsgE form a constitutive complex, and the CsgE conformational dynamics cycle between open and closed forms during the recruitment and secretion of CsgA. The currently published or available data cannot identify or propose any of the mobilization modes or derivative modes of these estimates. [Figure 55]This figure shows data acquisition statistics and electron density maps for CsgGC1S and CsgG. a. Data acquisition statistics for the X-ray structures of CsgGC1S and CsgG. b. Electron density map at 2.8 Å for CsgGC1S, calculated using NCS-averaged and density-corrected experimental SAD phases and displayed with 1.5σ contours. The map shows the channel construction region (CL; showing a single prototype) and is superimposed on the final refined model. c. Electron density map in the CsgG TM domain region calculated from NCS-averaged and density-corrected molecular substitution phases (TM loop was not present in the input model) (along reciprocal lattice vectors a*, b*, and c*, respectively, with resolutions of 3.6, 3.7, and 3.8 Å); factor B sharpened at -20 Ų and displayed with 1.0σ contours. The figure shows the TM1 (Lys135~Leu154) and TM2 (Leu182~Asn209) regions of a single CsgG protomer superimposed on the final refined model. [Figure 56] This graph shows a single-channel current trace (left) and its enlarged region (right) of a CsgG WT protein interacting with a DNA hairpin having a single-stranded DNA overhang. The trace shows the current changing in response to potentials measured at +50mV or -50mV intervals (indicated by arrows). The downward current cutoff in the last +50mV segment represents the simultaneous occurrence of lodging of the hairpin double helix within the pore lumen and penetration of the single-stranded hairpin end into the internal pore structure, resulting in a near-complete current cutoff. Reversing the electric field to -50mV causes the pore electrophoresis to be uncut. A new +50mV episode causes DNA hairpin lodging / penetration and pore cutoff again. On the +50mV segment, unfolding of the hairpin structure can lead to the termination of the current cutoff, as indicated by the reversal of the current cutoff. A hairpin having the sequence 3' GCGGGGA GCGTATT AGAGTTG GATCGGATGCA GCTGGCTACTGACGTCA TGACGTCAGTAGCCAGCATGCATCCGATC-5' was added to the cis side of the small chamber at a final concentration of 10 nM. [Figure 57]This figure shows the purification and channel characteristics of CsgG-ΔPYPA, a mutant CsgG pore in which the constricting residue Y51 of the PYPA sequence (residues 50-53) is mutated to GG. [Modes for carrying out the invention]
[0032] Explanation of the sequence list Sequence ID 1 is the amino acid sequence of wild-type Escherichia coli (E. coli) CsgG, including the signal sequence. (Uniprot deposit number P0AEA2)
[0033] Sequence ID 2 is a polynucleotide of wild-type Escherichia coli (E. coli) CsgG containing a signal sequence. The cydone sequence (Gene ID: 12932538) is shown.
[0034] Sequence ID 3 is the mycelium of wild-type Escherichia coli (E. coli) CsgG at positions 53-77 of Sequence ID 2. It shows an acid sequence. This is the mature wild-type Escherichia coli (E. coli) CsgG monomer (i.e., This corresponds to the amino acid sequence between positions 38 and 63 (which lacks a signal sequence).
[0035] Sequence IDs 4-388 are modified monomers of CsgG lacking a signal sequence, number 38-6. The amino acid sequence at position 3 is shown.
[0036] Sequence ID 389 is a substrain of Escherchia coli K-12 strain MC4100. Codon-optimized polynucleotide encoding the wild-type CsgG monomer from substr. MC4100. The ocidal sequence is shown. This monomer lacks a signal sequence.
[0037] Sequence ID 390 is a substrain of Escherchia coli K-12 strain MC4100. This shows the amino acid sequence of the mature form of the wild-type CsgG monomer from substr. MC4100). The monomer lacks a signal sequence. The abbreviation used for this CsgG is CsgG-E. co.
[0038] Sequence ID 391 is YP_001453594.1: Virtual protein CKO_0203 2 [Citrobacter koseri (ATCC BAA-895)] It shows an amino acid sequence of approximately 248, which is 99% identical to sequence number 390.
[0039] Sequence ID 392 is WP_001787128.1: Partial, Curli production aggregate ( curli production assembly) / transport component CsgG [Salmonella enterica This shows the amino acid sequence of 16-238 of ), which is 98% identical to sequence number 390.
[0040] Sequence ID 393 is KEY44978.1|:curli production aggregate / transport protein. CsgG [Citrobacter amalonaticus] 16~ It shows a 277-amino acid sequence, which is 98% identical to sequence number 390.
[0041] Sequence ID 394 is YP_003364699.1:curli production aggregate / transport component. [Citrobacter rodentium (ICC168)] 16~ It shows a 277-amino acid sequence, which is 97% identical to sequence number 390.
[0042] Sequence ID 395 is YP_004828099.1:curli production aggregate / transport component. [Enterobacter asburiae LF7a] 16-277 The amino acid sequence is shown, which is 94% identical to sequence number 390.
[0043] Sequence ID 396 is a polynucleotide sequence encoding Phi29 DNA polymerase. Show the columns.
[0044] Sequence ID 397 shows the amino acid sequence of Phi29 DNA polymerase.
[0045] Sequence ID 398 is a codon-optimized gene derived from the sbcB gene of Escherichia coli (E. coli). This shows a polynucleotide sequence. This is exonuclease I from Escherichia coli (E. coli). It codes for the enzyme (EcoExo I).
[0046] Sequence ID 399 is an exonuclease I enzyme (EcoEx) derived from Escherichia coli (E. coli). The amino acid sequence of o I) is shown.
[0047] Sequence ID 400 is a codon-optimized gene derived from the xthA gene of Escherichia coli (E. coli). This shows a polynucleotide sequence. This is exonuclease I from Escherichia coli (E. coli). Enzyme II is encoded.
[0048] Sequence ID 401 is an amino acid of the exonuclease III enzyme from Escherichia coli (E. coli). The acid sequence is shown. This enzyme extracts 5' monophosphate from one strand of double-stranded DNA (dsDNA). The enzyme performs partitioned digestion of nucleosides in the 3'-5' direction. Enzyme initiation on the chain occurs approximately 4 times. A 5' overhang is required for the cleotide.
[0049] Sequence ID 402 is derived from the recJ gene of the hyperthermophilic bacterium *T. thermophilus*. This shows a codon-optimized polynucleotide sequence. This is from a hyperthermophilic bacterium (T. thermophilus) or This encodes the RecJ enzyme (TthRecJ-cd).
[0050] Sequence ID 403 is a RecJ enzyme (TthRe) from the hyperthermophilic bacterium T. thermophilus. The amino acid sequence of cJ-cd is shown. This enzyme extracts the 5' monophosphate nucleus from ssDNA. The oside undergoes forward digestion in the 5'-3' direction. Enzyme initiation on the chain requires at least 4 nuclei. Ochido is needed.
[0051] Sequence ID 404 is derived from the bacteriophage lambdaexo (redX) gene. This shows the Donn-optimized polynucleotide sequence. This is the bacteriophage lambda exonucleotide. Codes rease.
[0052] Sequence ID 405 shows the amino acid sequence of bacteriophage lambda exonuclease. This array is one of three identical subunits that come together to form a trimmer. This enzyme performs high forward-moving nucleotide removal in the 5'-3' direction from one strand of dsDNA. Perform the modification (http: / / www.neb.com / nebecomm / product (s / productM0262.asp). The enzyme initiation on the chain has a 5' phosphate group. It preferentially requires the 5' overhangs of the other four nucleotides.
[0053] Sequence ID 406 shows the amino acid sequence of Hel308 Mbu.
[0054] Sequence ID 407 shows the amino acid sequence of Hel308 Csy.
[0055] Sequence ID 408 shows the amino acid sequence of Hel308 Tga.
[0056] Sequence ID 409 shows the amino acid sequence of Hel308 Mhu.
[0057] Sequence ID 410 shows the amino acid sequence of Tral Eco.
[0058] Sequence ID 411 shows the amino acid sequence of XPD Mbu.
[0059] Sequence ID 412 shows the amino acid sequence of Dda 1993.
[0060] Sequence ID 413 shows the amino acid sequence of Trwc Cba.
[0061] Sequence ID 414 is WP_006819418.1: Transporter [Yokenella lengesburg] The amino acid sequence of Yokenella regensburgei (gay) is shown from 19 to 280, and this is the sequence number. It is 91% identical to item number 390.
[0062] Sequence ID 415 is WP_024556654.1:curli production aggregate / transport tank. Protein CsgG [Cronobacter pulveris] 16-27 The amino acid sequence of 7 is shown, which is 89% identical to sequence number 390.
[0063] Sequence ID 416 is YP_005400916.1:curli production aggregate / transport tank Protein CsgG [Rahnella aquatilis HX2] 16-2 It shows a sequence of 77 amino acids, which is 84% identical to sequence number 390.
[0064] Sequence ID 417 is KFC99297.1:CsgG family curli production aggregate. / Transport component [Kluyvera ascorbata ATCC33433] The amino acid sequence of 20-278 is shown, which is 82% identical to sequence number 390.
[0065] Sequence ID 418 is a curli production set of the KFC86716.1|:CsgG family. Body / transport component [Hafnia alvei ATCC13337] 16-27 The amino acid sequence of 4 is shown, which is 81% identical to sequence number 390.
[0066] Sequence ID 419 is related to the formation of YP_007340845.1|:curli polymers. The contributing uncharacterized protein [Enterobacteriaceae bacterial strain FG] The amino acid sequence of I 57] shows amino acids 16-270, which is 76% identical to sequence number 390. be.
[0067] Sequence ID 420 is WP_010861740.1:curli production aggregate / transport tank. Protein CsgG [Plesiomonas shigelloides] 17 It shows an amino acid sequence of approximately 274, which is 70% identical to sequence number 390.
[0068] Sequence ID 421 is YP_205788.1:curli-producing aggregate / transport outer membrane lipota 23% of the protein component CsgG [Vibrio fischeri ES114] It shows an amino acid sequence of approximately 270, which is 60% identical to sequence number 390.
[0069] Sequence ID 422 is WP_017023479.1:curli-producing aggregate protein. CsgG [Aliivibrio logei] contains 23-270 amino acids The column is shown, and it is 59% identical to sequence number 390.
[0070] Sequence ID 423 is WP_007470398.1: Curli production aggregate / transport component. CsgG [Photobacterium sp. AK15] 22-275 The amino acid sequence is shown, which is 57% identical to sequence number 390.
[0071] Sequence ID 424 is WP_021231638.1:curli-producing aggregate protein. The amino acid sequence of CsgG [Aeromonas veronii] from 17 to 277 This indicates that it is 56% identical to sequence number 390.
[0072] Sequence ID 425 is WP_033538267.1:curli production aggregate / transport tank. Protein CsgG [Shewanella sp. ECSMB14101] 27-2 It shows a 65-amino acid sequence, which is 56% identical to sequence number 390.
[0073] Sequence ID 426 is WP_003247972.1:curli-producing aggregate protein. This shows the amino acid sequence of CsgG [Pseudomonas putida] from 30 to 262. This is 54% identical to sequence number 390.
[0074] Sequence ID 427 is YP_003557438.1:curli production aggregate / transport component. CsgG [Shewanella violacea DSS12] 1-234 The amino acid sequence is shown, which is 53% identical to sequence number 390.
[0075] Sequence ID 428 is WP_027859066.1:curli production aggregate / transport tank. Protein CsgG [Marinobacterium jannaschii] The amino acid sequence of 36-280 is shown, which is 53% identical to sequence number 390.
[0076] Sequence ID 429 is CEJ70222.1:Curli production aggregate / transport component CsgG [Chryseobacterium oranimense G311] It shows an amino acid sequence from 29 to 262, which is 50% identical to sequence number 390.
[0077] Sequence ID 430 shows the polynucleotide sequence used in Example 18.
[0078] Sequence ID 431 shows the polynucleotide sequence used in Example 18.
[0079] Sequence ID 432 shows the polynucleotide sequence used in Example 18.
[0080] Sequence ID 433 shows the polynucleotide sequence used in Example 18.
[0081] Sequence ID 434 shows the polynucleotide sequence used in Example 18. Sequence ID 434 Attached to the 3' end of 4 are six iSp18 spacers (two on the opposite end) It is the thymine-bound and 3' cholesterol TEG.
[0082] Sequence ID 435 shows the amino acid sequence of StrepII(C).
[0083] Sequence ID 436 shows the amino acid sequence of Pro.
[0084] Sequence IDs 437-440 show the primers from Example 1.
[0085] Sequence ID 441 shows a hairpin from Example 21.
[0086] Sequence numbers 442-448 show the sequence from Figure 4.
[0087] Sequence numbers 449 and 450 show the sequences from Figure 43.
[0088] Detailed description of the invention Unless otherwise indicated, the implementation of this invention is within the scope of the skills of those skilled in the art, including chemistry and molecular biology. It utilizes conventional techniques of science, microbiology, recombinant DNA technology, and chemical methods. The technique is described in literature, for example, MR Green, J. Sambrook, 2012, Molecular Cloning: A Labor atory Manual, Fourth Edition, Books 1-3, Cold Spring Harbor Laboratory Press, Co ld Spring Harbor, NY; Ausubel, FM et al., (1995 and periodic supplements, Cur Rent Protocols in Molecular Biology, ch. 9, 13, and 16, John Wiley & Sons, New Y ork, NY);B. Roe, J. Crabtree, and A. Kahn, 1996, DNA Isolation and Sequencin g: Essential Techniques, John Wiley & Sons;JM Polak and James O'D. McGee, 19 90, In Situ Hybridization: Principles and Practice, Oxford University Press;M. J. Gait (Editor), 1984, Oligonucleotide Synthesis: A Practical Approach, IRL Pre ss; and DMJ Lilley and JE Dahlberg, 1992, Methods of Enzymology: DNA S Structure Part A: Synthesis and Physical Analysis of DNA Methods in Enzymology, A This is also explained in academic press. Each of these general texts refers to this specification by reference. It is incorporated into.
[0089] Before describing the present invention, we provide a number of definitions that will help in understanding the present invention. All references mentioned in the footnotes are incorporated herein by reference in their entirety. Unless otherwise defined, all specialized and scientific terms used herein are defined as follows: This has the same meaning as generally understood by those skilled in the art in which the present invention pertains.
[0090] As used herein, the term "comprising" means any of the elements being referred to. It must include (include) (or), and may include (include) other elements depending on the case. To taste. "To essentially become from ~" must include all the elements being mentioned and the elements being listed. We will exclude elements that would substantially affect the basic and new features, and where applicable, It means that it may include other elements. "~consists of" means that it may include all elements other than those listed. This means excluding the element. The embodiments defined by each of these terms are This falls within the scope of the present invention.
[0091] As used herein, the term "nucleic acid" means that each nucleotide has a 3' and 5' end. Nucleotides linked by sphodiester bonds, single-stranded or double-stranded covalent bonds. It is a combined sequence. Polynucleotides are composed of deoxyribonucleotide bases. It can also be composed of ribonucleotide bases. As for nucleic acids, DN Examples include A and RNA. Nucleic acids are synthesized in vitro. They may also be isolated from natural sources. As nucleic acids, they may be modified DNA or RNA, for example, methylated DNA or RNA, or post-translational modifications, for example, 7 - 3' processes such as 5' capping, cleavage, and polyadenylation with methylguanosine. RNA that has been spliced or fused can be further listed. Nucleic acids and For example, synthetic nucleic acids (XNAs), such as hexitol nucleic acids (HNAs), cyclohexene nucleic acids. Acid (CeNA), threose nucleic acid (TNA), glycerol nucleic acid (GNA), locked nucleus Acids (LNA) and peptide nucleic acids (PNA) can also be mentioned. In this specification, "poly The size of nucleic acids, also called nucleotides, is usually measured in terms of double-stranded polynucleotides. n is expressed as the number of base pairs (bp), or in the case of a single-stranded polynucleotide, as the number of nucleotides (n). It is expressed as the number of t). 1000bp or nt is equal to kilobases (kb). Polynucleotides with fewer than approximately 40 nucleotides are usually called "oligonucleotides." This is a type of DNA used for manipulation such as polymerase chain reaction (PCR). It may also include Rhymer.
[0092] In relation to this invention, the term "amino acid" is used in its broadest sense, and refers to naturally occurring amino acids. It is intended to contain L α-amino chains or residues. Naturally occurring amino acids The commonly used one- and three-letter abbreviations for ano acids are also used herein: A = Ala , C=Cys, D=Asp, E=Glu, F=Phe, G=Gly, H=His, l=l le, K=Lys, L=Leu, M=Met, N=Asn, P=Pro, Q=Gln, R =Arg, S=Ser, T=Thr, V=Val, W=Trp, and Y=Tyr(Lehn inger, A. L, (1975) Biochemistry, 2d ed., pp. 71-92, Worth Publishers, New York. ). The general term "amino acid" also includes D-amino acids and retro-inversoamino acids. Of course, chemically modified amino acids, such as amino acid analogs, are usually incorporated into proteins. Naturally occurring amino acids that are not included, such as norleucine, and amino acids specific to amino acids A chemically synthesized compound having certain properties known in the art, for example, β -Also includes amino acids. For example, the same conformation as natural Phe or Pro for peptide compounds. Phenylalanine or proline analogs or mimics that enable restriction are amino acids It is included in the definition of such analogs and mimics. In this specification, the respective amino acids This is called a "functional equivalent." Another example of amino acids is found in Roberts and Vellaccio's *The Peptides*. Analysis, Synthesis, Biology, Gross and Meiehofer, eds., Vol. 5 p. 341, Academy is included in Press, Inc., N.Y. 1983, which is hereby incorporated by reference into this specification and is incorporated herein by reference.
[0093] A "polypeptide" is a polymer of amino acid residues linked by peptide bonds, whether produced naturally or in vitro by synthetic means. A polypeptide having a length of less than about 12 amino acid residues is usually referred to as a "peptide", and one having a length of about 12 to about 30 amino acid residues may be referred to as an "oligopeptide". As used herein, the term "poly peptide" refers to a naturally occurring polypeptide, a precursor form or a product of a proprotein. A polypeptide can undergo mutations or post-translational modification processes, such as, but not limited to, glyco sylation, proteolytic cleavage, lipidation, signal peptide cleavage, propeptide cleavage, phosphorylation and the like. As used herein, the term "protein" refers to a macromolecule containing one or more polypeptide chains and is used herein to refer to a macromolecule containing one or more polypeptide chains. A "biopore" is a transmembrane protein structure that defines a channel or pore that allows the transfer of molecules and ions from one side of a membrane to the other. The passage of ionic species through this pore can be driven by a difference in the electrical potential applied across both sides of the pore. A "nanopore" [[ID=At least 50%, 60%, 70%, and 80% of the wild-type Escherichia coli (E. coli) CsgG were found to be present in the same sample. It may contain polynucleotides having %, 90%, 95%, or 99% complete sequence identity. Yes, it is possible. Similarly, polypeptides are also found in wild-type Escherichia coli (E. coli) Cs, as shown in Sequence ID No. 1. At least 50%, 60%, 70%, 80%, 90%, 95%, or 99% complete with γG. The polypeptide may contain a polypeptide with sequence identity. The polypeptide is a CsgG-like tan. It contains a polypeptide that includes the PFAM domain PF03783, which is a characteristic of the protein. This is possible. A list of currently known CsgG homologs and CsgG structures can be found at http: / / p It can be found at fam.xfam.Org / / family / PF03783. For example, sequence identity refers to a fragment or portion of a full-length polynucleotide or polypeptide. It may also be sequence identity with the sequence of the present invention. Even if they only share the same identity, a particular area, domain, or subunit is a distinct entity. It is possible that the sequence may share 80%, 90%, or even 99% sequence identity with the original sequence. According to the present invention, homology with the nucleic acid sequence of Sequence ID No. 2 is not limited to mere sequence identity. Many nucleic acid sequences, despite seemingly having low sequence identity, are biologically significant to one another. It can exhibit homology. In this invention, homologous nucleic acid sequences are subjected to low stringency conditions. It is thought that they will hybridize with each other below (MR Green, J. Sambro ok, 2012, Molecular Cloning: A Laboratory Manual, Fourth Edition, Books 1-3, Col d Spring Harbor Laboratory Press, Cold Spring Harbor, NY).
[0096] The term "vector" refers to the process of incorporating another nucleic acid (usually DNA) sequence fragment of an appropriate size. It is used to indicate DNA molecules that are linear or circular in length. NA fragments result in the transcription of the gene encoded by their DNA sequence fragment. It may include segments such as promoters and transcriptional terminators. Transistor, enhancer, intra-sequence ribosome entry site, untranslated region, polyadenylated signal This may include, but is not limited to, elements such as the origin of replication, selection markers, and replication origins. Various suitable promoters for physical hosts (for example, for E. coli, [be [-Lactamase and lactose promoter system, alkaline phosphatase, tri (Protophan (trp) promoter system, lac, tac, T3, T7 promoter) and various suitable promoters for eukaryotic hosts (e.g., Simianvirus 40 early or late promoter, Roussarcoma virus terminal repeat sequence promoter, cytomegalo (Viral promoter, adenovirus late promoter, EG-1a promoter) Expression vectors are available in plasmid, cosmid, viral, and yeast forms. Often derived from artificial chromosomes, vectors contain DNA sequences from several sources. These are often recombinant molecules. The specific embodiments of the present invention are wild-type molecules as described herein. The expression vector provides encoding a modified CsgG polypeptide. "Linked" refers to the DNA sequence, for example, the DNA linkage in nucleic acid vectors such as those mentioned above. When applied to a column, the array functions cooperatively to achieve their intended purposes (i.e., to initiate transcription that proceeds through the relevant coding array to the termination array). This indicates that the arrays are arranged to function cooperatively.
[0097] The transmembrane protein structure of the biological pore can actually be monomeric or oligomeric. Typically, the pore is composed of multiple polypeptide subunits arranged around a central axis, resulting in a protein-lined channel that extends substantially perpendicular to the membrane in which the nanopore residues are present. The number of polypeptide subunits is not limited. Usually, the number of subunits ranges from 5 to 30, preferably, the number of subunits is 6 - 10. Alternatively, the number of subunits is not defined, as in the case of perfringolysin or related large membrane pores. The portion of the protein subunits within the nanopore that forms the protein-lined channel usually contains a secondary structure that may include one or more transmembrane β-barrels and / or α-helix portions.
[0098] It should be understood that the various applications of the disclosed products and methods can be adjusted to meet the specific requirements in the art. It should also be understood that the terminology used herein is for the purpose of describing particular embodiments of the invention only and is not intended to be limiting.
[0099] In addition, as used in this specification and the appended claims, the singular forms "a", "an", and "the" are to be construed to include the plural forms as well, unless the context clearly dictates otherwise. As long as it is included, it also includes multiple references. Therefore, for example, "polynucleotide (a polynucleol References to "eotide" include two or more polynucleotides, and "polynucleotide bonded The reference to "polynucleotide binding protein" indicates two or more such proteins. It contains protein, and a reference to "helicase" means it contains two or more helicases. Furthermore, a reference to "monomer" includes two or more monomers, and is a reference to "por e) A reference to this includes two or more pors, etc.
[0100] All publications, patents, and patent applications cited herein, whether in the preceding or following text, The wishes are incorporated herein by reference in their entirety.
[0101] The present invention, in part, relates to bacterial amyloid secretion channels (CsgG), methods for producing the same, and Regarding its use in nucleic acid sequencing applications and molecular sensing.
[0102] CsgG is a membrane lipoprotein (Uniprot) present in the outer membrane of Escherichia coli (E. coli). Deposit number P0AEA2; Gene ID: 12932538). This outer lipid membrane In this context, CsgG contains a nano-sized oligomeric complex of nine CsgG monomer subunits. It forms a pore. Due to the type II (lipoprotein) signal sequence, the CsgG protein The precursor is transferred via the SEC translocon, and then matured CsgG (i.e., Triacetylation occurs at the N-terminal Cys residue of CsgG (where the type II signal sequence is cleaved). Triacetylated, or "lipidized," CsgG is used in Gram-negative hosts. It is transported to the outer membrane, where it enters its bilayer as a nonamepore. It is in its non-lipidized form. CsgG, for example CsgG C1S , is a soluble protein in the prepore three-dimensional structure. It is located in the liplasm (Figure 42).
[0103] X-ray structure of wild-type CsgG nanopores (Goyal et al., Nature, 2014, 516(7530), 250-3 ) indicates that it has a width of 120 Å and a height of 85 Å (hereafter referred to herein, The term "width" of a nanopore refers to its dimension parallel to the film surface, and the nanopore The term "height" refers to its dimension perpendicular to the membrane. CsgG pore The complex transmembrane via 36-strand β-barrels, creating a channel with an inner diameter of 40 Å. Figure 1). Each monomer in the CsgG channel assembly forms a conserved 12-residue loop (C loop). It has a loop called "CL" (Figure 2), and this loop works together to form a channel with a diameter of approximately 9.0 Å. This creates a narrowing of the CsgG nanopore (Figures 1 and 2). The narrowing of the wild-type CsgG nanopore is due to the CsgG oligomer. The amino acid residues Tyr51, Asn55, and of each CsgG monomer present in the sesame oil It is made up of three stacked concentric rings formed by the side chains of Phe56 (Figure 3). This numbering of residues is due to the absence of the N-terminal native 15-amino acid signal sequence. It is based on mature proteins. Therefore, mature proteins are based on residues 16-2 of SEQ ID NO: 1. Corresponds to 77. Tyr51 is at position 66 of sequence number 1, and Asn is at position 7 of sequence number 1. Phe56 is at position 0, and Phe56 is at position 71 of sequence number 1.
[0104] Constriction acts to restrict the passage of ions and other molecules through the CsgG channel. Single-channel current recording of reconstituted CsgG within a planar phospholipid bilayer was performed under standard electrolyte conditions. And using potentials of +50mV or -50mV, 43.1±4.5pA, respectively. The steady current was either (n=33) or -45.1±4.0 pA (n=13) (Figure 5).
[0105] The current flow through the CsgG channel is due to the stoichiometric amount of the periplasmic factor CsgE(Un It can be effectively blocked by adding iprot accession number POAE95 (Example) 10-12). Although we do not wish to be constrained by theory, traces of current are found in Cs gE forms a complex with the CsgG pore and acts to cap one end of the channel. This strongly suggests a mechanism by which the flow of ions through the CsgG channel is significantly reduced, compared to standard single ions. Measurements can be taken using single-channel recording technology (Examples 12 and 13, Figure 6). The inventors have developed a method for measuring current flow parameters (maximum current and current fluctuations) The ability to perform nanopore sequencing and separation by an embodiment of the present invention We found that it is suitable for use in sub-sensing applications.
[0106] Therefore, the present invention, in part, addresses the variation in the electrical measurement of the current flowing through the nanopore. Nucleic acid sequencing methods based on and Cs in such nucleic acid sequencing Regarding the use of gG nanopore protein complexes.
[0107] Nucleic acids are particularly suitable for nanopore sequencing. DNA and RNA contain naturally occurring nucleic acids. Nucleic acid bases can be distinguished by their physical size. Alternatively, when individual bases pass through the nanopore channel, the size difference between those bases causes This results in a reduction that is directly correlated with the ion flow through the channel. Recording the fluctuations in the ion flow This can be done. Suitable electrical measurement techniques for recording ion flow fluctuations include, for example, International Publication No. 20 Pamphlet No. 00 / 28312 and D. Stoddart et al., Proc. Natl. Acad. Sci., 2010, 106, pp 7702-7 (Single-channel recording equipment), and, for example, International Publication No. 2009 It is described in pamphlet No. 077734 (Multi-channel recording technology). For proper calibration... Furthermore, using characteristic reduction of ion flow, specific nucleotides and Related bases can be identified in real time.
[0108] The narrowest constriction size of a transmembrane channel is generally suitable for nucleic acid sequencing applications. This is an important factor in determining the suitability of the nanopore. If the stenosis is too small, the sequence The molecules being sung will not be able to pass through. However, the ion flow through the channel To achieve maximum effect, the channel must be large at its narrowest point (i.e., at the point of constriction). It should not be excessive. Ideally, each constriction should be as small as possible in diameter to the size of the base through which it passes. It should be as close as possible. The diameter of the constriction suitable for sequencing nucleic acids and nucleic acid bases is Nanometer range (10 -9 (in the meter range). Preferably, the diameter is approximately 0.5 to 1 It should be 0.5 nm, and typically the diameter is approximately 0.7-1.2 nm. The constriction within *E. coli* CsgG has a diameter of approximately 9 Å (0.9 nm). The inventors found that the size and steric arrangement of the stenosis within the CsgG channel are related to nucleic acid sequencing. It was deduced that it is suitable for [the purpose].
[0109] CsgG nanopores for nucleic acid sequencing applications are used in their wild-type form. This may also be done to further improve the desired properties of the nanopores in use, for example Furthermore, it may be further modified by directional mutagenesis of specific amino acid residues. For example, in embodiments of the present invention, the mutations include the number, size, shape, and arrangement of stenosis within the channel. This is intended to change the orientation. The modified mutant CsgG nanopore complex is This results in the insertion, substitution, and / or deletion of specific target amino acid residues within a lipeptide sequence. It can be prepared by known genetic engineering techniques. In the case of oligomer CsgG nanopores. Mutations may occur in each polypeptide subunit of the monomer, or This can occur in any one of the monomers, or in all of the monomers. This can also be caused by... Preferably, in one embodiment of the present invention, the described mutation is made into an oligomer This occurs in all monomer polypeptides within the protein structure.
[0110] According to embodiments of the present invention, modified mutant Cs have a reduced number of channel constrictions within the pore. gG nanopores are provided.
[0111] The wild-type Escherichia coli (E. coli) CsgG pore contains two channel constrictions (see Figure 1). (To be done). These are part of a broader structure that includes an additional amino acid from position 54 to 53. (i) amino acid residues Phe56 and Asn55 and (ii) amino acid residue Ty r51, and the C-loop motif are formed (Figures 2 and 3).
[0112] In typical nanopore nucleic acid sequencing, the individual nucleotides of the target nucleic acid sequence are... As the nucleotides pass through the nopore channels sequentially, partial occlusion of the channels occurs due to nucleotides. Therefore, the ion flow in the open channel is reduced. Using the preferred recording techniques described above, ions This reduction in flow is measured. The reduction in ion flow is measured for known nucleotides passing through the channel. This can be used to correspond to the reduction in ion flow measured by, which consequently affects which nucleotide A means of determining whether it is passing through the channel, therefore, if done sequentially, This method determines the nucleotide sequence of nucleic acids passing through the nopore. For accurate determination, reduce the ion flow through the channel to a single narrowing (i.e., "reading" It is necessary to directly correlate the size of each nucleotide passing through the "head". This has become common. The pore is "passed through" by, for example, the action of related polymerases. I understand that sequencing can be performed on intact nucleic acid polymers. It will be done. Alternatively, nuclei that are sequentially removed from the target nucleic acid adjacent to the pore. The sequence can be determined by passing othido triphosphate (for example, International Publication No. 201) Please refer to the 4 / 187924 pamphlet.
[0113] When there are two or more narrowings that are separated from each other, the narrowings can interact with each other. In this situation, The reduction in ion flow through the channel is due to the flow at all constricted points containing nucleotides. This is the result of the overall limiting. Therefore, in some cases, double stenosis is a single combined current signal. It can also lead to problems. In certain situations, a narrowing in one place, that is, one "reading" If there are two such reading heads, the current readout will be performed individually. It may not be possible to make a determination.
[0114] The wild-type pore structure of CsgG is modified by recombinant gene technology to widen one of its two constrictions. Rework the channel to modify or remove a single narrow point within the channel. By leaving a narrowing, a single reading head can be defined. CsgG oligomer The pore constriction motif is within the polypeptide of wild-type E. coli (CsgG) monomer. It is located at amino acid residues 38-63. The wild-type amino acid sequence in this region is sequence number 3. It is provided as such. When considering this region, amino acid residue positions 50-53 and 54-56 It is believed that any of the mutations in 58-59 are intended to fall within the scope of the present invention. Based on sequence similarity with CsgG homologues (Figure 4), amino acid residue positions 38-49, Numbers 53, 57, and 61-63 are considered highly conserved and therefore, substitution or It is not well suited to other modifications. Tyr51, As in the channel of the wild-type CsgG structure. Due to the critical positioning of the side chains n55 and Phe56, mutations at these positions are significant. This can be advantageous for modifying or altering the characteristics of the reading head.
[0115] A mutation at a given position in the CsgG protein monomer is equivalent to the wild-type amino acid at that position. This may result in substitution with any other natural or non-natural amino acid. In the embodiment, it is desirable to widen or eliminate the stenosis, preferably modified CsgG The amino acid side chains in the protein are sterically hindered compared to the amino acid side chains in the wild-type structure that they substitute for. The selection will be made to minimize harm. Substitute amino acid residues at a given position will be similar It may have electrostatic properties, or it may have different electrostatic properties. Preferably, The replacement amino acid side chain is used to minimize disruption to the secondary structure or properties of the channel. This will result in the amino acid side chain having the same electrostatic charge as the amino acid side chain in the wild-type structure that is substituted.
[0116] The selection of substitution amino acids is based on large-scale multi-sequence alignment, where the amino acid is found to be a different amino acid. Based on the BLOSUM62 matrix, this provides a standard methodology for calculating the likelihood of substituting no acids. This can also be the case. An example of the BLOSUM62 matrix is freely available to those skilled in the art on the internet. Yes, it is possible. For example, the National Center for Biotechnology Information (NBI) The NCBI website (http: / / www.ncbi.nlm.nih.gov) See / Class / Structure / aa / aa_explorer.cgi I want to.
[0117] T in the wild-type CsgG structure acts to form a first stenosis within the CsgG channel. A substitution of any amino acid is made to yr51. In particular, in a specific embodiment of the present invention Morphologically, Tyr51 is composed of alanine, glycine, valine, leucine, isoleucine, and as It may be substituted with paragine, glutamine, and phenylalanine (SEQ ID NO: 39~) 318). In embodiments of the present invention, substitution of Tyr51 with alanine or glycine is particularly This is preferable (SEQ ID NOs: 39-108). In the embodiment, residues 50-53 (in the wild-type sequence) PYPA) can be substituted with glycine-glycine (GG) (SEQ ID NO: 354~ 388).
[0118] Asn55 contributes to the second narrowing within the CsgG channel, with any amino acid Substitution is performed. In particular, in certain embodiments of the present invention, Asn55 is alanine, gu It may be substituted with lysine, valine, serine, or threonine (SEQ ID NOs: 9-33) 44-68, 79-103, 114-138, 149-173, 184-208, 219 (~243, 254~278, 289~313 and 324~348)
[0119] Any amino acid is used to form part of the second stenosis within the CsgG channel for Phe56. Acid substitution is performed. In particular, in certain embodiments of the present invention, Phe56 is Alani Substituted with glycine, valine, leucine, isoleucine, asparagine, and glutamine. This can happen (sequences 5-13, 15-18, 20-23, 25-28, 30-3) 3, 40-43, 45-48, 50-53, 55-58, 60-63, 65-68, 75 ~78, 80~83, 85~88, 90~93, 95~98, 100~103, 110~ 113, 115-118, 120-123, 125-128, 130-133, 135- 138, 145~148, 150~153, 155~158, 160~163, 165~ 168, 170-173, 180-183, 185-188, 190-193, 195- 198, 200~203, 205~208, 215~218, 220~223, 225~ 228, 230~233, 235~238, 240~243, 250~253, 255~ 258, 260~263, 265~268, 270~273, 275~578, 285~ 288, 290~293, 295~298, 300~303, 305~308, 310~ 313, 320~323, 325~328, 330~333, 335~338, 340~ 343 and 345-348). In embodiments of the present invention, P in alanine and glycine Substitution of he56 is particularly preferred (SEQ ID NOs: 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 105, 110, 115, 120, 125, 130, 135, 140, 145, 150, 155, 160, 165, 170, 175, 180, 185, 190, 195, 200, 205, 210, 215, 220, 225, 230, 235, 240, 245, 250, 255, 260, 265, 270, 275, 280, 285, 290, 295, 300, 305, 310, 315, 320, 325, 330, 335, 340, 345, 350, 355, 360, 365, 370, 375, 380, 385, 6, 11, 16, 21, 2 6, 31, 36, 41, 46, 51, 56, 61, 66, 71, 76, 81, 86, 91 , 96, 101, 106, 111, 116, 121, 126, 131, 136, 141, 146, 151, 156, 161, 166, 171, 176, 181, 186, 191, 196, 201, 206, 211, 216, 221, 226, 231, 236, 241, 246, 251, 256, 261, 266, 271, 276, 281, 286, 291, 296, 301, 306, 311, 316, 321, 326, 331, 336, 341 (346, 351, 356, 361, 366, 371, 376, 381 and 386).
[0120] In a given mutant CsgG protein, the substitution at Tyr51 is at positions 55 and 56. It may also occur simultaneously with one of them (Sequence keys 44, 49, 54, 59, 64, 7) 9, 84, 89, 94, 99, 114, 119, 124, 129, 134, 149, 15 4, 159, 164, 169, 184, 189, 194, 199, 204, 219, 22 4, 229, 234, 239, 254, 259, 264, 269, 274, 289, 29 4, 299, 304, 309, 359, 364, 369, 374, 379, 40, 41, 42, 75, 76, 77, 110, 111, 112, 145, 146, 147, 180, 181, 182, 215, 216, 217, 250, 251, 252, 285, 286 (and 287), in this case, at least one narrowing of a suitable dimension within the channel is maintained. Alternatively, the substitution of Tyr51 is performed at positions Asn55 and Phe56. They are mutually exclusive with respect to exchange (Sequence codes 39, 74, 109, 144, 179, 214, 249, 284, 354, 10, 11, 12, 15, 16, 17, 20, 21, 22, 2 5, 26, 27, 30, 31 and 32).
[0121] Alternatively, Tyr51, Asn55, or Phe56 in the wild-type CsgG protein One or more may be missing (sequence numbers 319-353, 34-38, 69- 73, 104-108, 139-143, 174-178, 209-213, 244-2 48, 279-283, 314-318, 384-388, 8, 13, 18, 23, 28 , 33, 43, 48, 53, 58, 63, 68, 78, 83, 88, 93, 98, 103 , 113, 118, 123, 128, 133, 138, 148, 153, 158, 163 , 168, 173, 183, 188, 193, 198, 203, 208, 218, 223 ,228,233,238,243,253,258,263,268,273,278 , 288, 293, 298, 303, 308 and 313). At least one in the channel In order to maintain the narrowing of the site, in a given embodiment, the deletion of amino acid residue Tyr51 is The deletions of both amino acid residues Asn55 and Phe56 are mutually exclusive (sequence number) No. 319~322, 324~327, 329~332, 334~337, 339~342 , 344~347, 38, 73, 108, 143, 178, 213, 248, 283 (318) Certain adjacent amino acid residues at positions 53 and 54 and 48 and 49. These may also be missing.
[0122] The present invention allows the above modifications to be performed in isolation or in any combination. It should be understood that this provides a certain embodiment.
[0123] Either stenosis at Tyr51 or stenosis at Asn55 / Tyr56 As a result of removal, the stenosis within the CsgG channel becomes a single location. This is constrained by theory. While not desirable, stenosis in Asn55 / Tyr56 can sometimes be desirable. It is assumed that this will result in greater conformational stability than the stenosis in yr51. However, A sn55 / Tyr56 stenosis is higher compared to nucleotides (measured along the central pore axis). Sometimes this can be excessive. This can lead to poor resolution of individual base pairs within the DNA strand during migration. Sometimes.
[0124] The opposite is likely true for Tyr51 stenosis. Asn55 / Tyr56 stenosis After removal of the constriction, the remaining ring of the Tyr51 residue in the oligomer is in the native structure. Conformational stability may be lower. However, Tyr51 stenosis is (along the central pore axis) (By measuring this, it is possible to create a short, narrow constriction within the channel that can distinguish individual bases.) It is highly likely.
[0125] In both embodiments, there is the presence of a single narrow constriction (measured along the central axis). This reduces the complexity of current readings when using pores for nucleic acid sequencing applications. It is highly likely that the observed current generated while nucleic acids are moving through the pore is Modulation involves the passage of a single narrowing of a distinct nucleotide, i.e., a "read head". This will be reflected in the data.
[0126] Effective removal of a single stenosis also increases the opening channel current of that pore variant. This is possible. An increase in the aperture channel current results in higher background conductance. This is advantageous, and leads to better degradation of the current cutoff level for different nucleic acid base pair signals. Connected. Thus, modifications to the reading head are used for nucleic acid sequencing and other This can improve the suitability of biopores for both molecular sensing applications.
[0127] As an alternative embodiment, or in addition to the above sequence modification, Asn55 / Phe56 stenosis is The condition is that it can be further adapted to adjust the height (when measured along the central pore axis). As such, further adaptive forms of Asn55 / Phe56 stenosis are CsgG channel It may or may not be accompanied by mutations at Tyr51 or other locations within the chromosome. Preferably, a further adaptation of Asn55 / Phe56 stenosis is performed by the Tyr51 residue. It is intended as part of a mutation to widen or eliminate the narrowing that forms therein.
[0128] In the wild-type morphology, Asn55 / Phe56 channel constrictions are adjacent to each other in a perpendicular direction. It is made up of two amino acid rings located in that position. As a result, the constriction is longer than 1 nm. It possesses a 1 nm length constriction, which allows electrical signals generated from ion flow to be transferred to other parts of the nucleic acid chain during migration. It is not possible to break it down into individual bases. Normally, nucleic acid sequencing is The known nanopores used generally have a length of less than 1 nm. For example, DNA cells The MspA nanopore used for quenching is 0. It has a constriction height of 6 nm (Manrao et al., Nature Biotechnology, 2012, 30(4), 349-35 3).
[0129] Reduce the height of the Asn55 / Phe56 stenosis (measured along the central pore axis). To achieve this, one of the two residues is substituted or deleted to form the apex or base of the stenosis. You may expand it.
[0130] Asn55 forms part of the second stenosis within the CsgG channel, any amino Acid substitution is attempted, particularly with alanine, glycine, valine, serine, or threonine. Substitution in (sequences 9-33, 44-68, 79-103, 114-138, 149-1) 73, 184-208, 219-243, 254-278, 289-313 and 324 (~348).
[0131] Any amino acid is used to form part of the second stenosis within the CsgG channel for Phe56. Acid substitution is intended, particularly for alanine, glycine, valine, leucine, and isoleucine. , substitution with asparagine and glutamine (SEQ ID NOs. 5-13, 15-18, 20-23) , 25-28, 30-33, 40-43, 45-48, 50-53, 55-58, 60- 63, 65-68, 75-78, 80-83, 85-88, 90-93, 95-98, 1 00-103, 110-113, 115-118, 120-123, 125-128, 1 30-133, 135-138, 145-148, 150-153, 155-158, 1 60-163, 165-168, 170-173, 180-183, 185-188, 1 90-193, 195-198, 200-203, 205-208, 215-218, 2 20-223, 225-228, 230-233, 235-238, 240-243, 2 50-253, 255-258, 260-263, 265-268, 270-273, 2 75-578, 285-288, 290-293, 295-298, 300-303, 3 05-308, 310-313, 320-323, 325-328, 330-333, 3 35-338, 340-343 and 345-348). In embodiments of the present invention, Alani Substitution of Phe56 with glycine and is particularly preferred (SEQ ID NOs: 5, 10, 15, 2 0, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85 , 90, 95, 100, 105, 110, 115, 120, 125, 130, 135, 1 40, 145, 150, 155, 160, 165, 170, 175, 180, 185, 1 90, 195, 200, 205, 210, 215, 220, 225, 230, 235, 2 40, 245, 250, 255, 260, 265, 270, 275, 280, 285, 2 90, 295, 300, 305, 310, 315, 320, 325, 330, 335, 3 40, 345, 350, 355, 360, 365, 370, 375, 380, 385, 6 , 11, 16, 21, 26, 31, 36, 41, 46, 51, 56, 61, 66, 71, 76, 81, 86, 91, 96, 101, 106, 111, 116, 121, 126, 1 31, 136, 141, 146, 151, 156, 161, 166, 171, 176, 1 81, 186, 191, 196, 201, 206, 211, 216, 221, 226, 2 31, 236, 241, 246, 251, 256, 261, 266, 271, 276, 2 81, 286, 291, 296, 301, 306, 311, 316, 321, 326, 3 31, 336, 341, 346, 351, 356, 361, 366, 371, 376, 3 81 and 386).
[0132] Modifications are also intended to adjust the minimum diameter of stenosis in wild-type CsgG pores. The minimum diameter of both constrictions within the A is approximately 0.9 nm (9 Å), which is the DNA sigma. Diameter 1. For known narrowings in MspA nanopores that demonstrate usefulness for quenching. It is smaller than 2 nm (Manrao et al., Nature Biotechnology, 2012, 30(4), 349-353). The above results in residual stenosis within modified CsgG pores having a minimum diameter of 0.5 to 1.5 nm. One of the mutations would be preferable.
[0133] Any of the modifications listed above can be beneficial in terms of the hydrophilicity and charge distribution of amino acids at the constriction site. Modify to improve pass-through and non-covalent interactions with respect to the nucleic acid chain during transit to enable current reading. This can improve the visibility. Any of the mutations listed above is near channel constriction. By beneficially altering the hydrophilicity and charge distribution, the flow of electrolyte ions through the constriction is optimized. To achieve better identification between nucleotides during the passage of nucleic acid chains in transit. It's also possible.
[0134] Further changes in the surface charge distribution within the channel lumen of CsgG proteins may result in altered surface charge distribution. Modifications are intended. In one embodiment of the present invention, these modifications are performed on the nucleic acid during migration. This is done to avoid undesirable electrostatic adsorption to the flannel wall. Nucleic acids have a negative charge, Cs The γ channel lumen has a slight positive charge (Figure 1), so electrostatic interactions affect nucleic acid sequencing. It is assumed that it may pass through or interfere with the transition during singeing. Preferably, a positively charged Mino acid residues, such as lysine, histidine, and arginine, are neutral or negatively charged. By substituting with a strand, the efficiency of nucleic acid transfer through the pore and therefore the clarity of the current readout are further improved. It can be improved.
[0135] To facilitate the movement or passage of nucleic acid chains (or individual nucleotides) within pore constrictions. Furthermore, modification of the channel lumen in wild-type E. coli (CsgG) or mutant CsgG is also observed. The transmembrane portion of CsgG having internal constriction is central It resembles a barrel with a lid characterized by a hole. Additional lumens are located in the pores closest to the barrel and lid. Adding a loop can encourage people to pass through.
[0136] A specific embodiment of the present invention is a CsgG pore covalently bonded to one or more mono The condition is that it can consist of a mer, dimer, or oligomer. As an unrestricted example, Monomers can be genetically fused in any stereochemistry, for example, by their terminal amino acids. This can sometimes happen. In this case, the amino terminus of one monomer is the carbo terminus of another monomer. It may also be fused to the xy terminus.
[0137] According to embodiments of the present invention, the CsgG pore has properties that are beneficial for molecules to pass through that pore. It is also a condition that it can be adapted to accommodate further accessory proteins. Adaptation to the pore can facilitate the immobilization of nucleic acid processing enzymes. Examples of single enzymes include DNA or RNA polymerase, isomerase, and topoisomerase. Examples include enzymes such as gyrase, telomerase, and helicase. By associating one or more elements with nanopores, the pass-through of nucleic acids through the pores is enhanced. And, in terms of controlling the rate at which nucleic acid strands are translated by pores, It is possible (Manrao et al., Nature Biotechnology, 2012, 30(4), 349-353). It passes through the pore. For controlling the transfer rate of nucleic acid strands, reading out is preferable for measuring the current of ion flow. This has the advantage of providing a more uniform and improved response.
[0138] In embodiments of the present invention, modification in the extracellular region of the nanopore is performed by one or more bonds / By providing an immobilization site, DNA particles can be placed inside or adjacent to the lumen of the channel. It helps facilitate the docking of suitable nucleic acid processing enzymes such as remerase. It is expected to be able to stand. Suitable immobilization sites are electrostatic pads for electrostatic bonding of the enzyme. Chi, one or more cysteine residues to enable covalent coupling, This may include modifying the internal width of the transmembrane channel portion to provide a three-dimensional anchor. be.
[0139] In specific embodiments described in more detail below, the present invention is used in nucleic acid sequencing. This provides further adaptive forms of CsgG wild-type pores.
[0140] In embodiments of the present invention, the extramembrane region of the CsgG pore (the lower portion shown in Figure 1) is a channel. The extramembrane region may be shortened or removed to facilitate the exit of nucleic acid strands on the opposite side of the lumen. Shortening or removing the region improves the current resolution of the electrical signal from the ion flow in the channel. It is also possible to do so. The advantage of the latter is that it reduces the resistance to ion flow caused by the extra-membrane region. This is caused by the transmembrane channel, internal stenosis and cysts. In the case of this CsgG pore, transmembrane channels, internal stenosis and cysts are involved. The cap region corresponds to three resistance regions in series. One of these contributes to... Removing or reducing it would increase the aperture channel current. Therefore, the current resolution will be improved. Such a change is α-helix 2 (α 2) Deletion of the C-terminal α-helix (αC) (Figure 1) and / or combination thereof It is possible.
[0141] In a further embodiment of the present invention, membrane pairs on the outer surface of wild-type Escherichia coli (E. coli) CsgG pores Proactive amino acids can be modified to facilitate the insertion of pores into the membrane. In this embodiment, a single amino acid substitution is used to obtain a preferred analogue of the wild-type residue that is more hydrophobic. The condition is that it can be substituted with. For example, residues Ser136, Gly138, Gly1 40, Ala148, Ala188, or Gly202, one or more of these, Ala, Va It can be changed to l, Leu, or Ile. In addition, tyrosine or tryptophthol Aromatic residues such as ammonium compounds are positioned appropriately at the interface between the hydrophobic membrane and the hydrophilic solvent, as is typical of wild-type compounds. Mino acids can be substituted. For example, one of the residues Leu154 or Leu182. Alternatively, multiple instances can be replaced with Tyr, Phe, or Trp.
[0142] In embodiments of the present invention, the thermal stability of protein pores can be increased. As a result, the shelf life of nanopores within the sequencing device is advantageously increased. In this embodiment, the thermal stability of the protein pore is increased by modifying the beta-turn sequence. Alternatively, this can be achieved by improving electrostatic interactions on the protein surface. In one embodiment, a membrane Sharing of a β-hairpin in a penetrating region across two adjacent β-chains within or between adjacent β-hairpins. It can be stabilized by the formation of binding disulfides. Such cross-chain cysta Examples of pairings include Val139-Asp203, Gly139-Gly205, and Lys135. -Thr207, Glu201-Ala141, Gly147-Gly189, Asp1 49-Gln187, Gln151-Glu185, Thr207-Glu185, Gl y205-Gln187, Asp203-Gly189, Ala153-Lys135, Gly137-Gln151, Val139-Asp149, or Ala141-Gly It could be 147.
[0143] In a further embodiment of the present invention, the codon usage frequency of the polypeptide sequence is determined by Biotechnol. High levels of CsgG protein according to the method described in J., 2011, 6(6), 650-659 It can be modified to enable expression with a low error rate. This modification is mR It is possible to target all secondary structures of NA.
[0144] This invention also proposes modifications to the protein sequence to improve its protease stability. This provides the removal of the flexible loop region, for example, the α-helix 2 (α2), C-terminus. The deletion of the α-helix (αC) (Figure 1) and / or combinations thereof is what is accomplished. It is possible.
[0145] In embodiments of the present invention, the CsgG polypeptide sequence / expression system is a possible protein The changes will be made to avoid aggregation.
[0146] The present invention relates to polypeptide sequences for improving protein folding efficiency. Modifications are also provided. Suitable technologies are provided in Biotechnol. J., 2011, 6(6), 650-659. ru.
[0147] One embodiment of the present invention provides a bioaffinity for facilitating the purification of CsgG protein. Further provides substitution or addition of the t-tag. The published CsgG pore structure is St Contains repII tag (Goyal et al., Nature, 2014, 516(7530), 250-3). This invention The embodiment facilitates purification by metal chelate affinity chromatography. Includes other bioaffinity tags such as histidine tags. Alternative embodiments of the present invention In this case, the tag is a FLAG tag or an epitope tag, for example, a Myc tag or a HA tag. It may contain biotin or the like. In further embodiments of the present invention, nanopores are made of biotin or the like. By modifying with a similar compound (e.g., desthiobiotin), strep It may also enhance purification through interaction with toavidin.
[0148] In embodiments of the present invention, the net charge of the polypeptide is increased, and the polyacrylamide gel To enhance protein migration during electrophoresis, a negative charge is added to the protein ends. This can happen. This is because these chemical species of CsgG heterooligomers are the target chemical species. If this is the case, it can lead to improved separation (Howorka et al., Proc. Nat. Acad. Sci See ., 2001, 98(23), 12996-13001). Heterooligomers are present in each pore. This could be useful for introducing a single cysteine residue, and this introduction could be used in the aforementioned DNA polymerase. This may be advantageous in promoting the binding of suitable nucleic acid processing enzymes such as -ase.
[0149] This invention, in part, is based on molecular dynamics derived from variations in the electrical measurements of current flowing through nanopores. Regarding the use of wild-type or modified CsgG nanopores in testing applications.
[0150] Molecules within the channel of a CsgG pore or near one of the channel's openings The coupling will affect the ion flow through the open channel via the pore. Similar to nucleic acid sequencing applications, the current can be changed using a suitable measurement technique. It is possible to measure fluctuations in channel ion flow (e.g., International Publication No. 2000 / 28) Pamphlet No. 312 and D. Stoddart et al., Proc. Natl. Acad. Sci., 2010, 106. (7702-7, or International Publication No. 2009 / 077734 pamphlet). By reducing the current... The degree of ion flow reduction measured in this way depends on the magnitude of interference within or near the pore. It is related to the target molecule (also called the analyte) inside or near the pore. The binding of (bu) results in a detectable and measurable event, thus providing a basis for biosensors. It forms the foundation. Suitable molecules for nanopore sensing include nucleic acids, proteins, peptides, and Examples include small molecules such as pharmaceuticals, toxins, or cytokines.
[0151] The detection of biomolecules is used in personalized drug discovery, medicine, diagnostics, life science research, and environmental monitoring. It is applied to security and / or defense industries.
[0152] In embodiments of the present invention, wild-type or modified Escherichia coli (E. coli) is disclosed herein. i) CsgG nanopores or their homologs can serve as molecular sensors. The procedure for detecting precipitates is described in Howorka et al., Nature Biotechnology (2012) Jun 7; 30(6):506-7 It is described as follows: The analyte molecule to be detected is bound to either side of the channel. It may also bind to the lumen of the channel itself. The binding site is determined by sensing It can be determined by the size of the molecule being targeted. Wild-type CsgG pores act as sensors. In some cases, or in embodiments of the present invention, the CsgG pore may be recombinant or chemical Depending on the method, the strength, location, or specificity of the bond of the sensed molecule can be increased. It is modified to add to the structure of the molecule being sensed. Typical modifications include complementary modifications to the structure of the molecule being sensed. One example is the addition of specific binding sites. If the analyte molecule contains nucleic acids, this binding site is , may contain cyclodextrin or oligonucleotides; for small molecules, This is a known complementary binding region, for example, the antigen-binding portion of an antibody or non-antibody molecule (one). Chain variable fragment (scFv) region, or antigen recognition domain from T cell receptor (TCR) It may include; or, in the case of a protein, it may be the public of the target protein It can be a known ligand. Thus, wild-type or modified Escherichia coli (E. coli) Csg G nanopores or their homologs are used to detect the presence of suitable antigens (including epitopes) in a sample. It can be made to act as a molecular sensor, and suitable antigens for this purpose include: Cell surface antigens (receptors, solid tumors or blood cancer cells (e.g., lymphoma or leukemia)) (Including markers), viral antigens, bacterial antigens, protozoan antigens, allergens, allergy-related Linked molecules, albumin (e.g., human, rodent, or bovine), fluorescent molecules (fluorescein) (including), blood group antigens, small molecules, drugs, enzymes, catalytic sites of enzymes or enzyme substrates, and yeast We can list analogues of the transition states of elementary substrates.
[0153] Modifications can be carried out using known genetic engineering and recombinant DNA techniques. The positioning of the adaptive form of the displacement also depends on the properties of the molecules being sensed, such as size and three-dimensional structure. It will depend on the structure and its biochemical properties. The selection of the structure to be adapted will be computer-based. A structural design can be used. Each is particularly suited to its intended sensing application. A series of custom-made CsgG nanopores are being envisioned. Protein-protein interactions or Determination and optimization of protein-small molecule interactions is performed using surface plasmon resonance to analyze molecular phase BIAcore® (registered trademark) for detecting interactions (BIAcore, Inc., Piscataway, NJ;ww) You can investigate using technologies such as (see also w.biacore.com). ru.
[0154] In this embodiment, the CsgG pore is used to prevent lipidization of the protein N-terminus by the N-terminal C It is a water-soluble octamer form in which the ys residue is substituted with an alternative amino acid. This is possible. In an alternative embodiment, the N-terminus is used to avoid processing by the bacterial lipidization pathway. By removing the end leader sequence, the protein can be expressed in the cytoplasm. Cut.
[0155] CsgG monomer-soluble protein, octamer-soluble protein, and oligomer lipid. A method for producing crystalline CsgG pores is incorporated herein by reference as Goyal ef a l. (Nature, 2014; 516(7530): 250-3), as well as such a manufacturing method. Examples 1 and 2 describe this.
[0156] Mutant CsgG monomer This invention provides mutant CsgG monomers. The pore of the present invention can be formed using this method. The mutant CsgG monomer has a wild-type sequence. It is a monomer different from that of type CsgG monomer, and retains the ability to form pores. It is a monomer. Methods for confirming the ability of mutant monomers to form pores are available in this technology. This is a well-known fact, and such methods will be discussed in more detail below.
[0157] The mutant monomer has improved polynucleotide readout characteristics, that is, polynucleotide It shows improved creotide capture and nucleotide recognition. In detail, it is composed from mutant monomers. The pores that are formed capture nucleotides and polynucleotides more easily than those in the wild type. Furthermore, pores constructed from mutant monomers make it easier to distinguish between different nucleotides. This demonstrates an increase in the current range and a reduction in state fluctuations that increases the signal-to-noise ratio. Furthermore, this contributes to the electric current when polynucleotides move through pores constructed from mutants. The number of nucleotides decreases. This is because polynucleotides move through the pore. This makes it easier to identify the direct relationship between the current observed and the polynucleotide sequence. In addition, pores constructed from mutant monomers may also show increased throughput. In other words, it is more likely to interact with analytes such as polynucleotides. To facilitate the characterization of analytes using pores. Pores constructed from mutant monomers. It can enter the membrane more easily.
[0158] The mutant monomer of the present invention comprises a variant of the sequence shown in Sequence ID No. 390. Number 390 is Escherichia coli Str. K-12 substr. MC4100. It is a wild-type CsgG monomer from MC4100). The variant of sequence number 390 is sequence A polypeptide having a different amino acid sequence from that of number 390, which forms a pore. It is a polypeptide that retains that ability. The ability of the variant to form a pore is the technology The assay can be performed using any method known in the field. For example, Varian The t may be inserted into the amphiphilic layer along with other suitable subunits, and oligomerized to form a port. The ability to form A may be determined. Subunits are inserted into membranes such as amphiphilic layers. The method is known in the art. For example, the subunits are purified into triblocks. The subunits may be suspended in a solution containing a goji copolymer membrane, thereby dispersing the subunits within the membrane. Furthermore, by binding to the membrane, it is inserted into the membrane, aggregates, and becomes functional.
[0159] In all discussions in this specification, the standard single-letter notation for amino acids is used. The following are the components: alanine (A), arginine (R), asparagine (N), asparagus Ginate (D), cysteine (C), glutamate (E), glutamine (Q), glycine ( G), histidine (H), isoleucine (I), leucine (L), lysine (K), methionine Nin (M), phenylalanine (F), proline (P), serine (S), threonine (T) ), tryptophan (W), tyrosine (Y), and valine (V). Standard substitution notation is also used. In other words, Q42R means that the Q at position 42 is substituted with R.
[0160] In one embodiment, the mutant monomer of the present invention comprises sequence 3, which includes one or more of the following: Includes 90 variants: (i) One or more mutations at the following positions (i.e., next (Variations at one or more locations): N40, D43, E44, S54, S57, Q62, R97, E101, E124, E131, R142, T150 and R192, For example, one or more mutations at the following positions (i.e., one or multiple mutations at the following positions) Numerical variations): N40, D43, E44, S54, S57, Q62, E101, E1 31 and T150, or N40, D43, E44, E101 and E131, (ii ) Y51 / N55, Y51 / F56, N55 / F56 or Y51 / N55 / F56 (iii) Q42R or Q42K, (iv) K49R, (v) N102R, N102F, N102Y or N102W, (vi)D149N, D149Q or D1 49R, (vii)E185N, E185Q or E185R, (viii)D195N , D195Q or D195R, (ix) E201N, E201Q or E201R, ( x) E203N, E203Q or E203R, and (xi) one of the following positions Multiple deletions: F48, K49, P50, Y51, P52, A53, S54, N55, F 56 and S57. The variant may include any combination of (i) to (xi). In particular, the variant is
[0161] [ka]
[0162] [ka]
[0163] [ka]
[0164] [ka]
[0165] [ka]
[0166] [ka]
[0167] [ka]
[0168] [ka] It may contain.
[0169] If the variant contains any one of (i) and (iii)~(xi), then Y51 , one or more of N55 and F56, for example, Y51, N55, F56, Y51 / Mutations in N55, Y51 / F56, N55 / F56, or Y51 / N55 / F56 It may also include other things.
[0170] (i) The variants are N40, D43, E44, S54, S57, Q62, All R97, E101, E124, E131, R142, T150 and R192 This may include variations in number and combinations. In (i), the variant is preferably , one or more mutations at the following positions (i.e., one or more mutations at the following positions) (Variations include): N40, D43, E44, S54, S57, Q62, E101, E1 31 and T150. In (i), the variant is preferably 1 at the following position Includes one or more mutations (i.e., mutations at one or more of the following locations):N 40, D43, E44, E101 and E131. In (i), the variant is preferred or including mutations in S54 and / or S57. (i) In variant (a) S54 and / or S57 and (b) Y51, N55 and one or more of F56, for example, Y51, N55, F56, Y51 / N55, This includes mutations in Y51 / F56, N55 / F56, or Y51 / N55 / F56. If S54 and / or S57 are missing in xi), then it / they are in (i). It cannot undergo mutation, and vice versa. (i) In the variant, Preferably, the variant includes a mutation at T150, for example, T150I. Alternatively, the variant is Preferably, (a) T150 and (b) one of Y51, N55 and F56 There are multiple, for example, Y51, N55, F56, Y51 / N55, Y51 / F56, N55 / (i) Variants Preferably, this includes a mutation in Q62, such as Q62R or Q62K. The variants are preferably (a) Q62 and (b) Y51, N55 and F56 One or more of these, for example, Y51, N55, F56, Y51 / N55, Y51 / F5 6. Includes mutations in N55 / F56 or Y51 / N55 / F56. The variant is, D43, E44, Q62 or any combination thereof, for example, D43, E44, Q62 , D43 / E44, D43 / Q62, E44 / Q62 or D43 / E44 / Q62 It may contain mutations. Alternatively, the variant is preferably (a)D43,E4 4, Q62, D43 / E44, D43 / Q62, E44 / Q62 or D43 / E44 / Q62 and (b) one or more of Y51, N55 and F56, for example, Y51 , N55, F56, Y51 / N55, Y51 / F56, N55 / F56 or Y51 / N Includes the mutation in 55 / F56.
[0171] In this application, the symbol (ii) and elsewhere / means "and" Therefore, Y51 / N55 are Y51 and N55. In (ii), the variant is Preferably, the mutation includes a mutation at Y51 / N55. The constriction within CsgG is at residue Y51. It is proposed that it is made of three stacked concentric rings formed by the side chains of N55 and F56. It has been proposed (Goyal et al., 2014, Nature, 516, 250-253). Therefore, (ii) Mutations in these residues occur when polynucleotides are moving through the pore. By reducing the number of nucleotides contributing to the flow, (polynucleotides can make pores Identifying the direct relationship between observed currents (when moving through the material) and polynucleotides is easier. This can be easily done. Variants and pores useful for the method of the present invention are discussed below. Y56 may be mutated using one of the methods.
[0172] (v) The variant is N102R, N102F, N102Y or N102 It may contain W. The variants are preferably (a) N102R, N102F, N1 (b) 02Y or N102W, and one or more of Y51, N55 and F56 Numbers, for example, Y51, N55, F56, Y51 / N55, Y51 / F56, N55 / F5 Includes mutations in 6 or Y51 / N55 / F56.
[0173] In (xi), K49, P50, Y51, P52, A53, S54, N55, F5 Any number and combination of 6 and S57 may be missing. Preferably, One of K49, P50, Y51, P52, A53, S54, N55 and S57 or Multiple entries may be missing. (xi) is one of Y51, N55, and F56. If it is deleted, it / they cannot undergo the mutation in (ii) and the reverse The same applies to this as well.
[0174] (i) The variant is preferably the following substitutions: N40R, N40K, D43 N, D43Q, D43R, D43K, E44N, E44Q, E44R, E44K, S54 P, S57P, Q62R, Q62K, R97N, R97G, R97L, E101N, E1 01Q, E101R, E101K, E101F, E101Y, E101W, E124N, E124Q, E124R, E124K, E124F, E124Y, E124W, E131 One or more of D, R142E, R142N, T150I, R192E, and R192N Numbers, for example, N40R, N40K, D43N, D43Q, D43R, D43K, E44N , E44Q, E44R, E44K, S54P, S57P, Q62R, Q62K, E101 N, E101Q, E101R, E101K, E101F, E101Y, E101W, E1 One or more of 31D and T150I, or N40R, N40K, D43N, D 43Q, D43R, D43K, E44N, E44Q, E44R, E44K, E101N, E101Q, E101R, E101K, E101F, E101Y, E101W and E1 Includes one or more of 31D. Variants include any number and combination of these substitutions. (i) may include (i) the variant is preferably S54P and / or or includes S57P. In (i), the variant is preferably (a)S54P and (b) one or more of Y51, N55 and F56, For example, Y51, N55, F56, Y51 / N55, Y51 / F56, N55 / F56 or include mutations in Y51 / N55 / F56. One of Y51, N55 and F56 is also The variation in one or more of the following may be any of those discussed below. (i) The variants preferably include F56A / S57P or S54P / F56A. Ant preferably contains T150I. Alternatively, the variant preferably contains (a )T150I and (b)Y51, N55 and F56, one or more of these, for example, Y51, N55, F56, Y51 / N55, Y51 / F56, N55 / F56 or Y5 Includes mutations in 1 / N55 / F56.
[0175] (i) The variant preferably includes Q62R or Q62K. The variants are preferably (a) Q62R or Q62K, and (b) Y51 , one or more of N55 and F56, for example, Y51, N55, F56, Y51 / Mutations in N55, Y51 / F56, N55 / F56, or Y51 / N55 / F56 Includes. Variants include D43N, E44N, Q62R or Q62K or these. Any combination, for example, D43N, E44N, Q62R, Q62K, D43N / E44N , D43N / Q62R, D43N / Q62K, E44N / Q62R, E44N / Q62K This may include D43N / E44N / Q62R or D43N / E44N / Q62K. Alternatively, the variant is preferably (a) D43N, E44N, Q62R, Q62 K, D43N / E44N, D43N / Q62R, D43N / Q62K, E44N / Q62 R, E44N / Q62K, D43N / E44N / Q62R or D43N / E44N / Q 62K, and (b) one or more of Y51, N55 and F56, for example, Y5 1, N55, F56, Y51 / N55, Y51 / F56, N55 / F56 or Y51 / This includes mutations in N55 / F56.
[0176] In (i), the variant preferably includes D43N.
[0177] In (i), the variant is preferably E101R, E101S, E101F This includes E101N.
[0178] (i) The variant is preferably E124N, E124Q, E124R, E124K, E124F, E124Y, E124W, or E124D, for example, E124N Includes.
[0179] In (i), the variant preferably includes R142E and R142N.
[0180] (i) The variant preferably includes R97N, R97G, or R97L. nothing.
[0181] (i) The variant preferably includes R192E and R192N.
[0182] (ii) In the variant, preferably,
[0183] [ka]
[0184] [ka]
[0185] [ka] Includes.
[0186] (ii) The variants are preferably Y51R / F56Q and Y51N / F5 6N, Y51M / F56Q, Y51L / F56Q, Y51I / F56Q, Y51V / F5 6Q, Y51A / F56Q, Y51P / F56Q, Y51G / F56Q, Y51C / F5 6Q, Y51Q / F56Q, Y51N / F56Q, Y51S / F56Q, Y51E / F5 Includes 6Q, Y51D / F56Q, Y51K / F56Q, or Y51H / F56Q.
[0187] (ii) The variants are preferably Y51T / F56Q, Y51Q / F5 Includes 6Q or Y51A / F56Q.
[0188] (ii) In the variant, the variant is preferably Y51T / F56F, Y51T / F5 6M, Y51T / F56L, Y51T / F56I, Y51T / F56V, Y51T / F5 6A, Y51T / F56P, Y51T / F56G, Y51T / F56C, Y51T / F5 6Q, Y51T / F56N, Y51T / F56T, Y51T / F56S, Y51T / F5 6E, Y51T / F56D, Y51T / F56K, Y51T / F56H or Y51T / Includes F56R.
[0189] (ii) The variant is preferably Y51T / N55Q, Y51T / N5 Includes 5S or Y51T / N55A.
[0190] (ii) In the above, the variants are preferably Y51A / F56F and Y51A / F5 6L, Y51A / F56I, Y51A / F56V, Y51A / F56A, Y51A / F5 6P, Y51A / F56G, Y51A / F56C, Y51A / F56Q, Y51A / F5 6N, Y51A / F56T, Y51A / F56S, Y51A / F56E, Y51A / F5 Includes 6D, Y51A / F56K, Y51A / F56H, or Y51A / F56R.
[0191] (ii) The variants are preferably Y51C / F56A and Y51E / F5 6A, Y51D / F56A, Y51K / F56A, Y51H / F56A, Y51Q / F5 6A, Y51N / F56A, Y51S / F56A, Y51P / F56A or Y51V / Includes F56A.
[0192] In (xi), the variant is preferably a deletion of Y51 / P52, Y51 / P5 Missing sections 2 / A53, P50-P52, P50-A53, K49-Y51 Loss, deletions of K49-A53 and substitution with a single proline (P), deletions of K49-S54 and substitution in a single P, deletion of Y51~A53, deletion of Y51~S54, N55 / F5 Deletion of 6, deletion of N55~S57, deletion of N55 / F56 and substitution with a single P, N5 Deletion of 5 / F56 and substitution with a single glycine (G), deletion of N55 / F56 and single Substitution with one alanine (A), deletion of N55 / F56 and substitution with a single P and Y5 1N, N55 / F56 deletion and substitution with a single P and Y51Q, N55 / F56 Deletion and substitution in a single P and deletion of Y51S, N55 / F56 and in a single G Substitution and deletion of Y51N, N55 / F56 and substitution and Y51Q, N Deletion of 55 / F56 and replacement with a single G and deletion of Y51S, N55 / F56 and Substitution of a single A and deletion of Y51N, N55 / F56 and substitution of a single A / Y This includes deletions of 51Q, or N55 / F56, and substitutions with a single A, as well as Y51S.
[0193] The variant is comfortable,
[0194] [ka]
[0195] [ka] Includes.
[0196] The current as polynucleotides move through the pore causes fewer nucleotides to Preferred variants of the present invention that form contributing pores are Y51A / F56A, Y51A / F56N, Y51I / F56A, Y51L / F56A, Y51T / F56A, Y51I / F56N, Y51L / F56N or Y51T / F56N, or more preferably Y Includes 51I / F56A, Y51L / F56A, or Y51T / F56A. As discussed above. Thus, this is the electricity observed (when polynucleotides are moving through the pore). This makes it easier to identify direct relationships between fluids and polynucleotides.
[0197] Preferred variants that form pores exhibiting range increase include mutations at the following locations: Y51, F56, D149, E185, E201 and E203, N55 and F56, Y51 and F56, Y51, N55 and F56, or F56 and N102.
[0198] Preferred variants that form pores exhibiting range increase include: Y51N, F56A, D149N, E185R, E201N and E203N, N55S and F56Q, Y51A and F56A, Y51A and F56N, Y51I and F56A, Y51L and F56A, Y51T and F56A, Y51I and F56N, Y51L and F56N, Y51T and F56N, Y51T and F56Q, Y51A, N55S and F56A, Y51A, N55S and F56N, Y51T, N55S and F56Q, or F56Q and N102R.
[0199] The current as polynucleotides move through the pore causes fewer nucleotides to A preferred variant that forms a contributing pore is, Variations at positions N55 and F56, for example, N55X and F56Q (where X is Any amino acid), or Mutations at positions Y51 and F56, for example, Y51X and F56Q (where X is (Any amino acid) Includes.
[0200] The preferred variant that forms a pore showing increased throughput is the mutation at the following location. Includes: D149, E185 and E203, D149, E185, E201 and E203, or D149, E185, D195, E201, and E203.
[0201] Preferred variants that form pores exhibiting increased throughput include: D149N, E185N and E203N, D149N, E185N, E201N and E203N, D149N, E185R, D195N, E201N and E203N, or D149N, E185R, D195N, E201R, and E203N.
[0202] The preferred variants that form pores with increased polynucleotide capture are the following mutations. Includes: D43N / Y51T / F56Q, E44N / Y51T / F56Q, D43N / E44N / Y51T / F56Q, Y51T / F56Q / Q62R, D43N / Y51T / F56Q / Q62R, E44N / Y51T / F56Q / Q62R, or D43N / E44N / Y51T / F56Q / Q62R.
[0203] The preferred variant includes the following mutations: D149R / E185R / E201R / E203R or Y51T / F56Q / D1 49R / E185R / E201R / E203R, D149N / E185N / E201N / E203N or Y51T / F56Q / D1 49N / E185N / E201N / E203N, E201R / E203R or Y51T / F56Q / E201R / E203R, E201N / E203R or Y51T / F56Q / E201N / E203R, E203R or Y51T / F56Q / E203R, E203N or Y51T / F56Q / E203N, E201R or Y51T / F56Q / E201R, E201N or Y51T / F56Q / E201N, E185R or Y51T / F56Q / E185R, E185N or Y51T / F56Q / E185N, D149R or Y51T / F56Q / D149R, D149N or Y51T / F56Q / D149N, R142E or Y51T / F56Q / R142E, R142N or Y51T / F56Q / R142N, R192E or Y51T / F56Q / R192E, R192N or Y51T / F56Q / R192N
[0204] The preferred variant includes the following mutations: Y51A / F56Q / E101N / N102R, Y51A / F56Q / R97N / N102G, Y51A / F56Q / R97N / N102R, Y51A / F56Q / R97N, Y51A / F56Q / R97G, Y51A / F56Q / R97L, Y51A / F56Q / N102R, Y51A / F56Q / N102F, Y51A / F56Q / N102G, Y51A / F56Q / E101R, Y51A / F56Q / E101F, Y51A / F56Q / E101N, or Y51A / F56Q / E101G.
[0205] The present invention relates to a variant CsgG monomer containing a variant of the sequence shown in sequence 390. This also provides a variant CsgG monomer, which includes a mutation in T150. A preferred variant that forms a pore showing increased insertion includes T150I. The mutation, for example T150I, is combined with any of the mutations or combinations of mutations discussed above. They can be combined.
[0206] The present invention relates to sequence 3, which includes a combination of mutations present in the variant disclosed in the examples. We also provide variant CsgG monomers containing the sequence variant indicated by 90.
[0207] Methods for introducing or substituting naturally occurring amino acids are well known in the art. For example, methionine (M) is appropriately included in polynucleotides that encode mutant monomers. By replacing the methionine codon (ATG) in a specific position with the arginine codon (CGT) It can then be substituted with arginine (R). After that, the polynucleotide This can be expressed as discussed below.
[0208] Methods for introducing or substituting amino acids that do not exist in nature are also well known in this field. For example, synthetic aminoacyls in the IVTT system used to express mutant monomers. By including -tRNA, it is possible to introduce amino acids that do not exist in nature. Alternatively, in E. coli, which has nutritional requirements for specific amino acids, Synthesizing specific amino acids (i.e., mutant monomers in the presence of analogues that do not exist in nature) By expressing this gene, amino acids that do not exist in nature can be introduced. When introducing mutant monomers using plutido synthesis, naked ligation (naked It is also possible to introduce amino acids that do not exist in nature through ligation.
[0209] Other monomers of the present invention In another embodiment, the present invention relates to a mutation comprising a variant of the sequence shown in SEQ ID NO: 390. A CsgG monomer in which the variant is at position Y51, N55, and F56 The variant provides a CsgG monomer containing mutations in multiple locations. The variant is Y5 1, N55, F56, Y51 / N55, Y51 / F56, N55 / F56, or Y51 It may contain mutations in / N55 / F56. The variant is preferably Y51, This includes mutations at N55 or F56. The variant is at the positions Y51, N55 discussed above. Specific mutations in one or more of F56, and in any combination thereof It may include a difference. One or more of Y51, N55, and F56 are any of the A It may also be substituted with a amino acid. Y51 is F, M, L, I, V, A, P, G, C, Q, N, T, S, E, D, K, H, or R, for example, A, S, T, N, or Q are substituted. It may be. N55 is F, M, L, I, V, A, P, G, C, Q, T, S, E, D, K, H, or R may be substituted, for example, A, S, T, or Q. F56 is M, L, I, V, A, P, G, C, Q, N, T, S, E, D, K, H or R, for example However, it may be replaced with A, S, T, N, or Q.
[0210] The variant may further include one or more of the following: (i) in the following positions one or more mutations (i.e., mutations at one or more of the following locations) (i) N40, D43, E44, S54, S57, Q62, R97, E101, E124, E1 31, R142, T150 and R192; (iii) Q42R or Q42K; (iv )K49R;(v)N102R, N102F, N102Y or N102W;(vi)D 149N, D149Q or D149R; (vii) E185N, E185Q or E1 85R;(viii)D195N, D195Q or D195R;(ix)E201N, E201Q or E201R; (x) E203N, E203Q or E203R; and (xi) Deletion of one or more of the following positions: F48, K49, P50, Y51, P5 2, A53, S54, N55, F56 and S57. The variants are (i) discussed above. It may include any combination of (iii) to (xi). The variant is (i) And (iii) to (xi) may include any of the embodiments discussed above.
[0211] variant In addition to the specific mutations discussed above, variants may contain other mutations. Over the entire length of the 390-amino acid sequence, the variant preferably maintains amino acid identity. Based on this, the sequence will be at least 50% homologous. More preferably, Varian Based on amino acid identity, the amino acid sequence of sequence number 390 is used throughout the entire sequence. At least 55%, at least 60%, at least 65%, at least 70%, and 75%, at least 80%, at least 85%, at least 90%, and more preferably 75%, 80%, 85%, 90%, and 90%. They may be at least 95%, 97%, or 99% homologous. 100 or more consecutive, For example, across stretches of amino acids of 125, 150, 175, or 200 or more, At least 80%, for example, at least 85%, 90%, or 95% amino acid identity ("solid" There may also be "hard homology."
[0212] Homology can be determined using standard methods in the art. For example, U The WGCG Package can be used to calculate homology, for example, It provides the BESTFIT program used by default (Devereux et al., (1984) Nucleic Acids Research 12, pp. 387-395). For example, Altschul SF (1993) J Mol Evol 36:290-300; Altschul, S. F et al., (1990) J Mol Biol 215:403-10. As described, the PILEUP and BLAST algorithms are used to determine homology. Calculate and arrange the sequences (for example, equivalent residues or corresponding sequences (usually their differences) It can be identified (by default settings). The software for performing BLAST analysis is: National Center for Biological Information (http: / / www.ncbi.nlm.nih.go) It is publicly available in v / ).
[0213] Sequence ID 390 is a substrain of Escherichia coli K12 strain MC4100. It is the wild-type CsgG monomer from ubstr. MC4100). The variant of SEQ ID NO: 390 is , may include one of the substitutions present in other CsgG homologs. Preferred CsgG Homologous compounds are shown in sequence numbers 391-395 and 414-429. Variants are distributed The substitutions present in sequence numbers 391-395 and 414-429 compared to column number 390 This may include combinations of one or more items.
[0214] In addition to those discussed above, amino acid substitutions, for example, 1, 2, 3, 4, 5, 10, 20, The amino acid sequence of SEQ ID NO. 390 may also be subjected to substitutions of 30 or fewer units. Conservative substitutions. This refers to other amino acids with similar chemical structures, similar chemical properties, or similar side chain volumes. The amino acids to be introduced have similar polarity and hydrophilicity to the amino acids they substitute for. It can be hydrophobic, basic, acidic, neutral, or charged. Alternatively, conservative substitutions are Instead of existing aromatic or aliphatic amino acids, another amino acid that is aromatic or aliphatic. Acids may also be introduced. Conservative amino acid changes are well known in this field, as shown in the table below. Select such changes according to the characteristics of the 20 major amino acids as defined in 1. This is possible. If the amino acids have similar polarity, this also applies to the amino acid side chains in Table 2. This can be determined by referring to the Hydropathy Scale.
[0215] [Table 1]
[0216] [Table 2]
[0217] One or more amino acid residues of the amino acid sequence of Sequence ID No. 390 are the above-mentioned polypeptide Further deletions may occur from the initial number. 1, 2, 3, 4, 5, 10, 20, or 30 or more. Residues may be deleted in this case.
[0218] The variant may also contain the fragment of sequence number 390. Such fragments may form pores. It retains activity. The fragments are at least 50, at least 100, at least 150 in length. It may be at least 200 or at least 250 amino acids. Such fragments It may be used to generate pores. The fragment is preferably a transmembrane of SEQ ID NO: 390. This includes domains K135-Q153 and S183-S208.
[0219] One or more amino acids may be added to the polypeptide as an alternative or further addition. It also states: The amino acid sequence of SEQ ID NO: 390 or its polypeptide variant or fragment An elongation portion may be provided at the amino or carboxyl terminus of one of the molecules. They can be quite short, for example, they can be 1 to 10 amino acids long. Or, the elongated part These can be longer, for example, fewer than 50 or 100 amino acids. Body proteins may also be fused to the amino acid sequence according to the present invention. Other fusion proteins This will be discussed in more detail below.
[0220] As discussed above, the variant has a different amino acid sequence than that of sequence number 390. It is a polypeptide that forms pores and retains that ability. Liant typically contains the region of SEQ ID NO: 390, which is involved in pore formation. β-Vale The pore-forming ability of CsgG containing β-sheets is brought about by the β-sheets within each subunit. The variant of sequence number 390 is the region within sequence number 390 that forms a β-sheet. , that is, it usually includes K135~Q153 and S183~S208. If the variant retains its ability to form pores, then it will form β-sheets. One or more modifications can be applied to the region of sequence number 390. The variant preferably consists of one or more modifications, such as substitution, addition, or deletion. It is contained within the α-helix and / or loop region of the α-helix.
[0221] Monomers derived from CsgG are used to assist in their identification or purification, for example. By adding streptavidin tags, or to promote their secretion from cells. By adding a signal sequence (if the monomer does not naturally contain such a sequence), Modifications may be made. Other suitable tags are discussed in more detail below. Monomers are manifestation marks. It may be labeled with an identifier. The manifest label allows the monomer to be detected. Any sign is acceptable. Suitable signs are described below.
[0222] Monomers derived from CsgG can also be produced using D-amino acids. For example Furthermore, monomers derived from CsgG may contain a mixture of L-amino acids and D-amino acids. Yes, this is common in the technical field concerning the production of such proteins or peptides. It happened a while ago.
[0223] Monomers derived from CsgG are used to facilitate nucleotide identification, one or more It has specific modifications. Monomers derived from CsgG also have other nonspecific modifications, such as Modifications can be present if they do not interfere with pore formation. Numerous nonspecific side chain modifications. Such modifications are known in the art, and such modifications are used in the side chains of monomers derived from CsgG. Such modifications can be applied to, for example, a reaction with an aldehyde, and the subsequent Reduction with NaBH4 leads to reduced alkylation, and amidation with methyl acetimide, One example is acylation with acetic anhydride.
[0224] Monomers derived from CsgG are produced using standard methods known in the art. It is possible. Monomers derived from CsgG can also be synthesized, Alternatively, they may be produced by recombinant methods. For example, monomers can be converted in vitro. They may also be synthesized by translation and transcription (IVTT). Preferred methods for producing pores and monomers. The law is based on the specification of international application PCT / GB09 / 001690 (International Publication 2010 / 00 (Published as pamphlet No. 4273), and PCT specification No. 09 / 001679 ( (Published as International Publication No. 2010 / 004265 or PCT / GB) Specification No. 10 / 000133 (as International Publication No. 2010 / 086603 pamphlet) This is discussed in the publicly available publication. The method of inserting pores into the membrane is discussed.
[0225] In some embodiments, the mutant monomers are chemically modified. By this method, any site can be chemically modified. Mutant monomers are preferably This refers to the binding of a molecule to one or more cysteine molecules (cysteine bond), or one or multiple The binding of molecules to several lysine molecules, the binding of molecules to one or more non-natural amino acids, epithet The molecule is modified by enzymatic modification of the segment or by chemical modification of the terminal segment. Such modifications are preferred. Such implementation methods are well known in the art. Mutant mono is formed by the binding of any of the molecules. Mer may be chemically modified. For example, mutant monomers may be used as dyes or fluorophores. It may also be chemically modified by bonding.
[0226] In some embodiments, the mutant monomer is used to form a pore containing the monomer and a target nucleotide. Alternatively, they may be chemically modified with molecular adapters that facilitate interaction with the target polynucleotide sequence. The presence of an adapter allows the host to connect the pore with a nucleotide or polynucleotide sequence. • By improving guest chemistry, the pores formed from mutant monomers To improve sequencing capabilities. The principles of host-guest chemistry are relevant to this technology. It is well known that the adapter is compatible with nucleotide or polynucleotide sequences. The adapter affects the physical or chemical properties of the pore, which improves the interaction. , by changing the charge of the pore barrel or channel, or nucleotides as well Alternatively, by specifically interacting with or binding to polynucleotide sequences, the pore and This can promote that interaction.
[0227] The molecular adapter is preferably a cyclic molecule, cyclodextrin, or hybridize. Chemical species that can be combined, DNA binders or insertants, peptides or peptide analogs, compounds A polymer, an aromatic planar molecule, a small molecule with a positive charge, or a small molecule capable of hydrogen bonding. ru.
[0228] The adapter may be annular. The annular adapter is preferably the same pair as the pore. It has symmetry. CsgG typically has 8 or 9 subunits around its central axis. The adapter is preferably 8- or 9-fold symmetry. This will be explained in more detail below. To discuss.
[0229] Adapters are typically nucleotides or polynucleotides through host-guest chemistry. It interacts with nucleotide sequences. Adapters are typically nucleotides or polynucleotides. It can interact with rheotide sequences. The adapter can interact with nucleotides or polynucleotides. It contains one or more chemical groups that can interact with the rheotide sequence. The chemical groups are preferably non-covalent interactions, such as hydrophobic interactions and hydrogen bonds. , by van der Waals forces, π-cation interactions and / or electrostatic forces, nucle Interacts with nucleotide or polynucleotide sequences. One or more chemical groups that can interact with the sequence preferably have a positive charge. It can interact with one or a nucleotide or polynucleotide sequence. The multiple chemical groups more preferably include amino groups. The amino groups may be primary, secondary or primary It can be bonded to a tertiary carbon atom. The adapter is more preferably made of mesh. The adapter contains a ring of amino groups, for example, a ring with 6, 7, or 8 amino groups. It contains a ring of eight amino groups. The ring of protonated amino groups is a nucleotide or polynucleotide. It can interact with negatively charged phosphate groups in the rheotide sequence.
[0230] The correct positioning of the adapter within the pore is determined by the relationship between the adapter and the pore containing the mutant monomer. Host-guest chemistry can be enhanced. The adapter is preferably , one or more chemicals that can interact with one or more amino acids in the pore The group contains a non-covalent interaction, for example, a hydrophobic interaction. , by hydrogen bonding, van der Waals forces, π-cation interactions and / or electrostatic forces One or more chemicals that can interact with one or more amino acids in the pore Contains a group. Chemical groups that can interact with one or more amino acids in the pore are generally Usually, it is hydroxyl or amine. The hydroxyl group is primary, secondary or tertiary carbon It can bond to elementary atoms. The hydroxyl group can form hydrogen bonds with uncharged amino acids within the pore. A bond can be formed. Interaction between pores and nucleotide or polynucleotide sequences. Any adapter that facilitates use can be used.
[0231] Suitable adapters include cyclodextrins, cyclic peptides, and cucurbituryl. Examples include, but are not limited to, the adapter, which is preferably cyclodextrin. It is a cyclodextrin or its derivatives. Cyclodextrin or its derivatives are used by Eliseev, AV, and disclosed in Schneider, HJ. (1994) J. Am. Chem. Soc. 116, 6081-6088. It may be any of the following. The adapter is more preferably heptakis-6-amino- β-cyclodextrin (am7-βCD), 6-monodeoxy-6-monoamino-β- Cyclodextrin (am1-βCD) or heptakis-(6-deoxy-6-guani) It is guanidino-cyclodextrin (gu7-βCD). Guanidino in gu7-βCD The group has a much higher pKa than the primary amine in am7-βCD, and therefore is larger It has a strong positive charge. This gu7-βCD adapter is used when nucleotides remain in the pore. To increase the interval, to increase the accuracy of the measured residual current, as well as high temperature and low It can be used to increase the base detection rate with a high data collection rate.
[0232] As will be discussed in more detail below, succinimidi 3-(2-pyridyldithio)propionate When using the (SPDP) crosslinking agent, the adapter is preferably heptakis(6-de Oxy-6-amino)-6-N-mono(2-pyridyl)dithiopropanoyl-β-cyclo It is dextrin (am6amPDP1-βCD).
[0233] A more suitable adapter is γ containing 9 sugar units (and therefore having 9-fold symmetry). -Cyclodextrins are an example. γ-Cyclodextrins contain a linker molecule. This may occur, or the modified sugar units used in the example of β-cyclodextrin discussed above may be It may also be modified to include all or more.
[0234] The molecular adapter is preferably covalently bonded to the mutant monomer. The adapter can be covalently bonded to the pore using any known method. Molecular adapters are usually bonded by chemical bonds. When bound, one or more cysteines are substituted into the mutant, for example, within the barrel. Therefore, it is preferable that it be introduced. The mutant monomer is one of the mutant monomers Alternatively, it can be chemically modified by the binding of molecular adapters to multiple cysteine molecules. One or more cysteine compounds may be naturally occurring, namely, SEQ ID NO: 3 It may be at position 1 and / or 215 of 90. Alternatively, the mutant monomer may be at other positions. Chemical modification by the attachment of one or more molecules to cysteine introduced at the position. It can also be done. The cysteine at position 215 has a molecular adapter that is the cysteine at position 1 or To ensure that cysteine is not introduced at a different location and does not bind to that position. For example, it may be removed by substitution.
[0235] The reactivity of a cysteine residue can also be enhanced by modifications of adjacent residues. For example, adjacent The basic group of the arginine, histidine, or lysine residue in contact with it is the thiol group of cysteine. The pKa of the more reactive S - This will change the pKa to the original value. The reactivity may also be protected by thiol protecting groups such as dTNB. A linker is formed. Before combining, react one or more cysteine residues of the mutant monomers with their protecting groups. This can be done. The molecule may also be directly bound to the mutant monomer. The molecule is preferred Alternatively, linkers such as chemical crosslinking agents or peptide linkers can be used to connect mutant molecules. They are combined.
[0236] Suitable chemical crosslinking agents are well known in the art. A preferred crosslinking agent is 3 -(pyridine-2-yldisulfanyl)propanoate 2,5-dioxopyrrolidine-1- 2,5-Dioxopyrrolidine, 4-(pyridine-2-yldisulfanyl)butanoate -1-yl and 8-(pyridine-2-yldisulfanyl)octanoate 2,5-di Xopyrrolidine-1-yl is one example. The most preferred crosslinking agent is 3-(2-pyridyldi This is succinimidyl thiopropionate (SPDP). Typically, the molecule is bifunctionally crosslinked. After covalently bonding to the agent, the molecule / crosslinking agent complex is covalently bonded to the mutant monomer, After covalently bonding a bifunctional crosslinking agent to a monomer, the bifunctional crosslinking agent / monomer complex is It is also possible to bind it to molecules.
[0237] The linker is preferably resistant to dithiothreitol (DTT). Examples include iodoacetamide and maleimide linkers, but these Not limited to this.
[0238] In other embodiments, the monomer is bound to a polynucleotide-binding protein. Yes. This allows for modular sequencing that can be used in the sequencing method of the present invention. A binding system is formed. Polynucleotide-binding proteins are discussed below.
[0239] The polynucleotide-binding protein is preferably covalently bonded to the mutant monomer. A protein is covalently bonded to a monomer using any method known in the art. Monomers and proteins can be chemically fused, or they can be separated. They can also be fused genetically. The entire construct is expressed from a single polynucleotide sequence. In this case, the monomer and protein are genetically fused. Genetic fusion to a binding protein is described in International Patent Application No. PCT / GB09 / 001679. This is discussed in the detailed paper (published as international publication no. 2010 / 004265). It is.
[0240] When polynucleotide-binding proteins are bound by cysteine bonds, one or It is preferable that multiple cysteines are introduced into the mutant by substitution. One or Multiple cysteines are preferably low in conservation among homologs (this is due to mutation or insertion). It is introduced into the loop region (indicating that input may be permitted). Therefore, those regions are It is suitable for binding polynucleotide-binding proteins. In such embodiments, 251 The naturally occurring cysteine at the top position may be removed. The reactivity of the cysteine residue is as follows: This can be improved by modifying the text.
[0241] Polynucleotide-binding proteins can also be directly bound to mutant monomers, They may also be joined via one or more linkers. International Application No. PCT / G Specification No. B10 / 000132 (as International Publication No. 2010 / 086602 pamphlet) Using the hybridization linker described (published), the molecule is transformed into a mutant mono. It may be bound to a MER. Alternatively, a peptide linker may be used. Peptidine A linker is an amino acid sequence. The length, flexibility, and hydrophilicity of a peptide linker are typically... The peptide linker is designed so as not to interfere with the function of monomers and molecules. A flexible linker consists of 2 to 20 serine and / or glycine amino acids, for example, 4 , 6, 8, 10 or 16 stretches. A more preferred flexible linker is (SG)1, (SG)2, (SG)3, (SG)4, (SG)5 and (SG)8 are listed. In this case, S is serine and G is glycine. A preferred rigid linker is, A stretch of 2-30 proline amino acids, for example, 4, 6, 8, 16, or 24. A more preferred rigid linker is (P) 12 One example is proline, in which case P is proline. That is the case.
[0242] Mutant monomers are chemically modified in molecular adapters and polynucleotide-binding proteins. It can be modified.
[0243] Molecules (molecules that chemically modify monomers) can also be directly bonded to monomers. or International Application No. PCT / GB09 / 001690 Specification (International Publication No. 2010 / 00 (Published as pamphlet No. 4273), and PCT specification No. 09 / 001679 ( (Published as International Publication No. 2010 / 004265 or PCT / G) Specification No. B10 / 000133 (as International Publication No. 2010 / 086603 pamphlet) They may also be linked via a linker, such as one publicly disclosed.
[0244] Any of the proteins described herein, such as mutant monomers and pores of the present invention. This helps in the identification or purification of such residues, for example, histidine residues (his tags). , aspartic acid residue (asp tag), streptavidin tag, flag tag, SUM This can be detected by the addition of O tags, GST tags, or MBP tags, or by their secretion from cells. Addition of a signal sequence to promote (the polypeptide does not naturally contain such a sequence) It may be modified depending on the case. An alternative method for introducing a gene tag is tag By chemically reacting the protein at its native or engineered location. Yes, it exists. An example of this method is gel shift to engineered cysteine outside the protein. This method likely involves reacting reagents. This method is also a method for separating hemolysin heterooligomers. This has been proven (Chem Biol. 1997 Jul;4(7):497-505).
[0245] Any of the proteins described herein, such as mutant monomers and pores of the present invention. These proteins may be labeled with a manifestation label. The manifestation label makes the protein detectable. Any suitable label may be used. Suitable labels include fluorescent molecules, radioisotopes, for example, 125 I, 35 S, enzymes, antibodies, antigens, polynucleotides, and ligands, for example. Biotin is one example, but it is not limited to these.
[0246] Any of the proteins described herein, such as monomers or pores of the present invention, are compound They may be manufactured chemically or by recombinant means. For example Proteins are synthesized by in vitro translation and transcription (IVTT). There is. The amino acid sequence of a protein is modified to include amino acids that do not exist in nature. Proteins may also be modified to increase their stability. If the quality is produced by synthetic means, such amino acids can be introduced during production. Proteins can be modified after either synthetic production or recombinant production. It is also said.
[0247] Proteins can also be produced using D-amino acids. For example, proteins It may contain a mixture of L-amino acids and D-amino acids. This is the nature of such proteins. This is commonplace in the technical field of quality or peptide production.
[0248] If other nonspecific modifications do not interfere with the function of the protein, If so, it can be. Numerous nonspecific side-chain modifications are known in the art, and Such modifications can be applied to the side chains of proteins. For example, Reaction with aldehydes, followed by reduction with NaBH4, resulting in the reduction and alkylation of amino acids. Examples include amidation with methyl acetimide or acylation with acetic anhydride.
[0249] Any of the proteins described herein, comprising the monomers and pores of the present invention, It can be produced using standard methods known in the technical field. The polynucleotide sequences to be derived and duplicated are obtained using standard methods in the art. It can be manufactured. The polynucleotide sequence encoding the protein is in the present art. It can be expressed in bacterial host cells using standard techniques. The protein is Polypeptides produced in cells by in-situ expression of recombinant expression vectors Expression vectors are inducible promo vectors used to control polypeptide expression. These methods may have a ter depending on the case. These methods are described in Sambrook, J. and Russell, D. (2001). Mol ecular Cloning: A Laboratory Manual, 3rd Edition. Cold Spring Harbor Laboratory It is listed in the Press, Cold Spring Harbor, NY.
[0250] Proteins are produced by any protein liquid chromatography system. It can be produced on a large scale after purification from living organisms or after recombinant expression. Typical Examples of protein liquid chromatography systems include FPLC, AKTA system, and Bi o-Cad system, Bio-Rad BioLogic system and Gilson One example is an HPLC system.
[0251] Structures The present invention is a construct comprising two or more covalently bonded CsgG monomers, The present invention also provides constructs in which at least one of the monomers is a mutant monomer of the present invention. The structure retains its ability to form a pore. This can be determined as discussed above. It is possible to use one or more constructs of the present invention to feature polynucleotides. It can create pores for sequencing, for example. The structure is small At least two, at least three, at least four, at least five, at least six, few Contains at least 7, at least 8, at least 9, or at least 10 monomers Sometimes, the construct preferably contains two monomers. Two or more monomers may be the same It may be so, or it may be different.
[0252] At least one monomer in the construct is a mutant monomer of the present invention. Two or more, three or more, four or more, five or more, six or more, seven or more, eight or more, nine or more Alternatively, 10 or more monomers may be mutant monomers of the present invention. Preferably, construct All monomers in the substance are mutant monomers of the present invention. Mutant monomers are the same It may or may not be. In a preferred embodiment, the structure is the present invention. It contains two mutant monomers.
[0253] The mutant monomers of the present invention in the construct are preferably of approximately the same length or the same length The barrels of the mutant monomers of the present invention in the construct are preferably of approximately the same length. They are either the same length or the same length. Length is measured in units of the number and / or length of amino acids. It is possible.
[0254] The construct may contain one or more monomers that are not mutant monomers of the present invention. The CsgG mutant monomers that are not non-mutant monomers of the present invention include any of the above discussed. The amino acids / positions of sequence numbers 390, 391, 392, 393, and 394 have not been mutated. ,395,414,415,416,417,418,419,420,421,422 , 423, 424, 425, 426, 427, 428 or 429 or Sequence ID 39 0, 391, 392, 393, 394, 395, 414, 415, 416, 417, 41 8, 419, 420, 421, 422, 423, 424, 425, 426, 427, 42 Examples include monomers containing 8 or 429 comparative variants. At least in the construct One monomer corresponds to sequence numbers 390, 391, 392, 393, 394, 395, and 414. , 415, 416, 417, 418, 419, 420, 421, 422, 423, 424 , may include 425, 426, 427, 428 or 429, or Sequence ID 3 90, 391, 392, 393, 394, 395, 414, 415, 416, 417, 4 18, 419, 420, 421, 422, 423, 424, 425, 426, 427, 4 May contain a comparison variant of the sequence indicated by 28 or 429. Sequence ID 390 ,391,392,393,394,395,414,415,416,417,418 ,419,420,421,422,423,424,425,426,427,428 Alternatively, the 429 comparison variant is based on amino acid identity and is 3 across its entire sequence. 90, 391, 392, 393, 394, 395, 414, 415, 416, 417, 4 18, 419, 420, 421, 422, 423, 424, 425, 426, 427, 4 It is at least 50% homologous to 28 or 429, 40 or 41. More preferably, ratio The comparative variants, based on amino acid identity, are sequence numbers 390 and 399 across the entire sequence. 1, 392, 393, 394, 395, 414, 415, 416, 417, 418, 41 9, 420, 421, 422, 423, 424, 425, 426, 427, 428 or 429 amino acid sequences and at least 55%, at least 60%, at least 65%, less At least 70%, at least 75%, at least 80%, at least 85%, at least 9 0%, and more preferably at least 95%, 97%, or 99% homology. ru.
[0255] The monomers in the construct are preferably genetically fused. The monomers are throughout the construct. When a body is expressed from a single polynucleotide sequence, it is genetically fused. The coding sequences of monomers are used to form a single polynucleotide sequence that codes for a substance. They can be combined in various ways.
[0256] Monomers may be genetically fused in any stereochemistry. They may also be fused by their terminal amino acids. For example, the amino acid of one monomer The terminal end may also be fused to the carboxyl terminus of another monomer. (amino) (from the carboxyl direction) the second and subsequent monomers have methionine at their amino terminus. It may contain n (each of its ends is fused to the carboxyl terminus of the preceding monomer). For example, M is a monomer (without an amino-terminal methionine), and mM is an amino-terminal monomer. If it is a monomer with methionine, the construct has the sequence M-mM, M-mM-mM, and It may contain M-mM-mM-mM. The presence of these methionines is usually in the whole structure. Polynucleotide encoding a second or subsequent monomer in the polynucleotide encoding the structure This results from the expression of the 5' start codon (i.e., ATG) of the creotide. The first monomer in the substance (from amino to carboxyl direction) may also contain methionine. (Currently, mM-mM, mM-mM-mM, or mM-mM-mM-mM).
[0257] Two or more monomers can also be directly and genetically fused with each other. - Preferably, they are genetically fused using a linker. Monomer mobility The linker may be designed to restrict the amino acid sequence (i.e., (This is a peptide linker.) You can use any of the peptide linkers discussed above. stomach.
[0258] In another preferred embodiment, the monomers are chemically fused. The two monomers are , where two parts are chemically bonded, for example, by a chemical crosslinking agent. They are chemically fused. Any of the chemical crosslinking agents discussed above may be used. The lucer binds to one or more cysteine residues introduced into the mutant monomer of the present invention. This can also happen. Alternatively, the linker may be attached to one end of one of the monomers in the construct. They may be combined.
[0259] When a construct contains various monomers, the concentration of the linker in the excess monomer is maintained. This prevents crosslinking between monomers. Alternatively, two linkers —A "key and keyhole" arrangement may be used. Only one end of each linker is connected to the other. In response, it can form a longer linker. The other end of that linker is each These react with various monomers. Such linkers are described in International Application No. PCT / GB10 / Specification No. 000132 (published as International Publication No. 2010 / 086602) It is described there.
[0260] Polynucleotides The present invention also provides polynucleotide sequences encoding the mutant monomers of the present invention. The isomonomer may be any of those discussed above. The polynucleotide sequence is Based on nucleotide identity, at least 50% of the entire sequence is identical to sequence number 389. Preferably containing sequences with 60%, 70%, 80%, 90%, or 95% homology. Three consecutive sequences Stretching of nucleotides above 00, for example, 375, 450, 525 or 600 or more Over a certain period, at least 80%, for example, at least 85%, 90%, or 95% nucleo There may also be a strong homology ("solid homology"). Homology is calculated as explained above. It is possible. The polynucleotide sequence is based on the degeneracy of the genetic code, as shown in Sequence ID No. 389. It may contain a different sequence.
[0261] The present invention also includes polynucleotide sequences encoding any of the gene fusion constructs of the present invention. The polynucleotide is preferably two or more of the sequences shown in SEQ ID NO: 389. Includes variants of the polynucleotide sequence, based on nucleotide identity, the entire sequence. Over SEQ ID NO: 389 and at least 50%, 60%, 70%, 80%, 90% or Preferably, it contains two or more sequences having 95% homology. 600 or more consecutive sequences, for example Across stretches of 750, 900, 1050, or 1200 or more nucleotides, At least 80%, for example, at least 85%, 90%, or 95% nucleotide identity ( There may also be "solid homology." Homology can be calculated as explained above. ru.
[0262] The polynucleotide sequence was derived and replicated using standard methods in the art. It is possible to encode the chromosomal DNA that codes for wild-type CsgG in Escherichia coli. It can be extracted from pore-producing organisms such as ) using PCR with specific primers. Then, the gene encoding the pore subunit can be amplified. The sequence may be subject to directional mutagenesis. A preferred directional mutagenesis method is... Known in the technical field, including chain reactions, for example, used in combination. The construct of the present invention Polynucleotides used are based on well-known techniques, for example, Sambrook, J. and Russell, D. (2001) ). Molecular Cloning: A Laboratory Manual, 3rd Edition. Cold Spring Harbor Labor To be manufactured using the techniques described in the atory press, Cold Spring Harbor, NY. It is possible.
[0263] Subsequently, the obtained polynucleotide sequence is used in a cloning vector or other replicable device. It can be incorporated into a recombinant vector. Using that vector, it can be used in compatible host cells. Polynucleotides can be replicated in this way. Thus, polynucleotide sequences are Introducing polynucleotides into a replicable vector, and introducing the vector into a suitable host. By doing so, and by growing host cells under conditions that induce vector replication. It can be produced. The vector can be recovered from the host cell. Polynucleotide Suitable host cells for cloning are known in the art, and such host cells The cells are described in more detail below.
[0264] Polynucleotide sequences can be cloned into suitable expression vectors. Polynucleotide sequences within a cell typically lead to the expression of coding sequences by the host cell. It is operablely linked to a control sequence that can be used. This allows the pore subunit to be expressed.
[0265] The term "operably linked" means that the components described function in their intended manner. This refers to juxtaposition, a relationship that makes this possible. A control array that is "operably linked" to a code array. Ligation is achieved so that the expression of the coding sequence is matched with the control sequence. Multiple copies of the same or different polynucleotide sequences are introduced into the vector. Sometimes that happens.
[0266] Subsequently, the expression vector can be introduced into a suitable host cell. The mutant monomer or construct of the invention inserts a polynucleotide sequence into an expression vector. This involves introducing the vector into compatible bacterial host cells and developing the polynucleotide sequence. It can be produced by growing host cells under conditions that induce the phenomenon. The recombinant expression monomers or constructs thereof can self-assemble within the host cell membrane to form pores. Alternatively, the recombinant pores thus produced can be removed from the host cells and inserted into another membrane. When producing pores containing at least two different monomers or constructs, the different monomers or constructs can be expressed separately in the different host cells as described above, removed from the host cells, and assembled into pores in another membrane, such as a rabbit cell membrane or a synthetic membrane.
[0267] Vectors can be, for example, plasmids, viruses or phage vectors that have, for example, an origin of replication, optionally a promoter for the expression of the polynucleotide sequence, and optionally a regulator of the promoter. Vectors may contain one or more selectable marker genes, such as a tetracycline resistance gene. Promoters and other expression regulatory signals can be selected to be compatible with the host cell in which the expression vector is designed. T7, trc, lac, ara or λ L promoters are usually used.
[0268] Host cells usually express monomers or constructs at high levels. Host cells transformed with the polynucleotide sequence will be selected to be compatible with the expression vector used to transform the cells. Host cells are usually bacterial cells, preferably Escherichia coli cells. λDE3 lysogens, such as C41(DE3), B L21(DE3), JM109(DE3), B834(DE3), TUNER, Orig Any cell containing ami and Origami B contains a vector containing the T7 promoter It can express the gene. In addition to the conditions listed above, Cao et al., 2014, PNAS, St structure of the nonameric bacterial amyloid secretion channel, doi - 1411942111 Goyal et al., 2014, Nature, 516, 250-253, structural and mechanistic insights The method mentioned in the bacterial amyloid secretion channel CsgG The CsgG protein can be expressed using either of these methods.
[0269] The present invention also includes methods for producing mutant monomers or constructs of the present invention. The method includes the step of expressing the polynucleotide of the present invention in a suitable host cell. The polynucleotide is preferably part of the vector and preferably a promoter. It is movably connected.
[0270] Poa The present invention also provides various pores. The pores of the present invention allow for the detection of different nucleotides with high sensitivity. Because it can be identified, it is used for characterizing polynucleotide sequences, for example, sequencing. It is ideal for this purpose. The pore can remarkably distinguish between 4 nucleotides in DNA and RNA. This can be done. The pore of the present invention distinguishes between methylated nucleotides and unmethylated nucleotides. It can even do that. The base resolution of the pore of the present invention is surprisingly high. The pore has all four It shows almost complete separation of DNA nucleotides. Furthermore, the pores show the residence time within the pores and Based on the current flowing through the pore, deoxycytidine monophosphate (dCMP) and methyl Identify -dCMP.
[0271] The pore of the present invention can also distinguish between different nucleotides under a wide range of conditions. In this case, the pore is under conditions that are favorable for characterizing nucleic acids, such as sequencing. This will allow for the identification of nucleotides from different pores. The extent to which this can be achieved depends on the applied potential, salt concentration, buffer solution, temperature, and additives, such as urea. This can be controlled by changing the presence of betaine and DTT. This makes it possible to fine-tune the pore function, especially during sequencing. This will be discussed in more detail below. Using the pore of the present invention, not per nucleotide Furthermore, we can identify polynucleotide polymers from their interactions with one or more monomers. It can also be done this way.
[0272] The pores of the present invention may be isolated, or substantially isolated, It may be manufactured or substantially purified. The pore of the present invention is If any other components such as lipids or other pores are completely absent, then they have been isolated or purified. The pore is mixed with a carrier or diluent that will not interfere with its intended use. In combination, they are substantially isolated. For example, pores are less than 10%, less than 5%, less than 2%, or Forms containing less than 1% of other components, such as triblock copolymers, lipids or other pores. If present, it is substantially isolated or substantially purified. Alternatively, the pore of the present invention These can also be present within the membrane. Suitable membranes are discussed below.
[0273] The pores of the present invention may exist as individual or single pores. A "poa" in the Ming Dynasty can also exist as a homogeneous or heterogeneous group of two or more poas.
[0274] Homo oligomer pore This invention relates to homooligomers derived from CsgG containing the same mutant monomer as the present invention. A is also provided. Homo-oligomeric pores may contain any of the variants of the present invention. The invention of homo-oligomeric pores is used for characterizing polynucleotides, for example, in sequencing. This is ideal. The homo-oligomeric pore of the present invention has any of the advantages discussed above. It is possible.
[0275] Homo-oligomeric pores may contain any number of mutant monomers. The pore is Typically, at least 7, at least 8, at least 9, or at least 10 identical The mutant monomers include, for example, seven, eight, nine, or ten mutant monomers. Preferably, it comprises eight or nine identical variant monomers. One or more, for example For example, two, three, four, five, six, seven, eight, nine, or ten variant monomers are Preferably, it is chemically modified as discussed above.
[0276] The method for creating the pore will be discussed in more detail below.
[0277] Hetero-oligomeric pore This invention relates to a heterozygous We also provide oligomeric pores. The hetero-oligomeric pores of the present invention have the characteristics of polynucleotides. For example, it is ideal for sequencing. Hetero-oligomeric pores are in the art. Prepared using a known method (e.g., Protein Sci. 2002 Jul; 11 (7): 1813-24). It is possible.
[0278] The heterooligomeric pore contains enough monomers to form that pore. These monomers may be of any type. The pore is usually at least 7 , at least 8, at least 9 or at least 10 monomers, for example, 7, It contains 8, 9, or 10 monomers. The pore preferably contains 8 or 9 monomers. —Includes.
[0279] In a preferred embodiment, all of the monomers (for example, 10, 9, and 8 of the monomers) are used. (7) are mutant monomers of the present invention, and at least one of them is different from the others. In a more preferred embodiment, the pore contains eight or nine variant monomers of the present invention. They include, and at least one of them is different from the others. They are all different from one another. Sometimes.
[0280] The mutant monomers of the present invention in the pore are preferably of approximately the same length or the same length. The barrels of the mutant monomers of the present invention in the pore are preferably of approximately the same length. Or they are the same length. Length is measured in units of the number and / or length of amino acids. It is possible.
[0281] In another preferred embodiment, at least one of the mutant monomers is the mutant monomer of the present invention. It is not Mer. The remaining monomers in this embodiment are preferably the mutant monomers of the present invention. Therefore, the pores are 9, 8, 7, 6, 5, 4, and 3 according to the present invention. It may contain two or one variant monomer. Any number of monomers in the pore is The monomers do not have to be the variant monomers of the invention. The pores are preferably seven or eight of the monomers of the invention. The present invention comprises mutant monomers and monomers that are not monomers of the present invention. They may be the same, or they may be different.
[0282] The mutant monomers of the present invention in the construct are preferably of approximately the same length or the same length The barrels of the mutant monomers of the present invention in the construct are preferably of approximately the same length. They are either the same length or the same length. Length is measured in units of the number and / or length of amino acids. It is possible.
[0283] The pore may contain one or more monomers that are not mutant monomers of the present invention. The CsgG monomers that are not mutant monomers of the present invention include any of the amino acids discussed above. / Position also mutated / not replaced SEQ ID NOs: 390, 391, 392, 393, 394, 3 95, 414, 415, 416, 417, 418, 419, 420, 421, 422, 4 23, 424, 425, 426, 427, 428 or 429 or sequence number 390, 391, 392, 393, 394, 395, 414, 415, 416, 417, 418, 419, 420, 421, 422, 423, 424, 425, 426, 427, 428 also Examples include monomers containing 429 comparative variants. SEQ ID NOs: 390, 391, 392, 393, 394, 395, 414, 415, 416, 417, 418, 419, 420, 421, 422, 423, 424, 425, 426, 427, 428 or 42 The 9 comparison variants are based on amino acid identity and are the same as Sequence ID No. 39 across their entire sequence. 0, 391, 392, 393, 394, 395, 414, 415, 416, 417, 41 8, 419, 420, 421, 422, 423, 424, 425, 426, 427, 42 It is typically at least 50% homologous to 8 or 429. More preferably, comparative variant Based on amino acid identity, sequence numbers 390, 391, and 392 throughout the entire sequence 393, 394, 395, 414, 415, 416, 417, 418, 419, 420, 421, 422, 423, 424, 425, 426, 427, 428 or 429 No acid sequence and at least 55%, at least 60%, at least 65%, at least 70% , at least 75%, at least 80%, at least 85%, at least 90%, and More preferably, they may be at least 95%, 97%, or 99% homologous.
[0284] In all embodiments discussed above, one or more mutant monomers, for example, two, Three, four, five, six, seven, eight, nine, or ten are preferred, as discussed above. It is chemically modified.
[0285] The method for creating the pore will be discussed in more detail below.
[0286] Structure-containing pore The present invention also provides a pore comprising at least one construct of the present invention. It contains two or more covalently bonded monomers derived from CsgG, and at least one monomer Another is the mutant monomer of the present invention. In other words, the construct is one more than one monomer. It must contain a mer. The pore must have sufficient structure and to form that pore. , and optionally contain monomers. For example, octamerpore (a) each contains two (b) may include four constructs, each containing four monomers It may also include a construct, or (b) one construct containing two monomers and the construct It may also contain six monomers that do not form part of the molecule. For example, the nanomerpore is (a) Each of the four constructs contains two constructs, and one monomer does not form part of the construct. (b) two constructs, each containing four monomers, and a portion of the construct (b) a single construction comprising two monomers It may also include a substance and seven monomers that do not form part of the structure. Structure and monomers Those skilled in the art can predict other combinations.
[0287] At least two monomers in the pore are in the form of the construct of the present invention. Therefore, the pore contains at least one mutant monomer of the present invention. The pore is total And at least 7, at least 8, at least 9 or at least 10 monomers For example, it typically contains 7, 8, 9, or 10 monomers (at least one monomer). (Two must be in the structure). The pore preferably has 8 or 9 monomers. —Includes (at least two monomers must be present in the construct).
[0288] Pores containing the same construct may also be homooligomers (i.e., they may contain the same construct). or heterooligomers (i.e., in which at least one construct is different from the others) Sometimes it happens.
[0289] A pore typically consists of (a) one construct containing two monomers and (b) five, six, It contains seven or eight monomers. The construct is any of the ones discussed above. Good. The monomers are the mutant monomers of the present invention, SEQ ID NOs: 390, 391, 392, 393 ,394,395,414,415,416,417,418,419,420,421 Monoma including 422, 423, 424, 425, 426, 427, 428 or 429 - and sequence numbers 390, 391, 392, 393, 394, and 395 as discussed above. , 414, 415, 416, 417, 418, 419, 420, 421, 422, 423 , variants including comparison variants 424, 425, 426, 427, 428 or 429 This may include any of the above, including monomers.
[0290] Another typical pore is more than one of the structures of the present invention, for example, two, three of the present invention. or it includes four structures. If necessary, such pores are formed to form that pore. Further comprising sufficient additional monomers or constructs. The additional monomers are variations of the present invention. Monomers, SEQ ID NOs: 390, 391, 392, 393, 394, 395, 414, 41 5, 416, 417, 418, 419, 420, 421, 422, 423, 424, 42 Monomers containing 5, 426, 427, 428 or 429, and sequences as discussed above. Numbers 390, 391, 392, 393, 394, 395, 414, 415, 416, 41 7, 418, 419, 420, 421, 422, 423, 424, 425, 426, 42 7, including the variant monomers containing comparative variants of 428 or 429, as discussed above. It may be any of the above. Further constructions may be any of the above. , or sequence numbers 390, 391, 392, 393, 394, 395, and 41 discussed above. 4, 415, 416, 417, 418, 419, 420, 421, 422, 423, 42 Monomers containing 4, 425, 426, 427, 428, or 429 or SEQ ID NO: 39 0, 391, 392, 393, 394, 395, 414, 415, 416, 417, 41 8, 419, 420, 421, 422, 423, 424, 425, 426, 427, 42 Two or more covalently linked CsgG models, each containing 8 or 429 comparison variants. It may also be a structure that includes a nomar.
[0291] Further pores of the present invention include only a construct comprising two monomers, for example, the pore is It may contain four, five, six, seven, or eight constructs, each containing two monomers. At least one structure is a structure of the present invention, that is, in at least one structure At least one monomer, preferably each monomer in at least one construct, These are mutant monomers of the present invention. All constructs containing the two monomers are constructs of the present invention. It can also be the case.
[0292] The specific pore according to the present invention is one of four constructs of the present invention, each containing two monomers. And at least one monomer in each construct, preferably each monomer in each construct is It contains four constructs, which are mutant monomers of the original compound. The constructs are oligomerized, and each construct It is also possible to create a pore with a structure in which only one monomer contributes to the pore's channel. Yes, it's possible. Typically, the other monomers in the construct will be outside the channel of the pore. Example For example, the pore of the present invention is a structure of 7, 8, 9, or 10 buildings containing 2 monomers. This may include constructs containing 7, 8, 9, or 10 monomers in the channel. ru.
[0293] Mutations can be introduced into the above constructs. These mutations may alternate. Furthermore, the mutations of each monomer within the two monomer constructs are different, and the constructs are homooligomeric. It is assembled in this way, resulting in alternating modifications. In other words, it includes MutA and MutB. The monomers are fused together and assembled to form an AB:AB:AB:AB pore. Alternatively, mutations may be adjacent. That is, the same mutation may be present in the construct. It is introduced into two monomers, and then this construct is used with different mutant monomers or constructs. It is oligomerized using this method. In other words, monomers containing MutA are fused, and then Using MutB-containing monomers, ori It gets turned into sesame seeds.
[0294] One or more monomers of the present invention in the construct-containing pore are chemically as discussed above. It can also be modified.
[0295] Analyte characterization This invention provides a method for determining the presence, absence, or one or more characteristics of a target analyte. This method provides a target analyte and a CsgG pore or its variant, for example, the pore of the present invention. The target analyte is brought into contact with the pore, for example, so that it moves through the pore. The ap and analyte are moving through the pore, and one or more measurements are taken at each step. The process includes a step of determining the presence or absence of the analyte or one or more of its characteristics. The target analyte is sometimes called the template analyte or the analyte of interest.
[0296] The method involves using a target analyte and a CsgG pore or its variant, such as the pore of the present invention, as the target. The process includes bringing the analyte into contact with the pore so that it moves through it. The pore is typically small. at least 7, at least 8, at least 9 or at least 10 monomers, for example , containing 7, 8, 9 or 10 monomers. The pores are preferably 8 or 9 It contains the same monomer. Preferably, one or more, for example, 2, 3, 4, 5, 6, 7. The monomers 8, 9, or 10 are chemically modified as discussed above.
[0297] The CsgG pore may originate from any organism. The CsgG pore is sequence number 390. 391, 392, 393, 394, 395, 414, 415, 416, 417, 418, 419, 420, 421, 422, 423, 424, 425, 426, 427, 428 It may contain monomers that include the sequence shown at 429. CsgG pore is sequence number 390, 391, 392, 393, 394, 395, 414, 415, 416, 417, 418, 419, 420, 421, 422, 423, 424, 425, 426, 427, Each of the sequences is represented by 428 or 429, and each contains at least 7, at least 8, and a small number of at least nine or ten monomers, for example, seven, eight, nine or ten It may contain monomers (i.e., the pore contains the same monomer from the same organism). (It is a mooligomer). CsgG is sequence numbers 390, 391, 392, 393, 394 ,395,414,415,416,417,418,419,420,421,422 The sequences indicated by 423, 424, 425, 426, 427, 428, or 429, respectively It may also contain any combination of monomers, including those included in POA. For example, POA is in SEQ ID NO: 390 Seven monomers containing the sequence shown, and two monomers containing the sequence shown in sequence number 391. It may include "mar".
[0298] The CsgG mutant can have any number of mutant monomers, for example, at least 7, and at least Also 8, at least 9 or at least 10 monomers, for example 7, 8, 9 It may contain 10 monomers. Mutant monomers include SEQ ID NOs: 390, 391, 3 92, 393, 394, 395, 414, 415, 416, 417, 418, 419, 4 20, 421, 422, 423, 424, 425, 426, 427, 428 or 429 It may also include comparison variants of the sequence shown. Comparison variants are discussed above. The comparative variant must be able to form a pore, with respect to the pore of the present invention. It may possess one of the homology percentages discussed above.
[0299] The CsgG variant preferably comprises nine monomers, and at least of those monomers One is a variant of the sequence shown in sequence number 390, where (a) one of the following positions is also Variations in or in multiple locations: N40, Q42, D43, E44, K49, Y51, S54, N55, F56, S57, Q62, E101, N102, E124, E131, R142 , D149, T150, E185, R192, D195, E201 and E203, for example If so, mutations at one or more of the following positions: N40, Q42, D43, E44, K4 9, Y51, S54, N55, F56, S57, Q62, E101, N102, E131 , D149, T150, E185, D195, E201 and E203, or the following positions One or more mutations in one or more of the following: N40, Q42, D43, E44 , K49, Y51, N55, F56, E101, N102, E131, D149, T15 0, E185, D195, E201 and E203, and / or (b) the following positions One or more missing characters: F48, K49, P50, Y51, P52, A53, S54 The variant includes N55, F56 and S57. The variant includes (a) (a) may include (b), or it may include both (a) and (b). In this case, positions N40, Q42, D43, E44, K49, Y51, S54, N55, F 56, S57, Q62, E101, N102, R124, E131, R142, D149 Any number and combination of E185, R192, D195, E201 and E203 They may be mutated. As discussed above, one of Y51, N55, and F56 Alternatively, multiple mutations contribute to the current that occurs when polynucleotides move through the pore. By reducing the number of rheosides, polynucleotides are moving through the pore. This makes it easier to identify the direct relationship between the observed current and polynucleotides. .
[0300] (b) F48, K49, P50, Y51, P52, A53, S54, N55 Any number and combination of F56 and S57 may be missing. The expression refers to the specific mutations / substitutions or combinations thereof discussed above with respect to the mutant monomers of the present invention. It may include any of the following:
[0301] A preferred variant for use in the method of the present invention includes one or more of the following substitutions: ( a) F56N, F56Q, F56R, F56S, F56G, F56A or F56K These are F56A, F56P, F56R, F56H, F56S, F56Q, F56I, F56L , F56T or F56G, (b)N55Q, N55R, N55K, N55S, N55G , N55A or N55T, (c)Y51L, Y51V, Y51A, Y51N, Y51Q , Y51S or Y51G, (d)T150I, (e)S54P, and (f)S57P The variant may include any number and combination of (a) through (f).
[0302] Preferred variants for use in the method of the present invention include Q62R or Q62K.
[0303] Preferred variants for use in the method of the present invention are D43, E44, Q62 or these. Any combination of these, for example, D43, E44, Q62, D43 / E44, D43 / Q62, This includes mutations in E44 / Q62 or D43 / E44 / Q62.
[0304] The variant is, Examples of mutations at positions Y51, F56, D149, E185, E201, and E203. For example, Y51N, F56A, D149N, E185R, E201N and E203N, A mutation at position N55, for example, N55A or N55S, A mutation at position Y51, for example, Y51N or Y51T. A mutation at position S54, for example, S54P. A mutation at position S57, for example, S57P, Variations at position F56, e.g., F56N, F56Q, F56R, F56S, F56 G, F56A or F56K or F56A, F56P, F56R, F56H, F5 6S, F56Q, F56I, F56L, F56T or F56G, Mutations at positions Y51 and F56, for example, Y51A and F56A, Y51A and F56N, Y51I and F56A, Y51L and F56A, Y51T and F5 6A, Y51T and F56Q, Y51I and F56N, Y51L and F56N, also Alternatively, Y51T and F56N, preferably Y51I and F56A, Y51L and F 56A, or Y51T and F56A, more preferably Y51T and F56Q, More preferably, Y51X and F56Q (where X is any amino acid), Variations at positions N55 and F56, for example, N55X and F56Q (where X is (Any amino acid), Mutations at positions Y51, N55, and F56, e.g., Y51A, N55S, and F 56A, Y51A, N55S and F56N, or Y51T, N55S and F56 Q, Mutations at positions S54 and F56, for example, S54P and F56A, or S 54P and F56N, Variations at positions F56 and S57, for example, F56A and S57P, or F 56N and S57P, Mutations at positions D149, E185, and E203, e.g., D149N, E185 N and E203N, Mutations at positions D149, E185, E201, and E203, e.g., D149N , E185N, E201N and E203N, Mutations at positions D149, E185, D195, E201, and E203, for example, D149N, E185R, D195N, E201N and E203N, or D149 N, E185R, D195N, E201R and E203N, Mutations at positions F56 and N102, e.g., F56Q and N102R, (a) Variations at position Q62, e.g., Q62R or Q62K, and (b) position One or more of the positions Y51, N55, and F56, for example, positions Y51, N55, F5 6. Y51 / N55, Y51 / F56, N55 / F56 or Y51 / N55 / F56 Mutations in, for example, Y51T / F56Q / Q62R, (i) Positions D43, E44, Q62 or any combination thereof, for example, position D4 3, E44, Q62, D43 / E44, D43 / Q62, E44 / Q62 or D43 (ii) a mutation at / E44 / Q62, and (ii) one of the positions Y51, N55, and F56 Alternatively, multiple positions, for example, Y51, N55, F56, Y51 / N55, Y51 / F56 , mutations in N55 / F56 or Y51 / N55 / F56, for example, D43N / Y 51T / F56Q, E44N / Y51T / F56Q, D43N / E44N / Y51T / F 56Q, D43N / Y51T / F56Q / Q62R, E44N / Y51T / F56Q / Q 62R, or D43N / E44N / Y51T / F56Q / Q62R, A mutation at position T150, e.g., T150I It may contain.
[0305] The pores preferred for use in the method of the present invention are at least 7, at least 8, and at least Also 9 or at least 10 monomers, for example 7, 8, 9 or 10 monomers -Includes, and each monomer is one of the following substitutions: (a) F56N, F56Q, F56R, F56S F56G, F56A or F56K or F56A, F56P, F56R, F56H, F56S, F56Q, F56I, F56L, F56T or F56G, (b)N55Q, N55R, N55K, N55S, N55G, N55A or N55T, (c)Y51L, Y51V, Y51A, Y51N, Y51Q, Y51S or Y51G, (d)T150I Sequence ID 39, which includes one or more of (e)S54P and (f)S57P. Includes variants of 0. Variants include all numbers and combinations of (a) through (f). This may occur. The monomers in these preferred pores are preferably identical.
[0306] The pores preferred for use in the method of the present invention are at least 7, at least 8, and at least Also 9 or at least 10 monomers, for example 7, 8, 9 or 10 monomers —Includes, and each monomer is
[0307] [ka] Includes variant of sequence number 390, which includes [the specified character].
[0308] The CsgG mutant for use in the present invention preferably comprises nine monomers, and At least one monomer is located at one or more positions Y51, N55, and F56. A variant of the sequence shown in SEQ ID NO: 390, including a mutation. Use in the method of the present invention A preferred pore is at least 7, at least 8, at least 9 or at least It contains 10 monomers, for example, 7, 8, 9 or 10 monomers, and each monomer Sequence ID 39 contains mutations in one or more of the following positions: Y51, N55, and F56. It includes variants of 0. The monomers in these preferred pores are preferably identical. The variants are Y51, N55, F56, Y51 / N55, Y51 / F56, N55 / F May contain mutations in 56, or Y51 / N55 / F56. The variant is above In one or more of the positions Y51, N55 and F56 discussed, and every The combination may include one of the specific mutations: Y51, N55, and F56. One or more amino acids may be substituted with any amino acid. Y51 is F, M, L, I V, A, P, G, C, Q, N, T, S, E, D, K, H or R, for example, A, S, T It may be replaced with N or Q. N55 is F, M, L, I, V, A, P, G , replaced by C, Q, T, S, E, D, K, H or R, for example A, S or T or Q They sometimes exist. F56 is M, L, I, V, A, P, G, C, Q, N, T, S, E, D It may be replaced by K, H, or R, for example, A, S, T, N, or Q. An ant may further include one or more of the following: (i) one in the following positions or multiple mutations (i.e., mutations at one or more of the following positions) (i) N40 , D43, E44, S54, S57, Q62, R97, E101, E124, E131, R142, T150 and R192; (iii) Q42R or Q42K; (iv) K4 9R;(v)N102R, N102F, N102Y or N102W;(vi)D149 N, D149Q or D149R; (vii) E185N, E185Q or E185R (viii) D195N, D195Q or D195R; (ix) E201N, E20 1Q or E201R; (x)E203N, E203Q or E203R; and (x i) Deletion of one or more of the following positions: F48, K49, P50, Y51, P52, A 53, S54, N55, F56 and S57. The variants are (i) and ( iii) May include any combination of (xi). The variant is (i) and (iii) to (xi) may include any of the embodiments discussed above.
[0309] (1) A variant preferred for use in the method of the present invention is, or (2) in the method of the present invention The preferred pores for use are at least 7, at least 8, at least 9, or fewer. Each contains 10 monomers, for example, 7, 8, 9 or 10 monomers, Each of these is Y51R / F56Q, Y51N / F56N, Y51M / F56Q, Y51L / F56Q, Y51I / F56Q, Y51V / F56Q, Y51A / F56Q, Y51P / F56Q, Y51G / F56Q, Y51C / F56Q, Y51Q / F56Q, Y51N / F56Q, Y51S / F56Q, Y51E / F56Q, Y51D / F56Q, Y51K / Includes variants of sequence number 390, including F56Q or Y51H / F56Q.
[0310] (1) A variant preferred for use in the method of the present invention is, or (2) in the method of the present invention The preferred pores for use are at least 7, at least 8, at least 9, or fewer. Each contains 10 monomers, for example, 7, 8, 9 or 10 monomers, Each of these includes Y51T / F56Q, Y51Q / F56Q, or Y51A / F56Q. , including variant number 390.
[0311] (1) A variant preferred for use in the method of the present invention is, or (2) in the method of the present invention The preferred pores for use are at least 7, at least 8, at least 9, or fewer. Each contains 10 monomers, for example, 7, 8, 9 or 10 monomers, Each of these is Y51T / F56F, Y51T / F56M, Y51T / F56L, Y51T / F56I, Y51T / F56V, Y51T / F56A, Y51T / F56P, Y51T / F56G, Y51T / F56C, Y51T / F56Q, Y51T / F56N, Y51T / F56T, Y51T / F56S, Y51T / F56E, Y51T / F56D, Y51T / The sequence number 390 includes F56K, Y51T / F56H, or Y51T / F56R. Includes riant.
[0312] (1) A variant preferred for use in the method of the present invention is, or (2) in the method of the present invention The preferred pores for use are at least 7, at least 8, at least 9, or fewer. Each contains 10 monomers, for example, 7, 8, 9 or 10 monomers, Each of these includes Y51T / N55Q, Y51T / N55S, or Y51T / N55A. , including variant number 390.
[0313] (1) A variant preferred for use in the method of the present invention is, or (2) in the method of the present invention The preferred pores for use are at least 7, at least 8, at least 9, or fewer. Each contains 10 monomers, for example, 7, 8, 9 or 10 monomers, Each of these is Y51A / F56F, Y51A / F56L, Y51A / F56I, Y51A / F56V, Y51A / F56A, Y51A / F56P, Y51A / F56G, Y51A / F56C, Y51A / F56Q, Y51A / F56N, Y51A / F56T, Y51A / F56S, Y51A / F56E, Y51A / F56D, Y51A / F56K, Y51A / Includes variants of sequence number 390, including F56H or Y51A / F56R.
[0314] (1) A variant preferred for use in the method of the present invention is, or (2) in the method of the present invention The preferred pores for use are at least 7, at least 8, at least 9, or fewer. Each contains 10 monomers, for example, 7, 8, 9 or 10 monomers, Each of these is Y51C / F56A, Y51E / F56A, Y51D / F56A, Y51K / F56A, Y51H / F56A, Y51Q / F56A, Y51N / F56A, Y51S / The sequence number 390 includes F56A, Y51P / F56A, or Y51V / F56A. Includes riant.
[0315] (1) A variant preferred for use in the method of the present invention is, or (2) in the method of the present invention The preferred pores for use are at least 7, at least 8, at least 9, or fewer. Each contains 10 monomers, for example, 7, 8, 9 or 10 monomers, Each of them,
[0316] [ka] Includes variant of sequence number 390, which includes [the specified character].
[0317] (1) A variant preferred for use in the method of the present invention is, or (2) in the method of the present invention The preferred pores for use are at least 7, at least 8, at least 9, or fewer. Each contains 10 monomers, for example, 7, 8, 9 or 10 monomers, Each of them,
[0318] [ka] Includes variant of sequence number 390, which includes [the specified character].
[0319] The monomers in these preferred pores are preferably identical.
[0320] The pores preferred for use in the method of the present invention are at least 7, at least 8, and at least Also 9 or at least 10 monomers, for example 7, 8, 9 or 10 monomers —Includes each monomer, F56A, F56P, F56R, F56H, F56S, F Barriers of sequence number 390, including 56Q, F56I, F56L, F56T, or F56G It contains . The monomers in these preferred pores are preferably the same.
[0321] The CsgG variant is any of the variants of Sequence ID No. 390 disclosed in this embodiment. It may include, or may include, any of the pores disclosed in this embodiment.
[0322] The CsgG variant is most preferably the pore of the present invention.
[0323] Steps (a) and (b) are preferably carried out using a potential applied through the film. As will be discussed in more detail below, depending on the applied potential, pores and polynuclei are usually formed. This results in the formation of a complex with rheotide-binding proteins. The applied potential is voltage. It may be a position. Or, the applied potential may be a chemical potential. One example of this is the use of a salt gradient across the amphiphilic layer. The salt gradient is described by Holden et al. This is disclosed in ., J Am Chem Soc. 2007 Jul 11; 129(27):8650-5.
[0324] The method is a method for determining the presence, absence, or one or more characteristics of the target analyte. The method determines the presence or absence of at least one target analyte or one or more characteristics. This method may be determined by the presence or absence of two or more target analytes, or by the presence or absence of one. Or it may involve determining multiple features. The method can be any number of analytes, for example, 2 The presence or absence of 5, 10, 15, 20, 30, 40, 50, 100 or more analytes or 1 It may include the determination of one or more features. Any number of one or more analytes. It can determine the characteristics of a product, for example, 1, 2, 3, 4, 5, or 10 or more characteristics.
[0325] The target analytes are preferably metal ions, inorganic salts, polymers, amino acids, peptides, and polymers. Lipeptides, proteins, nucleotides, oligonucleotides, polynucleotides, pigments , bleach, pharmaceuticals, diagnostic agents, recreational drugs, explosives or environmental pollutants Yes, there is. The method involves two or more analytes of the same type, for example, two or more proteins, or two or more The presence or absence of the above nucleotides or two or more pharmaceuticals, or one or more characteristics It may also relate to the determination. Alternatively, the method may involve two or more different types of analytes, for example. , one or more proteins, one or more nucleotides and one or This may also involve determining the presence or absence of multiple drugs, or the characteristics of one or more drugs.
[0326] The target analyte can be secreted from the cell. Alternatively, the target analyte may be present within the cell. It may also be an analyte, and therefore, before carrying out the present invention, the analyte is extracted from the cells. It must.
[0327] The analytes are preferably amino acids, peptides, polypeptides, and / or proteins. Amino acids, peptides, polypeptides, or proteins can also be found in nature. They may or may not be present in nature. Polypeptides or proteins are synthesized or It can contain modified amino acids. There are many different types of modifications to amino acids. This is known in the art. The preferred amino acids and their modifications are as described above. The objective of the invention is to modify the target analyte by any method available in the art. It should be understood that it can be used for decoration.
[0328] Proteins include enzymes, antibodies, hormones, growth factors or growth regulatory proteins, for example, It can be an itokine. The cytokine is an interleukin, preferably an IFN. -1, IL-1, IL-2, IL-4, IL-5, IL-6, IL-10, IL-12 and IL-13, interferon, preferably IL-γ, and other cytokines, For example, TNF-α can be selected. Proteins include bacterial proteins and fungal proteins. It may be a protein, viral protein, or parasite-derived protein.
[0329] The target analyte is preferably a nucleotide, oligonucleotide, or polynucleotide. It is. Nucleotides and polynucleotides will be discussed below. Oligonucleotides are Usually, the number of nucleotides is 50 or less, for example, 40 or less, 30 or less, 20 or less, 10 or less. Oligonucleo The nucleotides include any of the nucleotides discussed below, including debased and modified nucleotides. Sometimes it happens.
[0330] Target analytes, such as target polynucleotides, are present in one of the preferred samples discussed below. It may exist.
[0331] Pores are typically located within membranes as discussed below. Using the method discussed below, target fractions are... The precipitate can be coupled to the membrane or delivered to the membrane.
[0332] Using any of the measurements discussed below, determine the presence, absence, or one or more of the target analytes. The numerical characteristics can be determined. Preferably, the target analyte and the pore are subjected to the analyte The steps of bringing the object into contact with the pore, for example, so that it moves through the pore, and analysis By measuring the current passing through a pore as an object moves through it, the analyte is identified. The process includes a step of determining the presence or absence of or one or more of the features.
[0333] The target analyte is determined when current flows through the pore in a manner specific to the analyte (i.e., pore If a specific current associated with the analyte is detected flowing through A, then it is present. The analyte is absent if the current does not flow through the pore in a manner specific to the nucleotides. Yes. A control experiment was conducted in the presence of the analyte, and the analyte's response to the current flowing through the pore was... It is possible to determine how the influence is exerted.
[0334] Using the present invention, the differences that similar structures impart to the current passing through the pores of analytes are obtained. They can be distinguished based on their influence. Individual analytes can be examined at the single-molecule level, and they can be distinguished. They can be identified from the current amplitudes when they interact with the pore. Using the present invention Furthermore, it is possible to determine whether or not a specific analyte is present in the sample. It is also possible to measure the concentration of a specific analyte in the sample. Other pores besides CsgG are used. The characterization of the analytes is well known in the art.
[0335] Characterization of polynucleotides This invention characterizes target polynucleotides, for example, by sequencing polynucleotides. A method for characterizing polynucleotides using nanopores or Two main strategies for sequencing: chain characterization / sequencing and exonuclease characterization / sequencing. The method of the present invention is both It may depend on the method.
[0336] In strand sequencing, DNA passes through nanopores either in response to or repelled by the applied potential. This is how it is transitioned. Exonucleas act progressively or progressively on double-stranded DNA. -se can be used on the cis side of the pore to deliver the remaining single strand under the applied potential. It can be used on the transformer side, or under reverse potential conditions. Similarly, double-stranded DN Helicases that rewind A can be used in the same way. Polymerases can also be used. It is possible. For sequencing applications, chain transfer may be required against the applied potential. However, under reversed potential or in the absence of potential, DNA is initially " It should be "captured". After that, when the potential switches after binding, the chain changes from cis to tra It passes through the pore into the pore and is held by the electric current in the extended higher-order structure. Single-stranded DNA Xonucleases, or single-stranded DNA-dependent polymerases, act as molecular motors. Then, the single strand that has just been transferred through the pore is repelled by the applied potential and moves from transformer to cis. It can be reversed in a controlled, gradual manner.
[0337] In one embodiment, a method for characterizing a target polynucleotide involves a target sequence and pores and hematologic cells. The method includes a step of contacting with a helicase enzyme. Any helicase can be used in the method. This is possible. Suitable helicases are discussed below. Helicases have two properties for the pore. It can act in modes. Firstly, the method is preferably to which helicase is applied. To control the movement of the target array through the pore in accordance with the field resulting from the voltage, helical This is done using an enzyme. In this mode, the 5' end of the DNA is first trapped in the pore, The enzyme moves the target sequence along with the environment until the target sequence is finally moved to the trans side of the bilayer. Control the movement of DNA into the pore so that it passes through A. Alternatively, the method is preferably, The helicase enzyme moves the target sequence through the pore in response to the field generated as a result of the applied voltage. It is done in a controlled manner. In this mode, the 3' end of the DNA is first trapped in the pore, The enzyme resists the applied field until the target sequence is eventually pushed back to the cis side of the bilayer. It controls the movement of DNA through the pore, as if it were being pulled out of it.
[0338] In exonuclease sequencing, the exonuclease targets the polynucleotide. The individual nucleotides are released from one end of the ocide, and these individual nucleotides It is identified as discussed below. In another embodiment, the method for characterizing the target polynucleotide The method includes the step of contacting the target sequence with the pore and exonuclease enzyme. Any of the exonucleases discussed below can be used in the method. As discussed, it is possible to covalently bond to the pore.
[0339] Exonucleases typically firmly hold onto one end of a polynucleotide. It is an enzyme that digests the sequence one nucleotide at a time from its end. Ze can digest polynucleotides in the 5'-to-3' direction or the 3'-to-5' direction. Yes, it is possible. The ends of the polynucleotide to which the exonuclease binds are usually the yeast used. This is determined by selection of elements and / or by using methods known in the art. Using a hydroxyl group or cap structure at any end of a polynucleotide, Preventing or promoting the binding of exonucleases to specific ends of a dinucleotide It's usually possible.
[0340] The method involves using an exonuclease on polynucleotides, and the nucleotides are polynucleotides. The rate at which the characterization or identification of nucleotides in the proportions discussed above can be made from the terminals The method includes bringing the substance into contact with the substance so that it can be digested. A method for doing this is available in the art. This is well known. For example, Edman degradation can be used to extract a single amino acid from the end of a polypeptide. They are digested sequentially, and as a result, high-performance liquid chromatography (HPLC) is used to combine them. It can be determined. Similar methods can be used in the present invention.
[0341] The rate at which exonuclease functions is usually the optimal rate for wild-type exonuclease. It is too slow. The preferred rate for exonuclease activity in the method of the present invention is 0.5 to 10 per second. 00 nucleotides, 0.6-500 nucleotides per second, 0.7-200 nucleotides per second , 0.8-100 nucleotides per second, 0.9-50 nucleotides per second, or 1-2 nucleotides per second This involves the digestion of 0 or 10 nucleotides. The rate is preferably 1, 10, or 1 per second. The nucleotides are 00, 500, or 1000. Preferred rate of exonuclease activity. This can be achieved in various ways. For example, a variant in which the optimal rate of activity is reduced. Exonuclease can be used in accordance with the present invention.
[0342] In the chain characterization embodiment, the method involves a polynucleotide in a CsgG pore or its variant. For example, the pore of the present invention and the polynucleotide move through the pore, for example. The steps involve bringing the pore into contact, and as the polynucleotide moves relative to the pore. Perform one or more measurements, and if the measurement reveals one or more characteristics of the polynucleotide The method includes the step of demonstrating and thereby characterizing the target polynucleotide.
[0343] In an embodiment of exonucleotide characterization, the method involves characterizing a polynucleotide with CsgG A or its variants, for example, the pore of the present invention, and the exonuclease, and exonuclease The enzyme digests individual nucleotides from one end of the target polynucleotide, and then the individual nucleotides The step of bringing the creotide into contact with the pore, for example, so that it moves through the pore, And one or more measurements are taken as individual nucleotides move toward the pore. The measurement is performed to reveal one or more characteristics of individual nucleotides, thereby targeting the polynucleotide. This includes a step of characterizing nucleotides.
[0344] Each individual nucleotide is a single nucleotide. Nucleo It is a nucleotide bond. A nucleotide bond is a bond between a sugar group of one nucleotide and another nucleotide. Each nucleotide typically contains at least 5, at least 1 phosphate groups. 0, at least 20, at least 50, at least 100, at least 200, at least Another polynucleotide of 500, at least 1000 or at least 5000 nucleotides Those that are not bound by nucleotides that are bound to ocid. For example, individually The nucleotides are digested from a target polynucleotide sequence such as a DNA or RNA chain. The nucleotide may be any of those discussed below.
[0345] Individual nucleotides can interact with pores in any way and at any site. Yes, it is possible. The nucleotides are preferably transferred by an adapter as discussed above. It reversibly binds to the pore with an adapter such as the nucleotide, most preferably, When they pass through the pore and through the membrane, by or with the adapter They also bind reversibly to the pore. Nucleotides pass through the membrane through the pore. Sometimes, by or together with an adapter, the pore barrel or channel is reversible. They can also be combined.
[0346] During the interaction between individual nucleotides and pores, the nucleotides are specific to that nucleotide. It is common for it to act on the current flowing through the pore in a specific manner. For example, a specific nucleus Otid reduces the current flowing through the pore to a specific degree over a specific average time. In other words, the current flowing through the pore is specific to a particular nucleotide. This is the result. A control experiment was conducted to determine the effect of a specific nucleotide on the electric current flowing through the pore. The effect can be determined. Subsequently, the results of implementing the method of the present invention on the test sample can be determined. By comparing with those derived from such control experiments, specific nucleotides in the sample can be identified. It is possible to determine whether a particular nucleotide is present in the sample. This is possible. A frequency that influences the current flowing through the pore in a way that indicates a specific nucleotide. Using this method, the concentration of that nucleotide in the sample can be determined. The ratio of nucleotides can also be calculated. For example, the ratio of methyl-dCMP in dCMP The ratio can be calculated.
[0347] The method includes the step of measuring one or more characteristics of a target polynucleotide. The target polynucleotide is called the template polynucleotide or the target polynucleotide. Sometimes they do.
[0348] This embodiment also uses a CsgG pore or a variant thereof, such as the pore of the present invention. With respect to the analyte, either of the pores and embodiments discussed above can be used.
[0349] Polynucleotides Polynucleotides, such as nucleic acids, are macromolecules containing two or more nucleotides. A cleotide or nucleic acid may contain any combination of any nucleotide. Cleotides can occur naturally or are artificially produced. One or more nucleotides in a nucleotide may also be oxidized or methylated. One or more nucleotides in a polynucleotide can also be damaged. For example, polynucleotides may contain pyrimidine dimers. These are generally damaged by ultraviolet radiation and are the main cause of skin melanoma. One or more nucleotides may be modified, for example, with labels or tags. Suitable labeling is described below. Polynucleotides contain one or more spacers. It is also said.
[0350] Nucleotides typically contain a nucleic acid base, a sugar, and at least one phosphate group. Nucleic acid bases and sugars form nucleosides.
[0351] Nucleic acid bases are usually heterocyclic. Examples of nucleic acid bases include purines and pyrimidins. More specifically, adenine (A), guanine (G), thymidine (T), uracil (U) Examples include, but are not limited to, cytosine (C).
[0352] Sugars are usually pentose sugars. Nucleotide sugars include ribose and dextrose. Examples include, but are not limited to, roses. The sugar is preferably deoxyribose. ru.
[0353] Polynucleotides preferably contain the following nucleoside: deoxyadenosine (dA ), deoxyuridine (dU) and / or thymidine (dT), deoxyguanosine ( dG) and deoxycytidine (dC).
[0354] Nucleotides are typically ribonucleotides or deoxyribonucleotides. Nucleotides typically contain monophosphate, diphosphate, or triphosphate. It may also contain more than three phosphates, for example, four or five phosphates. It can be attached to the 5' end of the nucleotide, or it can be attached to the 3' end. The nucleotides include adenosine monophosphate (AMP) and guanosine monophosphate (GM). P), thymidine monophosphate (TMP), uridine monophosphate (UMP), 5-methylcytidine Monophosphate, 5-hydroxymethylcytidine monophosphate, cytidine monophosphate (CMP), cyclic Adenosine monophosphate (cAMP), cyclic guanosine monophosphate (cGMP), deoxyadenosine Deoxyguanosine monophosphate (dAMP), deoxyguanosine monophosphate (dGMP), deoxythyme Dinine monophosphate (dTMP), deoxyuridine monophosphate (dUMP), deoxycytidine Examples include monophosphate (dCMP) and deoxymethylcytidine monophosphate, but these Not limited to. Nucleotides are preferably AMP, TMP, GMP, CMP, UMP Selected from dAMP, dTMP, dGMP, dCMP, and dUMP.
[0355] Nucleotides can be debasic (i.e., they can lack nucleic acid bases). Creotides may also lack nucleic acid bases and sugars (i.e., C3 spacers) be).
[0356] The nucleotides in a polynucleotide may be bonded to each other in any manner. Nucleotides are usually linked by their sugar and phosphate groups, similar to nucleic acids. Nucleotides are formed by their nucleic acid bases, as in the case of pyrimidine dimers. It may also be connected.
[0357] Polynucleotides can be single-stranded or double-stranded. At least a portion of the polynucleotide is preferably double-stranded.
[0358] Polynucleotides are nucleic acids, such as deoxyribonucleic acid (DNA) or ribonucleic acid (R It can be NA). Polynucleotides hybridize to a single DNA strand. It can contain one RNA chain. Polynucleotides are peptide nucleic acids (PNAs). Glycerol nucleic acid (GNA), threose nucleic acid (TNA), locked nucleic acid (LNA), or other synthetic polymers having nucleotide side chains, etc., which are not known in the art. It may be any synthetic nucleic acid. The PNA backbone is a repeating N chain linked by peptide bonds. It is composed of -(2-aminoethyl)-glycine units. The GNA main chain is phosphodiester It is composed of repeating glycol units linked by tel bonds. The TNA backbone is phosphate It is composed of repeated threose sugars linked together by hodiester bonds. As discussed above, it has an extra bridge connecting the 2' oxygen and 4' carbon of the ribose moiety. It is formed from ribonucleotides. Cross-linked nucleic acids (BNAs) are made from modified RNA nucleotides. Yes. BNA is sometimes called bound or isolated RNA. BNA monomers are "solid A 5-membered, 6-membered, or even 7-membered bridge having a defined C3'-endo sugar puckering. It can include a bridge structure. The bridge is synthetically incorporated at the 2',4' positions of the ribose. It produces 2',4'-BNA monomers.
[0359] Polynucleotides are most preferably ribonucleic acid (RNA) or deoxyribonucleic acid (D NA)
[0360] Polynucleotides can be of any length. For example, polynucleotides are , length at least 10, at least 50, at least 100, at least 150, less 200 each, at least 250, at least 300, at least 400 or at least It can be 500 nucleotides or a pair of nucleotides. Polynucleotides are long 1000 nucleotides or more nucleotide pairs, with a length of 5000 nucleotides or This refers to a number of nucleotide pairs or more, or a length of 100,000 nucleotides or more than a nucleotide pair. It is possible to be superior.
[0361] Any number of polynucleotides can be investigated. For example, the method of the present invention is 2 , 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 50, 100 or more polynucleotides This concerns the characterization of a polynucleotide. When characterizing two or more polynucleotides, These may be different polynucleotides, and two examples of the same polynucleotide are shown. Sometimes it happens.
[0362] Polynucleotides can be naturally occurring or artificially created. For example, the sequence of the manufactured oligonucleotide can be verified using this method. It is possible. The method is usually performed in vitro.
[0363] sample Polynucleotides are typically present in any suitable sample. The present invention is typically, Using samples that are known to contain or suspected to contain polynucleotides This is done by using a sample and determining what is known to be present in that sample. It may also be done to confirm the identity of existing or expected polynucleotides.
[0364] The sample may be a biological sample. The present invention relates to samples obtained from any organism or microorganism. Alternatively, it can be performed in vitro using the extracted sample. The organisms are usually archaeal, prokaryotic, or eukaryotic, and typically belong to one of the five kingdoms, namely the plant kingdom. It belongs to one of the following kingdoms: Animalia, Fungi, Prokaryotes, and Protists. This invention relates to all We This can be performed in vitro using samples obtained from or extracted from Russ. The sample is preferably a fluid sample. The sample usually contains the patient's body fluids. The sample may include urine, etc. It may be amniotic fluid, saliva, mucus, or amniotic fluid, but preferably blood, plasma, or serum. ru.
[0365] The samples are usually of human origin, but alternatively, they may be from other mammals, for example. , from commercially raised animals, such as horses, cows, sheep, fish, chickens or pigs. It is fine to have a pet, or alternatively, a cat or a dog. The samples are of plant origin, for example, cash crops, for example, cereals, legumes, fruits or Vegetables, for example, wheat, barley, oats, canola, corn, soybeans, Rhubarb, rhubarb, banana, apple, tomato, potato, grape, tobacco, bean, lentil Samples obtained from cinnamon, sugarcane, cacao tree, or cotton may also be used.
[0366] The sample may be a non-biological sample. The non-biological sample is preferably a fluid sample. Examples of body samples include surgical fluids, water, such as drinking water, seawater or river water, and laboratory samples. Examples include reagents for testing.
[0367] The sample is usually prepared before use in the present invention, for example, by centrifugation, or if desired. By passing through a membrane that filters out non-conforming molecules or cells, such as red blood cells, The sample is processed. The sample may be measured immediately after collection. The sample is preferred before assay. It is also common for them to be stored at -70°C.
[0368] Characterization The method includes a step of measuring two, three, four, or five or more features of a polynucleotide. This may occur. One or more features are preferably (i) the length of the polynucleotide, ( ii) identity of the polynucleotide, (iii) sequence of the polynucleotide, (iv) polynuclear Selection is based on the secondary structure of the creotide and whether or not the (v) polynucleotide is modified. It is done. {i}, {ii}, {iii}, {iv}, {v}, {i,ii}, {i,ii i}, {i,iv}, {i,v}, {ii,iii}, {ii,iv}, {ii,v}, {iii,iv}, {iii,v}, {iv,v}, {i,ii,iii}, {i,ii ,iv}, {i,ii,v}, {i,iii,iv}, {i,iii,v}, {i,iv ,v}, {ii,iii.iv}, {ii,iii,v}, {ii,iv,v}, {ii i,iv,v},{i,ii,iii,iv},{i,ii,iii,v},{i,ii ,iv,v}, {i,iii,iv,v}, {ii,iii,iv,v} or {i,i In accordance with the present invention, any combination of (i) to (v) such as {i, iii, iv, v} is measured It may be determined that different combinations of (i) to (v) that include any of the combinations listed above are It is possible to measure the first polynucleotide in comparison to the second polynucleotide. Cut.
[0369] (i) The length of the polynucleotide is, for example, between the polynucleotide and the pore. To determine the number of interactions, or the duration of the interaction between polynucleotides and pores. Therefore, it can be measured.
[0370] (ii) The identity of polynucleotides can be measured in a number of ways. The identity of polynucleotides is measured along with the measurement of the polynucleotide sequence. In some cases, the measurement may be performed without measuring the polynucleotide sequence. The measurement described is straightforward, involving sequencing polynucleotides and thereby identifying them. The measurements described later can be performed in several ways. For example, polynucleotides The presence of a specific motif within the polynucleotide can be measured (without measuring the remaining polynucleotide sequence). i. Or, by measuring specific electrical and / or optical signals in that method In some cases, polynucleotides can be identified as originating from a specific source.
[0371] (iii) Determine the polynucleotide sequence as previously described. This is possible. Suitable sequencing methods, especially those using electrical measurements, are available at Stoddart. D et al., Proc Natl Acad Sci, 12;106(19):7702-7, Lieberman KR et al., J Am Chem Soc. 2010;132(50):17961-72, and International Application Publication No. 2000 / 28312 PAN It is written on the fret.
[0372] (iv) The secondary structure can be measured in various ways. For example, If the method involves electrical measurement, it uses changes in residence time or changes in current flowing through the pore. The secondary structure may be measured using this method. This will allow the single-stranded polynucleotide region and the double-stranded polynucleotide region to be determined. This makes it possible to distinguish between polynucleotide regions in a chain.
[0373] (v) can be used to measure the presence or absence of any modification. Polynucleotides are damaged by methylation, oxidation, or other means, one or Modified with multiple proteins or with one or more labels, tags, or spacers. Preferably includes determining whether or not it is present. Specific modifications involve specific interactions with the pore. This leads to the measurement of such specific interactions using the method described below. It is possible to do this. For example, methylcytosine can be used in its interaction with each nucleotide, and the pores can be It can be distinguished from cytosine based on the current that passes through it.
[0374] The target polynucleotide is brought into contact with a CsgG pore or a variant thereof, for example, the pore of the present invention. Pores are usually located within the membrane. Suitable membranes are discussed below. The method can be performed using any equipment suitable for investigating membrane / pore systems. The method can be carried out using any device suitable for pore sensing. For example, The apparatus includes a chamber containing an aqueous solution and a partition wall that divides the chamber into two sections. The barrier wall typically has openings into which a pore-containing membrane is formed. Alternatively, the barrier wall , forming a film in which pores exist.
[0375] International Application No. PCT / GB08 / 000562 Specification (International Publication 2008 / 10212) The method may be carried out using the apparatus described in Pamphlet No. 0.
[0376] Various different types of measurements can be performed. This measurement is not limited to electrical measurements. This includes measurement and optical measurement. Possible electrical measurements include current measurement and impedance measurement. Tunneling effect measurement (Ivanov AP et al., Nano Lett. 2011 Jan 12; 11(1):279-85), Examples include FET measurement (brochure for International Publication Application No. 2005 / 124888). Optical measurements may be used in combination with electrical measurements (Soni GV et al., Rev Sci Instrum. 201). 0 Jan; 81(1):014301). The measurement is a transmembrane current measurement, for example, of ions flowing through a pore. It can also be a measurement of electric current.
[0377] Stoddart D et al., Proc Natl Acad Sci, 12; 106(19):7702-7, Lieberman KR et al. , J Am Chem Soc. 2010;132(50):17961-72, and International Application Publication No. 2000 / 283 Using standard single-channel recording equipment as described in Pamphlet No. 12, Aerodynamic measurements may be performed. Alternatively, see, for example, International Application Publication No. 2009 / 077734. As described in the brochure for International Application Publication No. 2011 / 067559. Electrical measurements may be performed using a multi-channel system such as the one described.
[0378] The method is preferably carried out using a potential applied through a film. The applied potential is The applied potential may be a voltage potential. Alternatively, the applied potential may be a chemical potential. An example of this method is the use of a salt gradient across a membrane, such as an amphiphilic layer. The salt gradient is Hold This is disclosed in en et al., J Am Chem Soc. 2007 Jul 11; 129(27):8650-5. This is achieved by using the current that passes through the pore as polynucleotides move toward the pore. This involves estimating or determining the sequence of polynucleotides. This is called strand sequencing.
[0379] The method involves measuring the current passing through the pore as polynucleotides move through it. The method may include a step of applying a potential. Therefore, the apparatus used in the method may apply a potential. An electrical circuit capable of measuring electrical signals across a film and pore. This may also include electrical circuits. The method is done using patch clamps or voltage fixing. This can also occur. The method preferably involves the use of a fixed voltage.
[0380] The method of the present invention involves the current passing through the pore when polynucleotides move relative to the pore. This may include a measurement step. It is suitable for measuring ion currents through transmembrane protein pores. Suitable conditions are known in the art and are disclosed in the examples. The method is typically This is done using a voltage applied through the film and pore. The voltage used is typically +5 V to -5V, for example, +4V to -4V, +3V to -3V, or +2V to -2V. The voltages used are typically -600mV to +600mV, or -400mV to +40 It is 0mV. The voltages used are preferably -400mV, -300mV, and -200mV. Select from mV, -150mV, -100mV, -50mV, -20mV and 0mV. The lower limit is +10mV, +20mV, +50mV, +100mV, +150mV, +200 The range has an upper limit that is independently selected from mV, +300mV, and +400mV. The voltage used is more preferably in the range of 100mV to 240mV, and most preferably The range is approximately 120mV to 220mV. By using an increased applied potential, the pore This makes it possible to increase the recognition between different nucleotides.
[0381] The method usually involves metal salts, such as alkali metal salts, halide salts, such as chloride salts. This is carried out in the presence of any charge carrier, such as alkali metal chloride salts. For example, an ionic liquid or organic salt, such as tetramethylammonium chloride or trimethylammonium chloride. Phenylphenylammonium, phenyltrimethylammonium chloride, or ethyl chloride -3-methylimidazolium can be cited. In the exemplary apparatus discussed above, the salt is It is present in the aqueous solution inside the chamber. Potassium chloride (KCl), sodium chloride (NaC) l) Cesium chloride (CsCl), or potassium ferrocyanide and potassium ferricyanide A mixture of potassium ferrocyanide is usually used. A mixture of potassium ferricyanide is preferred. The charge carrier is asymmetric with respect to the membrane. For example, the type and / or concentration of charge carriers may differ on both sides of the membrane.
[0382] The salt concentration may be that at saturation. The salt concentration can be less than 3M, but usually... is 0.1~2.5M, 0.3~1.9M, 0.5~1.8M, 0.7~1.7M, 0.9 The concentration is ~1.6M, or 1M~1.4M. The salt concentration is preferably 150mM~1M. Yes. The method preferably involves at least 0.3M, for example, at least 0.4M, less Both 0.5M, at least 0.6M, at least 0.8M, at least 1.0M, less Salt of 1.5M, at least 2.0M, at least 2.5M, or at least 3.0M This is done using concentration. High salt concentrations result in a high signal-to-noise ratio and normal current. This makes it possible to identify currents that indicate the presence of nucleotides against a background of fluctuations. do.
[0383] The method is usually carried out in the presence of a buffer. In the exemplary apparatus discussed above, the buffer is It is present in the aqueous solution within the chamber. Any buffer solution may be used in the method of the present invention. Typically, the buffer is a phosphate buffer. Other suitable buffers include HEPES and Tr It is an is-HCl buffer. The method is usually 4.0-12.0, 4.5-10.0, 5 0~9.0.5.5~8.8, 6.0~8.7, or 7.0~8.8, or 7.5 The procedure is carried out at a pH of ~8.5. The pH used is preferably around 7.5.
[0384] The methods are: 0°C to 100°C, 15°C to 95°C, 16°C to 90°C, 17°C to 85°C, and 18°C. It may be carried out at ~80°C, 19°C to 70°C, or 20°C to 60°C. The method is usually The procedure is carried out at room temperature. The method may, in some cases, be performed at a temperature that supports enzyme function, for example, around 37°C. It will continue.
[0385] Polynucleotide-binding proteins The chain characterization method involves using polynucleotides as polynucleotide-binding proteins and proteins. The quality is in contact with the pore of polynucleotides, for example, to control their movement through the pore. Preferably includes the step of pouring.
[0386] More preferably, the method is to (a) a polynucleotide into a CsgG pore or a variant thereof For example, the present invention involves a pore and a polynucleotide-binding protein, and the protein is polynucleotide-binding A step of bringing a rheotide into contact with a pore, for example, to control its movement through the pore. , and (b) once or more times when the polynucleotide is moving relative to the pore The measurement is performed, and the measurement shows one or more characteristics of polynucleotides, thereby indicating The process includes a step of characterizing the renucleotide.
[0387] More preferably, the method is to (a) a polynucleotide into a CsgG pore or a variant thereof For example, the present invention involves a pore and a polynucleotide-binding protein, and the protein is polynucleotide-binding A step of bringing a rheotide into contact with a pore, for example, to control its movement through the pore. , and (b) the current passing through the pore when the polynucleotide is moving relative to the pore The measurement indicates that the current exhibits one or more characteristics of polynucleotides, thereby indicating the polynucleotide This includes steps that characterize the leotide.
[0388] Polynucleotide-binding proteins bind to polynucleotides and pass through pores. It can be any protein that can control its movement. In the field, it is easy to determine whether or not a protein binds to a polynucleotide. Proteins typically interact with polynucleotides, and at least one of the polynucleotides It modifies one of its properties. Proteins cleave polynucleotides into individual nucleotides. By using shorter nucleotide chains, such as dinucleotides or trinucleotides... Proteins can modify polynucleotides at specific locations. By orienting or moving it to a specific position, that is, by controlling that movement, the polynucle It can also modify the word "Ochido".
[0389] The polynucleotide-binding protein is preferably a polynucleotide-handling enzyme. It is derived from (polynucleotide handling enzyme). Polynucleotide handling enzyme , interacts with polynucleotides to modify at least one property of the polynucleotide. It is a polypeptide that can do this. The enzyme cleaves polynucleotides into individual nucleos To make it a dinucleotide or shorter nucleotide chain, for example, a dinucleotide or a trinucleotide. Therefore, it can modify polynucleotides. The enzyme modifies polynucleotides at specific locations. Polynucleotides can also be modified by orienting or moving them to a specific position. Nucleotide handling enzymes can bind to polynucleotides and create pores. If the movement of the enzyme can be controlled, then it does not need to exhibit enzyme activity. For example, Enzymes can be modified to remove their enzymatic activity, or they can be modified to act as an enzyme. It may also be used under conditions that prevent it from being used. Such conditions will be discussed in more detail below. ru.
[0390] Polynucleotide handling enzymes are preferably derived from nucleolytic enzymes. The polynucleotide handling enzymes used in the constructs are more preferably enzyme-classified ( EC) groups 3.1.11, 3.1.13, 3.1.14, 3.1.15, 3.1.16, 3 1.21, 3.1.22, 3.1.25, 3.1.26, 3.1.27, 3.1.30 The enzyme is derived from any of the members of 3.1.31. The enzyme is from International Application PCT / GB Specification No. 10 / 000133 (as International Publication No. 2010 / 086603 pamphlet) Any of the publicly disclosed information may be used.
[0391] Preferred enzymes include polymerases, exonucleases, helicases, and topoisomers. Enzymes, such as gyrase, are suitable. Sonuclease (SEQ ID NO: 399), exonuclease II from Escherichia coli (E. coli) Enzyme I (SEQ ID NO: 401), RecJ (SEQ ID NO: 401) from hyperthermophilic bacteria (T. thermophilus) 403), and bacteriophage lambda exonuclease (SEQ ID NO: 405), T atD exonuclease and its variants are examples, but are not limited to them. It is not done. The sequence shown in sequence number 403 or its variants consist of three subunits. The knits interact to form trimer exonucleases. These exonucleases Polymerase can also be used in the exonuclease method of the present invention. Polymerase PyroPhage(registered trademark) 3173 DNA polymerase (Lucigen(registered trademark) (Registered trademark) Commercially available from Corporation, SD polymerase (Bioro The enzyme may be one of those commercially available from n(registered trademark), or a variant thereof. Preferably, Phi29 DNA polymerase (SEQ ID NO: 397) or a variant thereof The topoisomerase is preferably from subgroup (EC) 5.99.1.2 and They are a member of one of the categories 5.99.1.3.
[0392] The enzyme is most preferably a helicase, for example, Hel308 Mbu (SEQ ID NO: 40 6) Hel308 Csy (Sequence ID 407), Hel308 Tga (Sequence ID 40 8) Hel308 Mhu (SEQ ID NO: 409), Tral Eco (SEQ ID NO: 410) , derived from XPD Mbu (SEQ ID NO: 411) or its variants. Any helical Helicase may be used in this invention. Helicases include Hel308 helicase and RecD helicase. Cases, for example, Tral helicase or TrwC helicase, XPD helicase, or It may be a Dda helicase, or may be derived from it. Helicase This is the specification of international application PCT / GB2012 / 052579 (international publication 2013 / 0 (Published as pamphlet No. 57495), and also as PCT / GB2012 / 053274. Detailed document (published as international publication no. 2013 / 098562), PCT / G Specification No. B2012 / 053273 (International Publication No. 2013 / 098561 Brochure) (Published as) and the same PCT / GB2013 / 051925 specification (International Publication No. 2014 (Published as pamphlet No. / 013260), also PCT / GB2013 / 051924 Specification No. (published as International Publication No. 2014 / 013259), and PCT No. Specification No. GB2013 / 051928 (Pamphlet No. International Publication No. 2014 / 013262) (Published as a patent) and disclosed in PCT / GB2014 / 052736 It may be a helicase, a modified helicase, or a helicase construct.
[0393] The helicase preferably has the sequence represented by sequence number 413 (Trwc Cba). The variant is the sequence indicated by sequence number 406 (Hel308 Mbu) or That variant, or the sequence or variant indicated by sequence number 412 (Dda). It includes the variant in one of the points discussed below regarding transmembrane pores, which is the native sequence. This may differ. Preferred variants of Sequence ID No. 412 are (a) E94C and A 360C or (b)E94C, A360C, C109A and C136A, then in the case This includes (ΔM1)G1G2 (i.e., the deletion of M1, followed by the addition of G1 and G2). nothing.
[0394] Any number of helicases can be used in accordance with the present invention. For example, 1, 2, 3 Helicases of 4, 5, 6, 7, 8, 9, and 10 or more may also be used. Morphologically, different numbers of helicases may be used.
[0395] The method of the present invention preferably involves contacting a polynucleotide with two or more helicases. Includes a step. Two or more helicases are usually the same helicase. Licase can also be a different type of helicase.
[0396] The two or more helicases may be any combination of the helicases mentioned above. Two or more helicases may also be two or more Dda helicases. Licase consists of one or more Dda helicases and one or more TrwC helicases. This can also be the case. Two or more helicases are different variants of the same helicase. Sometimes.
[0397] Two or more helicases are preferably bound to each other. More preferably, they are covalently bonded to each other. Using any order and any method Helicase may be bonded to it. A helicase structure preferred for use in the present invention is, International application PCT / GB2013 / 051925 specification (International publication 2014 / 0132 (Published as Pamphlet No. 60), and the same PCT / GB2013 / 051924 specification ( (Published as International Publication No. 2014 / 013259, PCT / GB20) Specification No. 13 / 051928 (as International Publication No. 2014 / 013262) (Published) and as described in PCT / GB2014 / 052736 specification.
[0398] Sequence numbers 397, 399, 401, 403, 405, 406, 407, 408, 409 Variants 410, 411, 412 or 413 are sequence numbers 397, 399, 40 1, 403, 405, 406, 407, 408, 409, 410, 411, 412 or An enzyme having a different amino acid sequence from that of 413, and possessing polynucleotide binding ability. It is an enzyme that retains [something]. This ability can be measured using any method known in the art. It can be determined. For example, the variant can be brought into contact with a polynucleotide. It measures its ability to bind to polynucleotides and to move along polynucleotides. It is possible. Variants are modifications that facilitate the binding of polynucleotides, as well as / also It may contain modifications that enhance its activity at high salt concentrations and / or at room temperature. Barrier The nucleotide binds to polynucleotides (i.e., retains polynucleotide binding activity). ) does not function as a helicase (i.e., all necessary components for promoting movement, For example, ATP and Mg 2+ , when present, does not move along the polynucleotide. It may also be modified in such a way. Such modifications are well known in the art. Example For example, Mg in helicase 2+ Modification of the binding domain typically functions as a helicase. This results in variants that do not exist. These types of variants act as molecular brakes. (See below)
[0399] Sequence numbers 397, 399, 401, 403, 405, 406, 407, 408, 409 The variant is, across the entire length of the amino acid sequence of 410, 411, 412, or 413, Preferably, the sequence is at least 50% homologous based on amino acid identity. More preferably, the variant polypeptide is based on amino acid identity and the entire sequence. Sequence numbers 397, 399, 401, 403, 405, 406, 407, 40 8, 409, 410, 411, 412 or 413 amino acid sequences and at least 55%, At least 60%, at least 65%, at least 70%, at least 75%, and 80%, at least 85%, at least 90%, and more preferably at least 95% They can be 97% or 99% homologous. 200 or more consecutive numbers, for example 230, 250, 270, 280, 300, 400, 500, 600, 700, 800, 900 or 10 Over a stretch of 00 or more amino acids, at least 80%, for example, at least 85 There may also be %, 90%, or 95% amino acid identity ("rigorous homology"). Homology is , as discussed above. The variant is as discussed above with respect to Sequence ID No. 390. The sequence may differ from the wild-type sequence at one of the dots. The enzyme may also be covalently bonded to the pore. The enzyme may be covalently bonded to the pore by any method.
[0400] A preferred molecular brake is TrwC Cba-Q594A (a variant of Q594A). This is column number 413). This variant does not function as a helicase (i.e., po It binds to the renucleotide, but also to all the necessary components to facilitate its movement, such as ATP. and Mg 2+ (It does not move along the polynucleotide when it is present).
[0401] In chain sequencing, polynucleotides react to or repel the applied potential. It is translocated through A. It acts progressively or progressively on double-stranded polynucleotides. The exonuclease is then injected into the cis side of the pore under the applied potential, in order to deliver the remaining single strand. It can be used in that way, or it can be used on the transformer side under reverse potential. Similarly, helicases that unwind double-stranded DNA can be used in the same way. Lazes can also be used. For sequencing applications, chain transfer can be performed against the applied potential. It may require a row, but under reversed potential or in the absence of potential, DNA is enzyme It should first be "captured" by this. Then, when the potential switches after bonding, the chain It passes through the pore from cis to transform and is held by the current in the extended higher-order structure. Single-stranded DNA exonucleases or single-stranded DNA-dependent polymerases are molecular molecules. It acts as a trap, repelling the single strand that has just been transferred through the pore to the applied potential. It can be pulled back from lance to sys in a controlled, gradual manner.
[0402] Any helicase can be used in the method. The helicase is applied to the pore. It can operate in two modes. Firstly, the method is preferably by helicase So the polynucleotide moves through the pore in response to the field that is created as a result of the applied voltage. This is done using helicase. In this mode, the 5' end of the polynucleotide is First, the polynucleotides are captured within the pore, and then, by helicase, they are ultimately transfused into the membrane. Move into the pore so that it is passed through the pore along with the field until it moves to the side. Alternatively, method Preferably, polynucleotides are produced by helicase as a result of the applied voltage. It is done by bouncing off the trap and moving through the pore. In this mode, polynucleo The 3' end of the nucleotide is first trapped within the pore, and the polynucleotide is then processed by helicase. It is as if being pulled out of the pore by the applied field, repelling it until it is pushed back to the cis side of the membrane. They travel through the pore.
[0403] The method can also be performed in the reverse direction. The 3' end of the polynucleotide enters the pore first. Polynucleotides can be captured and, by helicase, ultimately transcend the membrane. Until it moves to the side, it may also move into the pore, passing through it as the situation changes.
[0404] If helicase lacks the components necessary to facilitate movement, or if helicase If the ze is modified to hinder or prevent its movement, the helicaze is polynu The polynucleotide binds to the creotide, and the applied site pulls the polynucleotide into the pore. This can act as a brake, slowing down the movement of polynucleotides. In the inactive mode, are polynucleotides captured from 3' downwards or from 5' downwards? Whether it is captured or not is not the issue; the issue is whether the enzyme acts as a brake and the pollen This is the applied field that draws the creotide towards the transformer side within the pore. (Inactive mode) In this case, the control of polynucleotide movement by helicase is gradual, sliding, and braking. This can be explained in numerous ways, including the following: helicar lacking helicase activity Zevaliant can also be used in this way.
[0405] In what order do you contact polynucleotides with polynucleotide-binding proteins and pores? It may be done. Polynucleotides are converted into polynucleotide-binding proteins such as helicase. When in contact with pores, polynucleotides first form complexes with proteins. It is preferable to do so. When a voltage is applied through the pore, polynucleotides / proteins The complex forms a complex with the pore and controls the movement of polynucleotides through the pore.
[0406] Every step in the method using polynucleotide-binding proteins is usually free nucleotides or free nucleotide analogs and polynucleotide-binding proteins The process is carried out in the presence of enzyme cofactors that enhance the action. Free nucleotides are the individual ones discussed above. It can be one or more of the nucleotides. Free nucleotides include A Denosine monophosphate (AMP), adenosine diphosphate (ADP), adenosine triphosphate (A TP), guanosine monophosphate (GMP), guanosine diphosphate (GDP), guanosine triphosphate Phosphate (GTP), thymidine monophosphate (TMP), thymidine diphosphate (TDP), thymid Uridine triphosphate (TTP), uridine monophosphate (UMP), uridine diphosphate (UDP), Lysine triphosphate (UTP), cytidine monophosphate (CMP), cytidine diphosphate (CDP) Cytidine triphosphate (CTP), cyclic adenosine monophosphate (cAMP), cyclic guanosine Deoxyadenosine monophosphate (cGMP), deoxyadenosine monophosphate (dAMP), deoxyadenosine Diphosphate (dADP), deoxyadenosine triphosphate (dATP), deoxyguanosine Deoxyguanosine monophosphate (dGMP), deoxyguanosine diphosphate (dGDP), deoxyguanosine Triphosphate (dGTP), deoxythymidine monophosphate (dTMP), deoxythymidine diphosphate Deoxyuridine triphosphate (dTDP), deoxythymidine triphosphate (dTTP), deoxyuridine monophosphate (dUMP), deoxyuridine diphosphate (dUDP), deoxyuridine triphosphate (d UTP), deoxycytidine monophosphate (dCMP), deoxycytidine diphosphate (dCD) Examples include P) and deoxycytidine triphosphate (dCTP), but are not limited to these. The free nucleotides are preferably AMP, TMP, GMP, CMP, UMP, dA Selected from MP, dTMP, dGMP, or dCMP. Free nucleotides are preferred. The enzyme cofactor is adenosine triphosphate (ATP). Enzyme cofactors enable the construct to function. This is a factor that causes [something]. The enzyme cofactor is preferably a divalent metal cation. The n is preferably Mg 2+ Mn 2+ Ca2+ or Co 2+ The enzyme cofactor is , most preferably Mg 2+ That is the case.
[0407] Helicases and molecular brakes In a preferred embodiment, the method is (a) one or more helicases and one or more molecular brakes are coupled Steps to prepare polynucleotides, (b) Contact the polynucleotide with the CsgG pore or a variant thereof, for example, the pore of the present invention. This causes one or more helicases and molecular brakes to be brought together, and both against the pore For example, to control the movement of polynucleotides through a pore, an electrical potential is generated across the pore. The step of applying, (c) Perform one or more measurements while the polynucleotide is moving against the pore. The measurement indicates that the polynucleotide exhibits one or more characteristics of a polynucleotide, thereby indicating a polynucleotide Steps that characterize Chido Includes.
[0408] This type of method is described in the brochure for international application PCT / GB2014 / 052737. This is discussed in detail in [the document].
[0409] One or more helicases may be any of those discussed above. These multiple molecular brakes bind to polynucleotides and pass through the pores of those polynucleotides. It can be any compound or molecule that slows down the movement. One or more The molecular brake preferably contains one or more compounds that bind to the polynucleotide. The compound is preferably one or more macrocyclic molecules. Macrocyclic molecules include cyclodextrins, calixarenes, cyclic peptides, and crystals. Examples include ethers, cucurbituryls, pyralarenes, their derivatives, or combinations thereof. These may be used, but are not limited to them. Cyclodextrins or their derivatives are used by Eliseev, Disclosed in AV, and Schneider, HJ. (1994) J. Am. Chem. Soc. 116, 6081-6088. It may be any of the following. This active substance is more preferably heptakis-6. -amino-β-cyclodextrin (am7-βCD), 6-monodeoxy-6-mono Mino-β-cyclodextrin (am1-βCD) or heptakis-(6-deoxy- It is 6-guanidino)-cyclodextrin (gu7-βCD).
[0410] One or more molecular brakes are preferably one or more single-chain binding proteins. It is a quality (SSB). One or more molecular brakes are carboxy compounds that do not have a net negative charge. (ii) a single-chain binding protein (SSB) containing a C-terminal region, or (ii) its C-terminus Modification S includes one or more modifications in the edge region that reduce the net negative charge of the C-terminal region. It is SB. One or more molecular brakes are most preferably international application PCT / G Specification No. B2013 / 051924 (International Publication No. 2014 / 013259 Brochure) It is one of the SSBs disclosed (publicly available as).
[0411] One or more molecular brakes are preferably one or more polynucleotide bonds. It is a compound protein. Polynucleotide-binding proteins bind to polynucleotides. It is any protein that can control the process and its movement through the pore. This can be done. In this technology, it is possible to determine whether or not a protein binds to a polynucleotide. This is easy. Proteins usually interact with polynucleotides, and polynucleotides Modify at least one property of the protein. Proteins cleave polynucleotides into individual Each nucleotide or shorter nucleotide chain, for example, di or trinucleotides, This can modify the polynucleotide. To orient or move Otid to a specific position, that is, to control its movement. Therefore, it can sometimes be used as a modifier.
[0412] The polynucleotide-binding protein is preferably a polynucleotide-handling enzyme. It originates from one or more molecular brakes, which are polynucleotide handlers discussed above. It may be derived from any of the enzymes. Phi2 acts as a molecular brake. A modified version of polymerase 9 (SEQ ID NO: 396) is licensed under the U.S. Patent No. 5,576,204. Disclosed in the specification. One or more molecular brakes are preferably helicases. To originate from.
[0413] Any number of molecular brakes derived from helicase can be used. For example, 1 2, 3, 4, 5, 6, 7, 8, 9, 10 or more helicases are used as molecular brakes. This can happen. When two or more helicases are used as molecular brakes, two of them More than one helicase is usually the same helicase. More than one helicase is different. It can also be helicase.
[0414] The two or more helicases may be any combination of the helicases mentioned above. Two or more helicases may be two or more Dda helicases. Licase consists of one or more Dda helicases and one or more TrwC helicases. This can also be the case. Two or more helicases are different variants of the same helicase. Sometimes.
[0415] Two or more helicases are preferably bound to each other. More preferably, they are covalently bonded to each other. In any order, using any method, Licase may be attached. One or more molecular brakes derived from helicase are Polynucleotides can dissociate from helicase in at least one conformational state. Preferably modified to reduce the size of the opening within the renucleotide binding domain. This is disclosed in International Publication No. 2014 / 013260.
[0416] A helicase construct preferred for use in the present invention is International Application No. PCT / GB2013 / 0 Specification No. 51925 (published as International Publication No. 2014 / 013260), PCT / GB2013 / 051924 specification (International Publication No. 2014 / 013259) (Published as a pamphlet), and the same PCT / GB2013 / 051928 specification (International (Published as brochure No. 2014 / 013262) and also as PCT / GB20 It is stated in the specification document No. 14 / 052736.
[0417] When one or more helicases are used in active mode (i.e., one or multiple) All the necessary components to facilitate the movement of several helicases, such as ATP and Mg 2+ When provided, one or more molecular brakes are preferably (a) inert molecular Used in the absence of necessary components to facilitate movement, (or unable to move actively), (b) one or more of the molecular brakes It is used in an active mode that moves in the opposite direction to multiple helicases, or (c) its One or more molecular brakes are in the same direction as one or more helicases, and not one Alternatively, it is used in an active mode that moves more slowly than multiple helicases.
[0418] When one or more helicases are used in inactive mode (i.e., one or All the necessary components to facilitate the movement of multiple helicases, such as ATP and Mg 2 + When it is not present or cannot be actively moved, one or more molecular bridges The key is preferably used in (a) inactive mode (i.e., to facilitate movement) (b) Used in the absence of necessary components, or unable to be actively transported) or one of them Alternatively, multiple molecular brakes may be aligned along the polynucleotide, in the same direction as the polynucleotide. It moves through the pore and is used in active mode.
[0419] One or more helicases and one or more molecular brakes, polynucleotides They can be joined at any position on the do, and as a result they come together, and both are polynucle Controls the movement of rheotide through the pore. One or more helicases and one or more The number of molecular brakes is at least one nucleotide apart, for example, at least 5, At least 10, at least 50, at least 100, at least 500, at least 10 00, at least 5000, at least 10,000, at least 50,000 nucleos They are more than 100cm apart. The method involves attaching a Y-adapter to one end and a hairpin to the other end. When it comes to characterizing double-stranded polynucleotides equipped with a link adapter, One or more helicases are coupled to the Y-adapter, preferably in a hairpin loop. One or more molecular brakes are coupled to the adapter. In this embodiment, one or Multiple molecular brakes preferably bind to polynucleotides but act as helicases. It is one or more helicases that have been modified to not function. The combined one or more helicases are preferably speculative, as will be discussed in more detail below. It is stopped by the ser. One or more molecular pulses are coupled to the hairpin loop adapter. - The key is preferably not stopped by a spacer. One or more helicases and one Alternatively, multiple molecular brakes are preferably one or more helicases that form hairpin holes. They come together when they reach the point. One or more helicases connect the Y adapter to the polynucleate. You can combine it with the Y adapter before combining it with Ochido, or combine the Y adapter with Polynu The molecules may be attached to the creotide and then to the Y adapter. Before attaching the hairpin loop adapter to the polynucleotide, the hairpin loop It may be attached to a push adapter, or a hairpin loop adapter with polynucleotide After connecting it to the hairpin loop adapter, you can also connect it to the hairpin loop adapter.
[0420] One or more helicases and one or more molecular brakes are preferably relative to each other. They do not bind. One or more helicases and one or more molecular brakes are preferable. They are not covalently bonded to each other. One or more helicases and one or more The molecular brake is preferably as described in International Application No. PCT / GB2013 / 051925. (Published as international publication no. 2014 / 013260, PCT / GB) Specification No. 2013 / 051924 (and brochure International Publication No. 2014 / 013259) (Published), and the same PCT / GB2013 / 051928 specification (International Publication No. 2014 / (Published as pamphlet No. 013262) and also PCT / GB2014 / 05273 They are not combined as described in Specification No. 6.
[0421] Spacer One or more helicases, International Application No. PCT / GB2014 / 050175 As discussed in the section on frets, you can also stop it with one or more spacers. i. One or more helicases and one or more sperm disclosed in the international application Any of the stereochemical configurations of the sar may be used in this invention.
[0422] A portion of the polynucleotide enters the pore and follows the field resulting from the applied potential. When moving through the pore, one or more helicases are involved in the process where polynucleotides are located within the pore. Moving through it causes it to move beyond the spacer by that pore. This is (one also (This includes multiple spacers) Polynucleotides move through the pore, one or more This is because helicase remains in the upper part of the pore.
[0423] One or more spacers are preferably parts of a polynucleotide, for example, Spacers are inserted into the polynucleotide sequence. One or more spacers Preferably, one or more blocking molecules hybridized to a polynucleotide, e.g. It is not part of Speed Bump.
[0424] A polynucleotide can have any number of spacers, for example, 1, 2, 3 There may be 4, 5, 6, 7, 8, 9, or 10 or more spacers. Preferably, poly There are two, four, or six spacers within a nucleotide. There may be one or more spacers within the region, for example, a Y adapter and / Alternatively, there may be one or more spacers inside the hairpin loop adapter.
[0425] One or more spacers, even if one or more helicases are in the active mode. Each of them creates an insurmountable energy barrier. One or more spacers - is achieved by removing a base from a nucleotide in a polynucleotide (for example, by removing a base from a nucleotide in a polynucleotide). By reducing the traction force of the gase, or (for example, by using a bulky chemical group) By physically blocking the movement of one or more helicases, one or more It is possible to stop helicase.
[0426] One or more spacers stop one or more helicases in any minute It may include combinations of children or molecules. One or more spacers may be one or Any molecule that prevents multiple helicases from moving along a polynucleotide. Or it may include a combination of molecules. One or more helicases may form a transmembrane pore. And to determine whether it is stopped by one or more spacers in the absence of the applied potential. This is easy. For example, a helicar that moves across a spacer and replaces the complementary strand of DNA. Ze's capabilities can be measured by PAGE.
[0427] One or more spacers typically contain linear molecules such as polymers. Multiple spacers typically have a structure different from that of polynucleotides. For example, If the renucleotide is DNA, then one or more spacers are usually not DNA. In more detail, polynucleotides are deoxyribonucleic acid (DNA) or ribonucleic acid (RN). In case A), one or more spacers are preferably peptide nucleic acids (PNA). Glycerol nucleic acid (GNA), threose nucleic acid (TNA), locked nucleic acid (LNA), or comprising a synthetic polymer having nucleotide side chains. One or more spacers are It may contain one or more nucleotides in the opposite direction to the polynucleotide. For example, If the polynucleotide is oriented from 5' to 3', then one or more spacers , may contain one or more nucleotides in the 3' to 5' direction. D may be any of the options discussed above.
[0428] One or more spacers contain one or more nitroindoles, for example, one or multiple 5-nitroindoles, one or more inosine, one or more Acridine, one or more 2-aminopurines, one or more 2,6-diamidines Nopurine, one or more 5-bromodoxyuridines, one or more inverted 5-bromodoxyuridines Thymidine (inverted dT), one or more inverted dideoxythymidines (ddT), one or multiple dideoxycytidines (ddC), one or more 5-methylcytidines, One or more 5-hydroxymethylcytidines, one or more 2'-O-methyl RNA base, one or more isodeoxycytidines (isodC), one or more A number of isodeoxyguanosine (isodG), one or more iSpC3 groups (i.e.) nucleotides lacking sugars and bases, one or more photocleavable (PC) groups, One or more hexanediol groups, one or more spacer 9 (iSp9) Base, one or more Spacer 18 (iSp18) bases, polymer or one or Preferably contains multiple thiol bonds. One or more spacers are located between these groups. This may include any combination. Many of these groups are IDT(registered trademark) (Integ It is commercially available from rated DNA Technologies (registered trademark).
[0429] One or more spacers may contain any number of these groups. For example , 2-aminopurine, 2,6-diaminopurine, 5-bromodoxyuridine, inverted dT, ddT, ddC, 5-methylcytidine, 5-hydroxymethylcytidine, 2'-O -MethylRNA base, iso dC, iso dG, iSpC3 group, PC group, hexanediol group And for thiol bonds, the tertiary spacers are preferably 2, 3, Includes 4, 5, 6, 7, 8, 9, 10, 11, and 12 or more. One or more spacers are Preferably, it contains 2, 3, 4, 5, 6, 7, or 8 or more iSp9 groups. The pacer preferably contains 2, 3, 4, 5 or 6 or more iSp18 groups. The new spacer consists of four iSp18 units.
[0430] The polymer is preferably a polypeptide or polyethylene glycol (PEG). The polypeptide is preferably 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12 The above amino acids are included. PEG is preferably 2, 3, 4, 5, 6, 7, 8, 9, 10 , containing 11 or 12 or more monomer units.
[0431] One or more spacers are preferably one or more debaset nucleotides ( That is, nucleotides lacking nucleic acid bases, for example, 2, 3, 4, 5, 6, 7, 8 It contains 9, 10, 11, 12 or more debaseated nucleotides. Nucleic acid bases are debaseated nucleos In tide, it may be substituted with -H(idSp) or -OH. This removes nucleic acid bases from multiple adjacent nucleotides, thereby creating a debase spacer. It can be inserted into a polynucleotide. For example, a polynucleotide can be inserted into a 3-methyladenonucleotide. Nin, 7-methylguanine, 1,N6-ethenoadenine, inosine, or hypoxanthine It may be modified to include, and the nucleic acid base is human alkyladenine DNA glycosyl These nucleotides can be removed using an enzyme (hAAG). Polynucleotides can be modified to include uracil, and nucleic acid bases are uracil It can also be removed by sil-DNA glycosylase (UDG). In one embodiment, one Alternatively, the spacers may not contain any debasic nucleotides.
[0432] One or more helicases are separated by their respective linear molecular spacers (i.e., they It may be stopped before the others, or stopped at each linear molecular spacer. There are also cases where linear molecular spacers are used, and polynucleotides have one or multiple spacers. A polynucleotide adjacent to the end of each spacer will be the one that the helicase will move across. A double-stranded region of the cyd is preferably provided. This double-stranded region is usually adjacent to the space It helps to stop one or more helicases in the sar. The presence of a double-stranded region is a method This is particularly preferable when the salt concentration is approximately 100 mM or less. Each double-stranded region is usually, The length is at least 10, for example, at least 12 nucleotides. Used in the present invention When a polynucleotide is single-stranded, the double-stranded region is a shorter polynucleotide. This can be formed by hybridizing to a region adjacent to the spacer. These shorter polynucleotides are usually formed from the same nucleotides as the polynucleotides. However, they can also be formed from different nucleotides. For example, shorter polynucleotides Rheotides can be formed from LNAs.
[0433] When linear molecular spacers are used, the polynucleotide preferably has one or the end of each spacer opposite to the end over which multiple helicases will move. It is equipped with blocking molecules. These blocking molecules have one or more helicases in each space It can help ensure that the stopped state is maintained. The blocking molecule is one Alternatively, if multiple helicases diffuse into a solution, they are retained on the polynucleotide. It can also help in maintaining. Blocking molecules physically block one or more helicases. The stopping group can be any of the chemical groups discussed below. The blocking molecule is a polynucleotide. It can also be a double-stranded region of ochide.
[0434] One or more spacers physically stop one or more helicases. Preferably contains one or more chemical groups. Preferably, one or more chemical groups Or multiple pendant chemical groups. One or more chemical groups in a polynucleotide. It may also be bonded to one or more nucleic acid bases. One or more chemical groups are They can also be bound to the polynucleotide backbone. 2, 3, 4, 5, 6, 7, 8, 9, Any number of chemical groups can exist, such as 10, 11, 12 or more. Preferred groups include These include fluorophores, streptavidin and / or biotin, cholesterol, and Chilen blue, dinitrophenol (DNP), digoxigenin and / or anti-digoxigenin Examples include, but are not limited to, shigenin and dibenzylcyclooctin groups.
[0435] Different spacers in a polynucleotide may contain different termination molecules. For example. One spacer may contain one of the linear molecules discussed above, while another spacer may contain , comprising one or more chemical groups that physically terminate one or more helicases There is. The spacer is one of the linear molecules discussed above, as well as one or more he One or more chemical groups that physically stop lycase, for example, one or more debases and It may contain fluorophores.
[0436] Depending on the type of polynucleotide and the conditions under which the method of the present invention is performed, a suitable spec Helicases can be designed. Most helicases bind to DNA and follow the DNA. Since they move, we can stop them using something other than DNA. The molecules discussed above.
[0437] The method of the present invention is preferably performed in the presence of free nucleotides and / or helicase This is carried out in the presence of cofactors. This will be discussed in more detail below. Transmembrane pore and application In the absence of potential, one or more spacers are preferably free nucleotides. Arresting one or more helicases in the presence of and / or helicase cofactors It is possible.
[0438] The method of the present invention involves free nucleotides and (one or more helicases in an active mode) When performed in the presence of helicase cofactors as discussed below, usually, Use one or more longer spacers to insert one or more helicases into polynucleate After ensuring they are completely stopped on the rheotide, they are brought into contact with the transmembrane pore and a potential is applied. Separating nucleotides and (one or more helicases in inactive mode) helicase In the absence of gauze, one or more shorter spacers can be used.
[0439] Salt concentration also affects the ability of one or more spacers to stop one or more helicases. It affects the force. In the absence of transmembrane pores and applied potential, one or more spacers - Preferably, one or more helicases are stopped at a salt concentration of about 100 mM or less. It is possible. The higher the salt concentration used in the method of the present invention, the more generally one is used. Alternatively, multiple spacers can be short, and vice versa.
[0440] Table 3 below shows preferred combinations of features.
[0441] [Table 3]
[0442] The method also involves moving two or more helicases beyond the spacer. In such cases, the length of the spacer is usually increased to eliminate the pore and the absence of the applied potential. The method is to prevent the trailing helicase from pushing the leading helicase beyond the spacer. Regarding moving two or more helicases across one or more spacers In this case, the spacer length discussed above is at least 1.5 times, for example, 2 times, 2.5 times, or 3 times. It can be doubled. For example, if the method uses more than one or more spacers, it can use more than two. Regarding the movement of the helicase above, the spacer length is shown in the third column of Table 3 above. It may be increased by 1.5 times, 2 times, 2.5 times, or 3 times.
[0443] film The pores of the present invention may be present within a membrane. In the method of the present invention, polynucleotides CsgG pores or their variants within the membrane, such as the pores of the present invention, are typically brought into contact with the membrane. A film can be used in accordance with the present invention. Suitable films are well known in the art. The film is preferably an amphiphilic layer. The amphiphilic layer contains both hydrophilic and lipophilic portions. It is a layer formed from amphiphilic molecules, such as phospholipids. Amphiphilic molecules are synthesized It may be a compound, or it may exist naturally. Amphiphilic compounds do not exist naturally. Materials and amphiphilic materials that form monolayers are well known in the art, for example, Contains block copolymer (Gonzalez-Perez et al., Langmuir, 2009, 25, 10447-10450) Block copolymers are polymers formed when two or more monomer subunits polymerize with each other to form a single polymer. It is a polymer material that generates polymer chains. Block copolymers are typically composed of each monomer subunit. Knitted properties are one of the contributing factors. However, block copolymers are individual subunits Block copolymers may possess unique properties not found in polymers formed from them. One of the monomer subunits is hydrophobic (i.e., lipophilic), but the other subunits It can be engineered to be hydrophilic when in an aqueous medium. Block copolymers possess amphiphilic properties and form structures that mimic biological membranes. It is possible. Block copolymers are dib (consisting of two monomer subunits). It can be a lock, but it behaves as an amphiphilic substance, forming more complex configurations. Copolymers can sometimes be constructed from more than two monomer subunits. The membrane may be a triblock, tetrablock, or pentablock copolymer. Preferably, it is a triblock copolymer film.
[0444] Archaeal bipolar tetraether lipids are naturally constructed so that the lipids form a monolayer membrane. These lipids are present in extremophilic bacteria that survive in harsh biological environments, and thermophilic bacteria. It is commonly found in bacteria, halophiles, and acidophiles. Their stability is due to the melting of the final bilayer. It is thought to originate from the combined properties. It has a general motif of hydrophilic-hydrophobic-hydrophilic. By producing triblock polymers that mimic these biological entities, It is easy to construct a mimicking block copolymer material. This material consists of a lipid bilayer and They behave similarly and form monomer films that encompass phase behavior ranging from vesicles to layered membranes. It is possible. The membranes formed from these triblock copolymers can be used in biological lipid membranes. It possesses several superior advantages. Because it synthesizes triblock copolymers, it is precisely this structure Carefully control the construction to form a membrane and interact with pores and other proteins. The precise chain length and properties required for use can be obtained.
[0445] Block copolymers are constructed from subunits that are not classified as lipid copolymers. It is also said that hydrophobic polymers are derived from siloxanes or other non-hydrocarbon monomers. It can also be manufactured. The hydrophilic portion of the block copolymer has low protein binding properties. It may also have low protein binding properties, and this low protein binding property means that it is exposed to raw biological samples. This enables the creation of a film that is highly resistant when exposed to cold. This head group unit is non-classical It may also originate from lipid head groups.
[0446] Triblock copolymer membranes exhibit increased mechanical and environmental stability compared to biological lipid membranes. It also has properties such as a higher operating temperature or pH range. The synthesizability of block copolymers is A platform for customizing polymer-based films for a wide range of applications. To bring about.
[0447] The film is most preferably according to International Application No. PCT / GB2013 / 052766 or the same. It is one of the membranes disclosed in PCT / GB2013 / 052767 pamphlet.
[0448] Amphiphilic molecules are chemically modified to facilitate the coupling of polynucleotides. It may be made into a sensual or functionalized product.
[0449] The amphiphilic layer may be monolayered or bilayered. The amphiphilic layer is usually planar. The amphiphilic layer can also be curved. The amphiphilic layer is supported Sometimes it is.
[0450] Amphiphilic membranes are typically naturally mobile and essentially function as two-dimensional fluids, approximately 10 8 c m·s -1 It acts at the lipid diffusion rate of the pore and coupled polynucleus. This means that rheotides can normally move within amphipathic membranes.
[0451] The membrane may be a lipid bilayer. The lipid bilayer is a model of the cell membrane, and in a certain range It serves as an excellent platform for experimental research. For example, lipid bilayers can be used as a single channel. It can be used for in vitro studies of membrane proteins by recording. Alternatively, lipids By using a bilayer as a biosensor, it is possible to detect the presence of a certain range of substances. The lipid bilayer can be any lipid bilayer. Suitable lipid bilayers include Examples include, but are not limited to, planar lipid bilayers, supporting bilayers, or liposomes. The lipid bilayer is preferably a planar lipid bilayer. A suitable lipid bilayer is described in the International Application. PCT Specification No. GB08 / 000563 (International Publication No. 2008 / 102121 Pamphlet) (Published as a Let), International Application No. PCT / GB08 / 004127 Specification (International Publication No. (Published as brochure No. 2009 / 077734) and international application PCT / GB20 Specification No. 06 / 001057 (as International Publication No. 2006 / 100484 pamphlet) It is publicly disclosed.
[0452] Methods for forming lipid bilayers are known in the art. Lipid bilayers are composed of lipids. The layer is supported at the aqueous solution / air interface, extending beyond both sides of the opening perpendicular to that interface. Montal and Mueller's method (Proc. Natl. Acad. Sci. US) It is generally formed by (A., 1972; 69: 3561-3566). Typically, lipids are formed in an aqueous electrolyte solution. First, the lipids are dissolved in an organic solvent on the surface, and then droplets of the solvent are placed on the surface of the aqueous solution on both sides of the opening. It is added by evaporation across the surface. Once the organic solvent has evaporated, a bilayer is formed. Then, the solution / air interface on both sides of the opening is physically moved up and down beyond the opening. Lipid bilayers may form across the openings of the membrane, or across the openings in recesses. It can also form within something.
[0453] The Montal and Müller method is suitable for protein pore insertion into high-quality lipid bilayers. It is a cost-effective and relatively easy method of forming, and therefore has a good reputation. The methods for forming the multilayer include tip immersion, coating of the double layer, and patch clamping of the lipid bilayer. These are some examples.
[0454] The tip immersion bilayer formation involves the opening surface (e.g.) on the surface of the test solution supporting the lipid monolayer. For example, it is necessary to touch the tip of the pipette to the solvent. In this case as well, it is dissolved in an organic solvent. The lipid monolayer is first removed from the solution / air interface by evaporating the dissolved lipid droplets at the solution surface. It is then generated by Langmuir-Schaefer. It is formed by a method and requires mechanical automation to move the opening relative to the solution surface.
[0455] For the double layer to be painted, droplets of lipids dissolved in an organic solvent are applied directly to the opening. The opening is then submerged in the test aqueous solution. The lipid solution is applied using a paint brush or equivalent. It is then thinly coated over the entire opening. As the solvent becomes thinner, a lipid bilayer is formed. Yes. However, complete removal of the solvent from the bilayer is difficult, and therefore, this method is not effective. The resulting bilayer has lower stability and is more prone to generating noise during electrochemical measurements. .
[0456] Patch clamps are commonly used in the study of biological cell membranes. The cell membrane is then pipetized by aspiration. The end of the net is clamped, and the membrane patch is attached to the entire opening. This method, The liposomes are clamped (they then rupture and seal the entire opening of the pipette). This method is suitable for the formation of lipid bilayers by leaving behind a lipid bilayer. This requires the creation of large, monolayered liposomes and small openings in materials with a glass surface. do.
[0457] Liposomes are extracted by sonication, extraction, or the Mozafari method (Colas et al., (2 007) It can be formed by Micron 38:841-847).
[0458] In a preferred embodiment, the lipid bilayer is provided in International Application No. PCT / GB08 / 00412 As described in Specification No. 7 (published as International Publication No. 2009 / 077734) It is formed in such a way. Advantageously, in this method, the lipid bilayer is formed from dry lipids. In the most preferred embodiment, the lipid bilayer is as described in International Publication No. 2009 / 077734 Pan As described in the fret (PCT / GB08 / 004127), the shape crosses the opening. It will be accomplished.
[0459] The lipid bilayer is formed from two opposing lipid layers. The two lipid layers are hydrophobic. The hydrophilic tail groups are arranged to face each other, forming a hydrophobic interior. The head groups are oriented outward toward the aqueous environment on both sides of the bilayer. These are not limited to, but include lipid disorder phase (fluid layer), liquid disorder phase, and solid disorder phase. (Layered gel phase, interdigitated gel phase) and planar bilayer crystals (layered subgels) It can be found in numerous lipid phases, including layered crystalline phases.
[0460] Any lipid composition that forms a lipid bilayer can be used. Key properties, such as surface charge, ability to support membrane proteins, packing density, or mechanical properties. A lipid bilayer having is selected to be formed. The lipid composition is one or more different It may contain lipids such as 100 or less. This can be done. The lipid composition preferably contains 1 to 10 lipids. The lipid composition is natural It may contain lipids present in and / or artificial lipids.
[0461] Lipids typically consist of two parts, a head group and an interface, which may be the same or different. It includes a hydrophobic tail group. Suitable head groups include neutral head groups, such as diacylg Ricerides (DG) and ceramides (CM), zwitterionic head groups, such as phosphatidyl Lucoline (PC), phosphatidylethanolamine (PE), and sphingomyelin (SM), a head group having a negative charge, for example, phosphatidylglycerol (PG), phosphatidylglycerol Phatidylserine (PS), phosphatidylinositol (PI), phosphate (PA) and Cardiolipin (CA), and positively charged head groups, such as trimethylammonium Examples include, but are not limited to, nium-propane (TAP). Suitable interface portion and Examples include naturally occurring interface parts, such as glycerol-based or ceramide-based parts. However, it is not limited to these. Suitable hydrophobic tail groups include saturated hydrocarbon chains, for example. For example, lauric acid (n-dodecanoic acid), myristic acid (n-tetradecanoic acid), palmitin Acid (n-hexadecanoic acid), stearic acid (n-octadecanoic acid) and arachidonic acid ( n-eicosanoic acid), unsaturated hydrocarbon chains, such as oleic acid (cis-9-octadecanoic acid) ), as well as branched hydrocarbon chains, such as phytanoyl, but not limited to these. i. The chain length of an unsaturated hydrocarbon chain and the position and number of double bonds in the unsaturated hydrocarbon chain are: It can vary. This includes the chain length of branched hydrocarbon chains and the branching of methyl groups within branched hydrocarbon chains. The position and number of these groups can vary. Hydrophobic tail groups are placed at the interface with ether or ester. It can be linked together as a link. The lipid may be mycolic acid.
[0462] Lipids can also be chemically modified. The head group or tail group of a lipid can be chemically modified. They may be decorated. Suitable lipids with chemically modified head groups include PEG-modified lipids. Decorative lipids, for example, 1,2-diacyl-sn-glycero-3-phosphoethanolamine-N- [Methoxy(polyethylene glycol)-2000], functionalized PEG lipids, e.g., 1,2 -Distearoyl-sn-glycero-3-phosphoethanolamine-N-[biotinyl( Polyethylene glycol (2000), as well as lipids modified for binding, e.g., 1,2 -Dioleil-sn-glycero-3-phosphoethanolamine-N-(succinyl) 1,2-Dipalmitoyl-sn-glycero-3-phosphoethanolamine-N-(bio Examples include, but are not limited to, cynyl. Suitable lipids include polymerizable lipids, such as 1,2-bis(10,12-tricosadiinoyl). )-sn-glycero-3-phosphocholine, fluorinated lipids, e.g., 1-palmitoyl-2- (16-Fluoropalmitoyl)-sn-glycero-3-phosphocholine, denatured lipid, for example ba 1,2-dipalmitoyl-D62-sn-glycero-3-phosphocholine, and A Tel-linked lipids, such as 1,2-di-O-phytanyl-sn-glycero-3-phosphocholine. These include, but are not limited to, lipids assist in the coupling of polynucleotides. They may also be chemically modified or functionalized to enhance their properties.
[0463] The amphiphilic layer, for example, a lipid composition, is typically one or more that affect the properties of the layer. Contains additives. Suitable additives include fatty acids, such as palmitic acid, myristic acid, and Oleic acid, fatty alcohols, such as palmitic acid alcohol, myristate alcohol Alcohols and oleic acid alcohols, sterols, such as cholesterol and ergosterols. lanosterol, sitosterol and stigmasterol, lysophospholipids, for example 1 Acyl-2-hydroxy-sn-glycero-3-phosphocholine and ceramides are among the examples. These are some examples, but they are not limited to these.
[0464] In another preferred embodiment, the film includes a solid layer. The solid layer is not limited to these. There are no microelectronic materials, insulating materials, such as Si3N4, Al2O3 and SiO, organic and inorganic polymers, such as polyamides, plastics, such as Tef lon® or elastomer, such as two-component addition-cured silicone rubber, and It can be formed from both organic and inorganic materials, including glass. The solid layer is graphite It may also be formed from n. Suitable graphene layers are described in International Patent Application No. PCT / US2008. Specification No. / 010637 (published as International Publication No. 2009 / 035647) ) is disclosed. When the film includes a solid layer, the pore is usually within the solid layer, for example, amphiphilic membranes contained within holes, walls, gaps, channels, trenches, and slits in solid layers or It is present in the layer. Those skilled in the art can prepare a suitable solid / amphiphilic hybrid system. The preferred systems are International Publication No. 2009 / 020682 and International Publication No. 2 This is disclosed in pamphlet number 012 / 005857. The amphiphilic membrane or layer discussed above. You can use any of the following.
[0465] The method typically involves (i) an artificial amphiphilic layer containing pores, and (ii) an isolated self containing pores. This is performed using a naturally occurring lipid bilayer, or (iii) cells in which a pore has been inserted. The method typically involves using an artificial amphiphilic layer, such as an artificial triblock copolymer layer. The layer adds other transmembrane and / or intramembrane proteins and other molecules to the pore. This may include. Preferred apparatus and conditions are discussed below. The method of the present invention is usually in This procedure is performed in vitro.
[0466] Coupling The polynucleotide is preferably coupled with a membrane containing a pore. The method is as follows: The process may include a step of coup...
Claims
1. An apparatus for molecular sensing comprising a transmembrane protein pore inserted into a membrane in vitro, wherein the transmembrane protein pore comprises at least one CsgG monomer having amino acid mutations at two or more positions corresponding to Y51, N55 and F56 of the sequence shown in Sequence ID No.
390.
2. The apparatus according to claim 1, wherein the at least one CsgG monomer includes a mutation at a position corresponding to Y51 / N55, Y51 / F56, N55 / F56, or Y51 / N55 / F56.
3. The at least one CsgG monomer is F56N / N55Q, F56N / N55R, F56N / N55K, F56N / N55S, F56N / N55G, F56N / N55A, F56N / N55T, F56Q / N55Q, F56Q / N55R, F56Q / N55K, F56Q / N55S, F56Q / N55G, F56Q / N55A, F56Q / N55T, F56R / N55Q, F56R / N55R, F56R / N55K, F56R / N55S, F56R / N55G, F56R / N55A, F56R / N55T, F56S / N55Q, F56S / N55R, F56S / N55K, F56S / N55S, F56S / N55G, F56S / N55A, F56S / N55T, F56G / N55Q, F5 6G / N55R, F56G / N55K, F56G / N55S, F56G / N55G, F56G / N55A, F56G / N55T, F56A / N55Q, F56A / N55R, F56A / N55K, F56A / N55S, F56A / N55G, F56A / N55A, F56A / N5 5T, F56K / N55Q, F56K / N55R, F56K / N55K, F56K / N55S, F56K / N55G, F56K / N55A, F56K / N55T, F56N / Y51L, F56N / Y51V, F56N / Y51A, F56N / Y51N, F56N / Y51Q, F5 6N / Y51S, F56N / Y51G, F56Q / Y51L, F56Q / Y51V, F56Q / Y51A, F56Q / Y51N, F56Q / Y51Q, F56Q / Y51S, F56Q / Y51G, F56R / Y51L, F56R / Y51V, F56R / Y51A, F56R / Y5 1N, F56R / Y51Q, F56R / Y51S, F56R / Y51G, F56S / Y51L, F56S / Y51V, F56S / Y51A, F56S / Y51N, F56S / Y51Q, F56S / Y51S, F56S / Y51G, F56G / Y51L, F56G / Y51V, F5 6G / Y51A, F56G / Y51N, F56G / Y51Q, F56G / Y51S, F56G / Y51G, F56A / Y51L, F56A / Y51V, F56A / Y51A, F56A / Y51N, F56A / Y51Q, F56A / Y51S, F56A / Y51G, F56K / Y5 1L, F56K / Y51V, F56K / Y51A, F56K / Y51N, F56K / Y51Q, F56K / Y51S, F56K / Y51G,N55Q / Y51L、N55Q / Y51V、N55Q / Y51A、N55Q / Y51N、N55Q / Y51Q、N55Q / Y51S、N55Q / Y51G、N55R / Y51L、N55R / Y51V、N55R / Y51A、N55R / Y51N、N55R / Y51Q、N55R / Y51S、N55R / Y51G、N55K / Y51L、N55K / Y51V、N55K / Y51A、N55K / Y51N、N55K / Y51Q、N55K / Y51S、N55K / Y51G、N55S / Y51L、N55S / Y51V、N55S / Y51A、N55S / Y51N、N55S / Y51Q、N55S / Y51S、N55S / Y51G、N55G / Y51L、N55G / Y51V、N55G / Y51A、N55G / Y51N、N55G / Y51Q、N55G / Y51S、N55G / Y51G、N55A / Y51L、N55A / Y51V、N55A / Y51A、N55A / Y51N、N55A / Y51Q、N55A / Y51S、N55A / Y51G、N55T / Y51L、N55T / Y51V、N55T / Y51A、N55T / Y51N、N55T / Y51Q、N55T / Y51S、N55T / Y51G、F56N / N55Q / Y51L、F56N / N55Q / Y51V、F56N / N55Q / Y51A、F56N / N55Q / Y51N、F56N / N55Q / Y51Q、F56N / N55Q / Y51S、F56N / N55Q / Y51G、F56N / N55R / Y51L、F56N / N55R / Y51V、F56N / N55R / Y51A、F56N / N55R / Y51N、F56N / N55R / Y51Q、F56N / N55R / Y51S、F56N / N55R / Y51G、F56N / N55K / Y51L、F56N / N55K / Y51V、F56N / N55K / Y51A、F56N / N55K / Y51N、F56N / N55K / Y51Q、F56N / N55K / Y51S、F56N / N55K / Y51G、F56N / N55S / Y51L、F56N / N55S / Y51V、F56N / N55S / Y51A、F56N / N55S / Y51N、F56N / N55S / Y51Q、F56N / N55S / Y51S、F56N / N55S / Y51G、F56N / N55G / Y51L、F56N / N55G / Y51V、F56N / N55G / Y51A、F56N / N55G / Y51N、F56N / N55G / Y51Q、F56N / N55G / Y51S、F56N / N55G / Y51G、F56N / N55A / Y51L、F56N / N55A / Y51V、F56N / N55A / Y51A、F56N / N55A / Y51N、F56N / N55A / Y51Q、F56N / N55A / Y51S、F56N / N55A / Y51G、F56N / N55T / Y51L、F56N / N55T / Y51V、F56N / N55T / Y51A、F56N / N55T / Y51N、F56N / N55T / Y51Q、F56N / N55T / Y51S、F56N / N55T / Y51G、F56Q / N55Q / Y51L、F56Q / N55Q / Y51V、F56Q / N55Q / Y51A、F56Q / N55Q / Y51N、F56Q / N55Q / Y51Q、F56Q / N55Q / Y51S、F56Q / N55Q / Y51G、F56Q / N55R / Y51L、F56Q / N55R / Y51V、F56Q / N55R / Y51A、F56Q / N55R / Y51N、F56Q / N55R / Y51Q、F56Q / N55R / Y51S、F56Q / N55R / Y51G、F56Q / N55K / Y51L、F56Q / N55K / Y51V、F56Q / N55K / Y51A、F56Q / N55K / Y51N、F56Q / N55K / Y51Q、F56Q / N55K / Y51S、F56Q / N55K / Y51G、F56Q / N55S / Y51L、F56Q / N55S / Y51V、F56Q / N55S / Y51A、F56Q / N55S / Y51N、F56Q / N55S / Y51Q、F56Q / N55S / Y51S、F56Q / N55S / Y51G、F56Q / N55G / Y51L、F56Q / N55G / Y51V、F56Q / N55G / Y51A、F56Q / N55G / Y51N、F56Q / N55G / Y51Q、F56Q / N55G / Y51S、F56Q / N55G / Y51G、F56Q / N55A / Y51L、F56Q / N55A / Y51V、F56Q / N55A / Y51A、F56Q / N55A / Y51N、F56Q / N55A / Y51Q、F56Q / N55A / Y51S、F56Q / N55A / Y51G、F56Q / N55T / Y51L、F56Q / N55T / Y51V、F56Q / N55T / Y51A、F56Q / N55T / Y51N、F56Q / N55T / Y51Q、F56Q / N55T / Y51S、F56Q / N55T / Y51G、F56R / N55Q / Y51L、F56R / N55Q / Y51V、F56R / N55Q / Y51A、F56R / N55Q / Y51N、F56R / N55Q / Y51Q、F56R / N55Q / Y51S、F56R / N55Q / Y51G、F56R / N55R / Y51L、F56R / N55R / Y51V、F56R / N55R / Y51A、F56R / N55R / Y51N、F56R / N55R / Y51Q、F56R / N55R / Y51S、F56R / N55R / Y51G、F56R / N55K / Y51L、F56R / N55K / Y51V、F56R / N55K / Y51A、F56R / N55K / Y51N、F56R / N55K / Y51Q、F56R / N55K / Y51S、F56R / N55K / Y51G、F56R / N55S / Y51L、F56R / N55S / Y51V、F56R / N55S / Y51A、F56R / N55S / Y51N、F56R / N55S / Y51Q、F56R / N55S / Y51S、F56R / N55S / Y51G、F56R / N55G / Y51L、F56R / N55G / Y51V、F56R / N55G / Y51A、F56R / N55G / Y51N、F56R / N55G / Y51Q、F56R / N55G / Y51S、F56R / N55G / Y51G、F56R / N55A / Y51L、F56R / N55A / Y51V、F56R / N55A / Y51A、F56R / N55A / Y51N、F56R / N55A / Y51Q、F56R / N55A / Y51S、F56R / N55A / Y51G、F56R / N55T / Y51L、F56R / N55T / Y51V、F56R / N55T / Y51A、F56R / N55T / Y51N、F56R / N55T / Y51Q、F56R / N55T / Y51S、F56R / N55T / Y51G、F56S / N55Q / Y51L、F56S / N55Q / Y51V、F56S / N55Q / Y51A、F56S / N55Q / Y51N、F56S / N55Q / Y51Q、F56S / N55Q / Y51S、F56S / N55Q / Y51G、F56S / N55R / Y51L、F56S / N55R / Y51V、F56S / N55R / Y51A、F56S / N55R / Y51N、F56S / N55R / Y51Q、F56S / N55R / Y51S、F56S / N55R / Y51G、F56S / N55K / Y51L、F56S / N55K / Y51V、F56S / N55K / Y51A、F56S / N55K / Y51N、F56S / N55K / Y51Q、F56S / N55K / Y51S、F56S / N55K / Y51G、F56S / N55S / Y51L、F56S / N55S / Y51V、F56S / N55S / Y51A、F56S / N55S / Y51N、F56S / N55S / Y51Q、F56S / N55S / Y51S、F56S / N55S / Y51G、F56S / N55G / Y51L、F56S / N55G / Y51V、F56S / N55G / Y51A、F56S / N55G / Y51N、F56S / N55G / Y51Q、F56S / N55G / Y51S、F56S / N55G / Y51G、F56S / N55A / Y51L、F56S / N55A / Y51V、F56S / N55A / Y51A、F56S / N55A / Y51N、F56S / N55A / Y51Q、F56S / N55A / Y51S、F56S / N55A / Y51G、F56S / N55T / Y51L、F56S / N55T / Y51V、F56S / N55T / Y51A、F56S / N55T / Y51N、F56S / N55T / Y51Q、F56S / N55T / Y51S、F56S / N55T / Y51G、F56G / N55Q / Y51L、F56G / N55Q / Y51V、F56G / N55Q / Y51A、F56G / N55Q / Y51N、F56G / N55Q / Y51Q、F56G / N55Q / Y51S、F56G / N55Q / Y51G、F56G / N55R / Y51L、F56G / N55R / Y51V、F56G / N55R / Y51A、F56G / N55R / Y51N、F56G / N55R / Y51Q、F56G / N55R / Y51S、F56G / N55R / Y51G、F56G / N55K / Y51L、F56G / N55K / Y51V、F56G / N55K / Y51A、F56G / N55K / Y51N、F56G / N55K / Y51Q、F56G / N55K / Y51S、F56G / N55K / Y51G、F56G / N55S / Y51L、F56G / N55S / Y51V、F56G / N55S / Y51A、F56G / N55S / Y51N、F56G / N55S / Y51Q、F56G / N55S / Y51S、F56G / N55S / Y51G、F56G / N55G / Y51L、F56G / N55G / Y51V、F56G / N55G / Y51A、F56G / N55G / Y51N、F56G / N55G / Y51Q、F56G / N55G / Y51S、F56G / N55G / Y51G、F56G / N55A / Y51L、F56G / N55A / Y51V、F56G / N55A / Y51A、F56G / N55A / Y51N、F56G / N55A / Y51Q、F56G / N55A / Y51S、F56G / N55A / Y51G、F56G / N55T / Y51L、F56G / N55T / Y51V、F56G / N55T / Y51A、F56G / N55T / Y51N、 F56G / N55T / Y51Q、F56G / N55T / Y51S、F56G / N55T / Y51G、F56A / N55Q / Y51L、F56A / N55Q / Y51V、F56A / N55Q / Y51A、F56A / N55Q / Y51N、F56A / N55Q / Y51Q、F56A / N55Q / Y51S、F56A / N55Q / Y51G、F56A / N55R / Y51L、F56A / N55R / Y51V、F56A / N55R / Y51A、F56A / N55R / Y51N、F56A / N55R / Y51Q、F56A / N55R / Y51S、F56A / N55R / Y51G、F56A / N55K / Y51L、F56A / N55K / Y51V、F56A / N55K / Y51A、F56A / N55K / Y51N、F56A / N55K / Y51Q、F56A / N55K / Y51S、F56A / N55K / Y51G、F56A / N55S / Y51L、F56A / N55S / Y51V、F56A / N55S / Y51A、F56A / N55S / Y51N、F56A / N55S / Y51Q、F56A / N55S / Y51S、F56A / N55S / Y51G、F56A / N55G / Y51L、F56A / N55G / Y51V、F56A / N55G / Y51A、F56A / N55G / Y51N、F56A / N55G / Y51Q、F56A / N55G / Y51S、F56A / N55G / Y51G、F56A / N55A / Y51L、F56A / N55A / Y51V、F56A / N55A / Y51A、F56A / N55A / Y51N、F56A / N55A / Y51Q、F56A / N55A / Y51S、F56A / N55A / Y51G、F56A / N55T / Y51L、F56A / N55T / Y51V、F56A / N55T / Y51A、F56A / N55T / Y51N、F56A / N55T / Y51Q、F56A / N55T / Y51S、F56A / N55T / Y51G、F56K / N55Q / Y51L、F56K / N55Q / Y51V、F56K / N55Q / Y51A、F56K / N55Q / Y51N、F56K / N55Q / Y51Q、F56K / N55Q / Y51S、F56K / N55Q / Y51G、F56K / N55R / Y51L、F56K / N55R / Y51V、F56K / N55R / Y51A、F56K / N55R / Y51N、F56K / N55R / Y51Q、F56K / N55R / Y51S、F56K / N55R / Y51G、F56K / N55K / Y51L, F56K / N55K / Y51V, F56K / N55K / Y51A, F56K / N55K / Y51N, F56K / N55K / Y51Q, F56 K / N55K / Y51S, F56K / N55K / Y51G, F56K / N55S / Y51L, F56K / N55S / Y51V, F56K / N55S / Y51A, F56K / N5 5S / Y51N, F56K / N55S / Y51Q, F56K / N55S / Y51S, F56K / N55S / Y51G, F56K / N55G / Y51L, F56K / N55G / Y51V, F56K / N55G / Y51A, F56K / N55G / Y51N, F56K / N55G / Y51Q, F56K / N55G / Y51S, F56K / N55G / Y51G , F56K / N55A / Y51L, F56K / N55A / Y51V, F56K / N55A / Y51A, F56K / N55A / Y51N, F56K / N55A / Y51Q, F5 6K / N55A / Y51S, F56K / N55A / Y51G, F56K / N55T / Y51L, F56K / N55T / Y51V, F56K / N55T / Y51A, F56K / N The apparatus according to claim 1, comprising mutations corresponding to 55T / Y51N, F56K / N55T / Y51Q, F56K / N55T / Y51S, F56K / N55T / Y51G, F56E / N55R, F56E / N55K, F56D / N55R, F56D / N55K, F56R / N55E, F56R / N55D, F56K / N55E, or F56K / N55D.
4. The apparatus according to claim 1, wherein the transmembrane protein pore is a heterooligomeric pore.
5. The apparatus according to claim 4, wherein the heterooligomer pore comprises nine CsgG monomers, and at least one of the nine CsgG monomers is different from the other CsgG monomers.
6. The apparatus according to claim 1, wherein the transmembrane protein pore is a homooligomeric pore.
7. The apparatus according to claim 6, wherein the homooligomeric pore consists of nine identical CsgG monomers.
8. A transmembrane protein pore comprising at least one CsgG monomer, wherein the CsgG monomer comprises amino acid mutations at positions corresponding to Y51 and F56 of the sequence shown in Sequence ID No.
390.
9. The pore according to claim 8, further comprising a variation at the position corresponding to N55.
10. The at least one CsgG monomer is F56N / Y51L, F56N / Y51V, F56N / Y51A, F56N / Y51N, F56N / Y51Q, F56N / Y51S, F56N / Y51G, F56Q / Y51L, F56Q / Y51V, F56Q / Y51A, F56Q / Y51N, F56 Q / Y51Q, F56Q / Y51S, F56Q / Y51G, F56R / Y51L, F56R / Y51V, F56R / Y51A, F56R / Y51N, F56R / Y51Q, F56R / Y51S, F56R / Y51G, F56S / Y51L, F56S / Y51V, F56S / Y51A, F56S / Y5 1N, F56S / Y51Q, F56S / Y51S, F56S / Y51G, F56G / Y51L, F56G / Y51V, F56G / Y51A, F56 G / Y51N, F56G / Y51Q, F56G / Y51S, F56G / Y51G, F56A / Y51L, F56A / Y51V, F56A / Y51A, The pore according to claim 8, comprising mutations corresponding to F56A / Y51N, F56A / Y51Q, F56A / Y51S, F56A / Y51G, F56K / Y51L, F56K / Y51V, F56K / Y51A, F56K / Y51N, F56K / Y51Q, F56K / Y51S, or F56K / Y51G.
11. The pore according to claim 8, wherein the pore is a heterooligomeric pore.
12. The pore according to claim 8, wherein the heterooligomeric pore comprises nine CsgG monomers, and at least one of the nine CsgG monomers is different from the other CsgG monomers.
13. The pore according to claim 8, wherein the transmembrane protein pore is a homo-oligomeric pore.
14. The pore according to claim 13, wherein the homooligomeric pore consists of nine identical CsgG monomers.
15. (i) To obtain a transmembrane protein pore comprising at least one CsgG monomer, wherein the CsgG monomer has amino acid mutations at one or more positions corresponding to Y51, N55 and F56 of the sequence shown in Sequence ID No. 390, (ii) Bringing the pore into contact with the membrane in vitro so that the pore is inserted into the membrane in vitro, A device for molecular sensing, manufactured by a method including [a specific method].
16. The at least one CsgG monomer is F56N / N55Q / Y51L, F56N / N55Q / Y51V, F56N / N55Q / Y51A, F56N / N55Q / Y51N, F56N / N55Q / Y51Q, F56N / N55Q / Y51S, F56N / N55Q / Y51G, F56N / N55R / Y51L, F56N / N55R / Y51V, F56N / N55R / Y51A, F56N / N55R / Y51N, F56N / N55R / Y51Q, F56N / N55R / Y51S, F56N / N55R / Y51G, F56N / N55K / Y51L, F56 N / N55K / Y51V, F56N / N55K / Y51A, F56N / N55K / Y51N, F56N / N55K / Y51Q, F56N / N55K / Y51S, F56N / N55K / Y51G, F56N / N55S / Y51L, F56N / N55S / Y51V, F56N / N5 5S / Y51A, F56N / N55S / Y51N, F56N / N55S / Y51Q, F56N / N55S / Y51S, F56N / N55S / Y51G, F56N / N55G / Y51L, F56N / N55G / Y51V, F56N / N55G / Y51A, F56N / N55G / Y5 1N, F56N / N55G / Y51Q, F56N / N55G / Y51S, F56N / N55G / Y51G, F56N / N55A / Y51L , F56N / N55A / Y51V, F56N / N55A / Y51A, F56N / N55A / Y51N, F56N / N55A / Y51Q, F 56N / N55A / Y51S, F56N / N55A / Y51G, F56N / N55T / Y51L, F56N / N55T / Y51V, F56 N / N55T / Y51A, F56N / N55T / Y51N, F56N / N55T / Y51Q, F56N / N55T / Y51S, F56N / N 55T / Y51G, F56Q / N55Q / Y51L, F56Q / N55Q / Y51V, F56Q / N55Q / Y51A, F56Q / N55 Q / Y51N, F56Q / N55Q / Y51Q, F56Q / N55Q / Y51S, F56Q / N55Q / Y51G, F56Q / N55R / Y51L, F56Q / N55R / Y51V, F56Q / N55R / Y51A, F56Q / N55R / Y51N, F56Q / N55R / Y5 1Q, F56Q / N55R / Y51S, F56Q / N55R / Y51G, F56Q / N55K / Y51L, F56Q / N55K / Y51V,F56Q / N55K / Y51A、F56Q / N55K / Y51N、F56Q / N55K / Y51Q、F56Q / N55K / Y51S、F56Q / N55K / Y51G、F56Q / N55S / Y51L、F56Q / N55S / Y51V、F56Q / N55S / Y51A、F56Q / N55S / Y51N、F56Q / N55S / Y51Q、F56Q / N55S / Y51S、F56Q / N55S / Y51G、F56Q / N55G / Y51L、F56Q / N55G / Y51V、F56Q / N55G / Y51A、F56Q / N55G / Y51N、F56Q / N55G / Y51Q、F56Q / N55G / Y51S、F56Q / N55G / Y51G、F56Q / N55A / Y51L、F56Q / N55A / Y51V、F56Q / N55A / Y51A、F56Q / N55A / Y51N、F56Q / N55A / Y51Q、F56Q / N55A / Y51S、F56Q / N55A / Y51G、F56Q / N55T / Y51L、F56Q / N55T / Y51V、F56Q / N55T / Y51A、F56Q / N55T / Y51N、F56Q / N55T / Y51Q、F56Q / N55T / Y51S、F56Q / N55T / Y51G、F56R / N55Q / Y51L、F56R / N55Q / Y51V、F56R / N55Q / Y51A、F56R / N55Q / Y51N、F56R / N55Q / Y51Q、F56R / N55Q / Y51S、F56R / N55Q / Y51G、F56R / N55R / Y51L、F56R / N55R / Y51V、F56R / N55R / Y51A、F56R / N55R / Y51N、F56R / N55R / Y51Q、F56R / N55R / Y51S、F56R / N55R / Y51G、F56R / N55K / Y51L、F56R / N55K / Y51V、F56R / N55K / Y51A、F56R / N55K / Y51N、F56R / N55K / Y51Q、F56R / N55K / Y51S、F56R / N55K / Y51G、F56R / N55S / Y51L、F56R / N55S / Y51V、F56R / N55S / Y51A、F56R / N55S / Y51N、F56R / N55S / Y51Q、F56R / N55S / Y51S、F56R / N55S / Y51G、F56R / N55G / Y51L、F56R / N55G / Y51V、F56R / N55G / Y51A、F56R / N55G / Y51N、F56R / N55G / Y51Q、F56R / N55G / Y51S、F56R / N55G / Y51G、F56R / N55A / Y51L、F56R / N55A / Y51V、F56R / N55A / Y51A、F56R / N55A / Y51N、F56R / N55A / Y51Q、F56R / N55A / Y51S、F56R / N55A / Y51G、F56R / N55T / Y51L、F56R / N55T / Y51V、F56R / N55T / Y51A、F56R / N55T / Y51N、F56R / N55T / Y51Q、F56R / N55T / Y51S、F56R / N55T / Y51G、F56S / N55Q / Y51L、F56S / N55Q / Y51V、F56S / N55Q / Y51A、F56S / N55Q / Y51N、F56S / N55Q / Y51Q、F56S / N55Q / Y51S、F56S / N55Q / Y51G、F56S / N55R / Y51L、F56S / N55R / Y51V、F56S / N55R / Y51A、F56S / N55R / Y51N、F56S / N55R / Y51Q、F56S / N55R / Y51S、F56S / N55R / Y51G、F56S / N55K / Y51L、F56S / N55K / Y51V、F56S / N55K / Y51A、F56S / N55K / Y51N、F56S / N55K / Y51Q、F56S / N55K / Y51S、F56S / N55K / Y51G、F56S / N55S / Y51L、F56S / N55S / Y51V、F56S / N55S / Y51A、F56S / N55S / Y51N、F56S / N55S / Y51Q、F56S / N55S / Y51S、F56S / N55S / Y51G、F56S / N55G / Y51L、F56S / N55G / Y51V、F56S / N55G / Y51A、F56S / N55G / Y51N、F56S / N55G / Y51Q、F56S / N55G / Y51S、F56S / N55G / Y51G、F56S / N55A / Y51L、F56S / N55A / Y51V、F56S / N55A / Y51A、F56S / N55A / Y51N、F56S / N55A / Y51Q、F56S / N55A / Y51S、F56S / N55A / Y51G、F56S / N55T / Y51L、F56S / N55T / Y51V、F56S / N55T / Y51A、F56S / N55T / Y51N、F56S / N55T / Y51Q、F56S / N55T / Y51S、F56S / N55T / Y51G、F56G / N55Q / Y51L、F56G / N55Q / Y51V、F56G / N55Q / Y51A、F56G / N55Q / Y51N、F56G / N55Q / Y51Q、F56G / N55Q / Y51S、F56G / N55Q / Y51G、F56G / N55R / Y51L、F56G / N55R / Y51V、F56G / N55R / Y51A、F56G / N55R / Y51N、F56G / N55R / Y51Q、F56G / N55R / Y51S、F56G / N55R / Y51G、F56G / N55K / Y51L、F56G / N55K / Y51V、F56G / N55K / Y51A、F56G / N55K / Y51N、F56G / N55K / Y51Q、F56G / N55K / Y51S、F56G / N55K / Y51G、F56G / N55S / Y51L、F56G / N55S / Y51V、F56G / N55S / Y51A、F56G / N55S / Y51N、F56G / N55S / Y51Q、F56G / N55S / Y51S、F56G / N55S / Y51G、F56G / N55G / Y51L、F56G / N55G / Y51V、F56G / N55G / Y51A、F56G / N55G / Y51N、F56G / N55G / Y51Q、F56G / N55G / Y51S、F56G / N55G / Y51G、F56G / N55A / Y51L、F56G / N55A / Y51V、F56G / N55A / Y51A、F56G / N55A / Y51N、F56G / N55A / Y51Q、F56G / N55A / Y51S、F56G / N55A / Y51G、F56G / N55T / Y51L、F56G / N55T / Y51V、F56G / N55T / Y51A、F56G / N55T / Y51N、F56G / N55T / Y51Q、F56G / N55T / Y51S、F56G / N55T / Y51G、F56A / N55Q / Y51L、F56A / N55Q / Y51V、F56A / N55Q / Y51A、F56A / N55Q / Y51N、F56A / N55Q / Y51Q、F56A / N55Q / Y51S、F56A / N55Q / Y51G、F56A / N55R / Y51L、F56A / N55R / Y51V、F56A / N55R / Y51A、F56A / N55R / Y51N、F56A / N55R / Y51Q、F56A / N55R / Y51S、F56A / N55R / Y51G、F56A / N55K / Y51L、F56A / N55K / Y51V、F56A / N55K / Y51A、F56A / N55K / Y51N、F56A / N55K / Y51Q、F56A / N55K / Y51S、F56A / N55K / Y51G、F56A / N55S / Y51L、F56A / N55S / Y51V、F56A / N55S / Y51A、F56A / N55S / Y51N、F56A / N55S / Y51Q、F56A / N55S / Y51S、F56A / N55S / Y51G、F56A / N55G / Y51L、F56A / N55G / Y51V、F56A / N55G / Y51A、F56A / N55G / Y51N、F56A / N55G / Y51Q、F56A / N55G / Y51S、F56A / N55G / Y51G、F56A / N55A / Y51L、F56A / N55A / Y51V、F56A / N55A / Y51A、F56A / N55A / Y51N、F56A / N55A / Y51Q、F56A / N55A / Y51S、F56A / N55A / Y51G、F56A / N55T / Y51L、F56A / N55T / Y51V、F56A / N55T / Y51A、F56A / N55T / Y51N、F56A / N55T / Y51Q、F56A / N55T / Y51S、F56A / N55T / Y51G、F56K / N55Q / Y51L、F56K / N55Q / Y51V、F56K / N55Q / Y51A、F56K / N55Q / Y51N、F56K / N55Q / Y51Q、F56K / N55Q / Y51S、F56K / N55Q / Y51G、F56K / N55R / Y51L、F56K / N55R / Y51V、F56K / N55R / Y51A、F56K / N55R / Y51N、F56K / N55R / Y51Q、F56K / N55R / Y51S、F56K / N55R / Y51G、F56K / N55K / Y51L、F56K / N55K / Y51V、F56K / N55K / Y51A、F56K / N55K / Y51N、F56K / N55K / Y51Q、F56K / N55K / Y51S、F56K / N55K / Y51G、F56K / N55S / Y51L、F56K / N55S / Y51V、F56K / N55S / Y51A、F56K / N55S / Y51N、F56K / N55S / Y51Q、F56K / N55S / Y51S、F56K / N55S / Y51G、F56K / N55G / Y51L、F56K / N55G / Y51V、F56K / N55G / Y51A、F56K / N55G / Y51N、F56K / N55G / Y51Q、F56K / N55G / Y51S、F56K / N55G / Y51G、F56K / N55A / Y51L、F56K / N55A / Y51V、F56K / N55A / Y51A、F56K / N55A / Y51N、F56K / N55A / Y51Q、F56K / N55A / Y51S、F56K / N55A / Y51G、F56K / N55T / Y51L、F56K / N55T / Y51V、F56K / N55T / Y51A、F56K / N55T / Y51N、 The pore according to claim 9, comprising a mutation corresponding to F56K / N55T / Y51Q, F56K / N55T / Y51S, or F56K / N55T / Y51G.